Further Thoughts
Table of Contents
Parameter Estimation
Let us look at the distribution we want to estimate , in this case, is not the distribution average but a different parameter of the distribution. are the observations again, two approaches seem natural:
- using the relation between the expected value of this distribution and the parameter: , we then might estimate with
- we could also try to estimate by the maximum of values we have seen so far:
We know that converges almost surely to b and has expected value , so lets look at :
for .
Thus the density of is for
and we get the expected value:
So is not unbiased, however is and for large .
We also find:
which means . Since is a constant we know that this means . Where “” means convergence in distribution and “” convergence in probability. For a quick introduction of these concepts see Convergence of random variables, it will not be necessary for this article, however.
In total we get , which means: , by -continuity of probability measures and thus
If the estimator converges in probability against the value it is estimating, we call the estimator consistent:
- and are consistent estimators for mean and variance respectively
- and are consistent estimators for the upper boundary of a uniform distribution interval.
The definition of convergence in probability is:
for a series of random variables and a random variable.
In our case, that means that no matter how small we allow the difference of or to to be, if we have enough observations the estimators are (almost) never further from than the allowed difference.This is weaker than almost sure convergence but, an important property to have for an estimator.
In summary, we now have:
- and
as unbiased and consistent estimators for . But which one is better? To answer this we might want to calculate their variances and choose the one with the variance that converges faster to zero:
the proof of which is left as an exercise. Now we see that we might want to choose for estimating in this case.
Choice of Significance Level
So far we always used a significance level around , in fact, this is the go-to choice in statistical testing. We would assume that this choice is somehow backed by some underlying fact of stochastic. It is not, it is more or less an arbitrary choice to quantify what we should consider plausible.
This does not mean that making any choice about in general or for a specific field is unreasonable and should be disregarded. In the chapter about hypothesis testing, I state that we should pick before collecting and testing our data. Having a general agreement about the value of within a field of scientific research ensures this is done in the right order. The reason this is important is, that there is more value for researchers in studies that can reject the zero hypothesis. So if a study barely misses the significance level, someone might just decide to increase it to publish a paper with significant results. Having a default value for prevents this.
There are a multitude of other ways to make data seem more significant, however. This is not a method to prevent all abuse of statistics, it just prevents this basic case!
p-Value interpretation
Often a very small p-value is interpreted as the zero hypothesis is very likely to be wrong and upon naive inspection of the meaning of the p-value we might think this is a good way to interpret it. Remember that the p-value is the probability that given the zero hypothesis, we get a less plausible observation than the observation we got.
Let us first deal with the word “likely” of this naive interpretation. The zero hypothesis is not a stochastic event and as such has no probability of occurring, this is why we have been talking about plausibility in the context of hypothesis tests: we are not calculating how likely is, rather we are trying to evaluate whether it is plausible!
Also, think about the following. With the p-value we are assuming that is true when calculating it, just to reject when it is small. If we were to be calculating some kind of probability for the likelihood of should we not have first picked the true probability measure? Instead, we pick a probability measure, the one that applies in the case of , just to conclude: “Well, that is not the right one.” This interpretation does not make sense. Instead, think of the p-value in terms of a proof by contradiction. The small p-value is the “contradiction” or in this case the implausibility. Assuming something false, the zero hypothesis, we have concluded something impossible, the small p-value!
What we are left with is the smaller the p-value the less plausible the zero hypothesis for our observations; if the p-value was zero there would be (almost) no way to get less plausible observations under zero hypothesis - there is (almost) no way to be more wrong about the zero hypothesis.
This is a good interpretation. The problem is when this is confused with the reality being vastly different than the zero hypothesis if the p-value is small. That is not the case!
Consider the following case that we will later verify in Python: we have an abundance of observations, they are all independent and stem from a -distribution. We use the t-test as introduced above to see if the observations have a mean smaller or equal to , since this is not the case we would want to see a very small p-value for enough observations. We would even like to see the p-value decrease as we increase the observations. After all the more data we collect the less plausible it should be for it to be produced by the wrong distribution. However, the distribution our data stems from really almost has a mean of zero so our zero hypothesis is not much different from reality.
But when does this even matter? Even in the case of only a small difference, we still correctly see an effect, for example, if this was some kind of medication, would we not like to see it prescribed even if it helps only a little bit?
Let us say we have a sleeping pill, a group of insomniac test subjects, and we know the average sleep duration per day of insomniacs exactly. We then give our sleeping pill to our test group, measure their sleeping duration and subtract the true average of insomniac sleep duration. Now it is sensible to test whether this change of sleep duration is greater than zero. In the situation as before we would conclude that the pill works. But if we told someone to buy a product that helps them with their sleep issues but only increases their sleep duration by hours, we probably would not sell many pills. The solution to this problem is measuring effect size rather than significance.
Let us test our intuition by implementing a simulation in Python:
We set up the imports and set the alpha level to . I decided to do simulations each in which I increase the observations in steps from to . Lastly, I prepare an array to calculate the average of the p-value for each step across all simulations.
import scipy.stats as st
import numpy as np
np.random.seed(234)
alpha = 0.05
simulations = 1000
steps = 5
observation_counts = np.linspace(10**4, 10**5, steps).astype(int)
avg_p = np.zeros(steps)Now for each simulation, we realize iid random variables in all_observations and then calculate S and p in a loop for the first observation_count observations of all_observations. Finally, we calculate the averages of the p values we got.
for _ in range(simulations):
all_observations = st.norm.rvs(size=np.max(observation_counts), loc=0.01, scale=1)
for i, observation_count in enumerate(observation_counts):
observations = all_observations[:observation_count]
emp_mean, emp_var = np.mean(observations), np.var(observations)
S = (emp_mean - 0) * np.sqrt(observation_count / emp_var)
p = 1 - stats.t.cdf(S, df = observation_count - 1)
avg_p[i] += p
avg_p = avg_p/simulationsThe following table displays the output.
| observation_count | avg_p |
|---|---|
| * | |
| * |
As expected the average p-value is decreasing, and our results become more significant for increasing observations. The * symbolizes when p is below where we can reject the zero hypothesis.
We can do a similar thing for the center confidence intervals. As we consider more observations the centered intervals get tighter around . For larger observations on average, we do not include in the interval.
| observation_count | center_ci |
|---|---|
| 10000 | |
| 32500 | |
| 55000 | |
| 77500 | |
| 100000 |
Of course, we can also approach this problem mathematically. Consider random variables with finite mean and variance , we know and . This means for any , and . Now consider:
The first case because of the central limit theorem and the second case because of the former considerations and by noticing that a.s. convergence implies convergence in distribution.
If and will go to infinity and thus be greater than eventually as we increase , otherwise if we will eventually not be able to reject . Analougous for .
Connecting the p-Value and Confidence Intervals
Let and and remember:
- For the approximative hypothesis test the p-value is
- The centered approximate confidence interval is
- if the p-value is larger than we cannot reject as seen in the p-value chapter.
Now the calculation is straighforward: