Out of proportion
TLDR
XXXX
The age of inference
We are living, it seems, in the age of inference. The word has really entered the agenda. So I wonder if in this moment of peak AI-hype we are ready to have better discussions than we were before. Are the new tools helping us? Or do old habits die hard?
It's fine
Some years ago, I was managing a project where a CEO decided to short-circuit to usual development process and ask a directly a developer to implement an absolutely terrible UI behavior. When I asked the developer about the eye-offending monstrosity he committed, he quickly blamed the CEO. When I asked the CEO about the animated horror he requested he said it was fine, claiming to have checked with five members of the team. Five of his employees and no written evidence. And no ticket. Obviously.
Not every CEO knows statistics. Turns out that there is a statistical method to determine whether something is fine or not. It's called a survey. We might be interested in determining the proportion of the relevant population that considers that something is fine. Because studying a population is rarely a realistic possibility, we must accept to study only a sample, and infer the real proportion from the proportion found in the survey, hopefully with reasonable accuracy.
That's it.
A statistic is a mathematical function that maps a certain quantity X from n surveyed items of an N-sized population to a number: the Mean, the Median, the Standard Deviation, etc. Statistical inference is the process of estimating the statistics of X for the case n=N (population) when in fact n << N (much smaller sample). For this case, if we consider X=1 when "it's fine", and X=0 otherwise, the proportion we are looking for is the sample mean of X, where n is the number of customers surveyed and N is the universe of potential customers. Developers in the payroll are not to be surveyed by the CEO in what comes to the software they are developing!
The rest is details.
The rest
In order to find the sample mean of X, we need to find the possible values k for the number of "it's fine" answers and assign a probability to each case. The variable representing these possible values will be called Y. Let's consider the number of people that say "it's fine" in the population is an unknown value K. As the survey goes on, the probability of finding "it's fine" answers decays slowly every time a fine answer is found and increases slightly every time a less-than-fine answer is found. This survey process is also known as sampling without replacement.
The probability that the j-th extraction results in "it's fine" (a success) is:
\(\displaystyle \frac{K - \text{successes so far}}{N - j} \)
whereas the probability of insucess is:
\(\displaystyle \frac{N - j - (K - \text{successes so far})}{N - j} \)
For a sequence of k successes we will have the following terms in the numerator:
\(\displaystyle K(K-1)\dots(K-k+1) \cdot (N - K)(N-K-1)\dots(N-K-(n-k)+1) \)
and the following terms in the denominator:
\(\displaystyle N(N-1)\dots(N-n+1) \)
The probability for a generic sequence of survey results can be expressed as:
\(\displaystyle \frac{K(K-1)\dots(K-k+1) \cdot (N - K)(N-K-1)\dots(N-K-(n-k)+1)}{N(N-1)\dots(N-n+1)} \)
The numerators and denominators can be rewritten as combinations, at the cost of a left over term 1/C(n,k):
\(\displaystyle \frac{1}{\binom{n}{k}} \cdot \frac{\binom{K}{k}\binom{N-K}{n-k}}{\binom{N}{n}} \)
But because the sequence is generic for k successes and there are many possible sequences with the same number of successes, we need to multiply it precisely by coefficient C(n,k), which cancels the left over term. We arrived at the 300 year old Hypergeometric distribution:
\(\displaystyle P(Y=k) = \frac{\binom{K}{k}\binom{N-K}{n-k}}{\binom{N}{n}} \)
which in this case measures the probability of finding k "it's fine" answers in a population where K people think "it's fine". The estimated proportion is k/n.
Problem solved, right?
A much wanted simplification
The problem is still not solved because we don't know how good the estimation k/n is in relation to the real proportion p=K/N. And ideally, we would like a not too complicated solution.
Whenever the population is very large (e.g. in the magnitude of millions) compared the sample we able to collect (e.g., in the magnitude of hundreds) the expression for a generic sequence of survey results can be greatly simplified as:
\( \displaystyle \frac{K(K-1)\dots(K-k+1) \cdot (N - K)(N-K-1)\dots(N-K-(n-k)+1)}{N(N-1)\dots(N-n+1)} \simeq \left(\frac{K}{N}\right)^k \left(1 - \frac{K}{N}\right)^{n-k} \)
or, equivalently:
\( \displaystyle \frac{K(K-1)\dots(K-k+1) \cdot (N - K)(N-K-1)\dots(N-K-(n-k)+1)}{N(N-1)\dots(N-n+1)} \simeq p^k (1-p)^{n-k} \)
where p=K/N.
The fact that removing a small fraction of the population, little by little, as the survey advances, hardly influences the associated probabilities is something that connects with intuition, but can also be proven mathematically. We work under the assumption that even in the order of magnitude of respectable samples (hundreds of units) the approximation described above is valid.
As before, the the sequence is generic for k successes and there are many possible sequences with the same number of successes. Once we multiply by the combinations coefficient C(n,k) we obtain the much simpler Binomial Distribution:
\( \displaystyle P(Y=k) = {\binom{n}{k}} p^k (1-p)^{n-k} \)
It is simple to verify that the expected value is given by:
\( \displaystyle E(Y) = np \)
and the Variance is given by
\( \displaystyle Var(Y) = np(1-p) \)
Samples aren't simple
We derived above a variable called Y, that counts the number of successes for a sequence of observations of the "It's fine" variable, called X. Y follows approximately the Binomial Distribution. Because we want to count the proportion of people that say "It's fine" rather than the absolute number successes, we will need to define new variable W = Y/n. This variable follows what is usually called the Sampling Distribution. We can now use a property of the Variance to note that:
\( \displaystyle Var(W) = \frac{1}{n^2} Var(Y) = \frac{1}{n} p(1-p) \)
The corresponding Standard Deviation is then:
\( \displaystyle S(W) = \sqrt{\frac{p(1-p)}{n}} \)
This means the determination the of the population proportion from the sample proportion shows dispersion, which we measure via the usual metric: the Standard Deviation. We expect dispersion to go away as we add more and more observations, the sample proportion slowly converging to the population proportion. A interesting exercise is writing the expression for W explicitly. It can be easily obtained by substituting w=k/n in the definition of Y, from which we find:
\( \displaystyle P(W = w) = \binom{n}{n w} p^{n w} (1-p)^{n(1-w)} \)
defined for
\( \displaystyle w = \left\{ 0, \, \frac{1}{n}, \, \frac{2}{n}, \, \dots, \, \frac{n-1}{n}, \, 1 \right\} \)
In formal terms, this expression does not introduce any new information - it expresses the same information as the Binomial Distribution followed by Y. But seeing its form and its domain explicitly written highlights one obvious fact that might not be obvious: the amount of possible different values for the estimated proportion is exactly n. That is, the proportion space is split into intervals of width 1/n which are very coarse grained for low n. For the n=5 anecdote, the possible values for the estimated proportion could be no others than {0, 0.2, 0.4, 0.6, 0.8, 1 }. If the real proportion is, say 0.3, the nearest observable sample values (0.2 and 0.4) are 0.1 units off.
That is, on top of the natural dispersion that causes small samples to not represent the population well, there is also a problem of granularity of representation. This problem too, goes away with larger samples, because as n increases the domain of W becomes denser. With n=5 possible values are separated by spaces of 0.2 (or 20%, in percentage) whereas for n=20 possible values are separated by 0.05 (or %5 in percentage) and for n=100, possible values are separated by 0.01 (or 1%, in percentage). The maximum error introduced due lack of representation granularity would be 0.1, 0.025 and 0.0005 or, in percentages, 10%, 2.5%, and 0.5%, respectively.
In general, the granularity error is given by:
\( \displaystyle E_g >= \frac{1}{2n} \)
This idea alone already sets a minimum requirement for the sample size:
\( \displaystyle n_g >= \frac{1}{2E} \)
where E is the tolerable error. We will see later that this criterion is necessary but not sufficient - the reason is that the granularity problem goes away with the inverse of n, faster than the dispersion caused by sampling randomness, which fades out with the inverse of the square root of n. Still, this simple criterion is enough for ruling out small samples (e.g. n<10) where the granularity error is significant.
Finally, we also see from the above expressions for Y and W, that dispersion depends on the population proportion itself. If the population proportion is 0 or 1, dispersion will be 0, because all population items are equal. On the other hand, it is easy to see that dispersion is maximized when p=1/2.
Let's now let everything sink in for a moment: the determination is not straightforward, because dispersion decays slowly, with the the inverse of the square root of n and we expect it to be harder when the population's opinions are split 50/50. In order to prepare a survey we must be ready for the worst case: p=1/2.
But how? Ideally, we would find a way to locate the population proportion within acceptable bound around the estimated proportion.
Error and Confidence
Let's use variable W to formalize the desired result:
\( \displaystyle P\left(\frac{k}{n} - E \le p \le \frac{k}{n} + E\right) \ge 1 - \alpha \)
This condition says that with probability 1-alpha the population proportion is inside the estimated proportion +-E , where E is the error we decided to tolerate.
We can break it down in two separate conditions:
\( \displaystyle P\left(\frac{k}{n} \le p + E\right) \ge 1 - \frac{\alpha}{2} \)
\( \displaystyle P\left(\frac{k}{n} \ge p - E\right) \ge 1 - \frac{\alpha}{2} \)
but because the Binomial Distribution is symmetric they have the same information. In order to consider the worst case dispersion we set p=1/2:
\( \displaystyle P\left(\frac{k}{n} \le \frac{1}{2} + E\right) \ge 1 - \frac{\alpha}{2} \)
These is the generic expression that we need. Once we decide on the tolerated error and the probability we can theoretically determine the minimum necessary sample size. The required sample size will be larger for larger values of the required probability and for smaller values of the tolerated error. For E = 0.05 and alpha = 0.05 we would find:
\( \displaystyle P\left(\frac{k}{n} \le 0.55 \right) \ge 0.975 \)
There is, however, a minor practical problem: the equation above can't be inverted for n. Luckily, looking for the right value of n takes less than half a second on a computer. Here is an example program that calculates the minimum sample size:
#!/usr/bin/python3
from scipy.stats import binom
MAX_SAMPLE=100000
def verify_sample_size(n, E, alpha):
# we calculate for the proportion that maximizes dispersion, i.e., the worst case
p = 0.5
# k_upper is the largest integer outcome within the allowable margin E
k_upper = int(n * (p + E))
# Evaluate exact Binomial CDF at p = 0.5
coverage_prob = binom.cdf(k_upper, n, p)
# Must satisfy two-sided coverage (1 - alpha/2)
return coverage_prob >= (1 - alpha / 2)
for n in range(1, MAX_SAMPLE):
result = verify_sample_size(n, 0.05, 0.05)
if result:
print("the minimum necessary sample size is", n)
exit()
print("no value found for sample below", MAX_SAMPLE)