ISS 2016 Statistics Paper-2 Solution: Question 44 (Unbiasedness and Consistency for Bernoulli p)
This post continues our question-by-question walkthrough of the ISS Statistics Paper-2 previous year papers — today we pick up the 2016 paper at Question 44, which tests one of the most foundational ideas in estimation theory using the humble Bernoulli distribution.
Quick Summary
Topic:Unbiasedness and consistency of the sample mean X̄ as an estimator of p, for a random sample from b(1, p)
Question Reference:ISS 2016, Statistics Paper-2, Question 44
Correct Answer:(c) Sample mean X̄ is unbiased and consistent for p
The Question, As Asked
If X1, X2, X3, ..., Xn is a random sample from b(1, p), then :
Sample mean X̄ is unbiased for p but not consistent
Sample mean X̄ is consistent for p but not unbiased
Sample mean X̄ is unbiased and consistent for p
Sample mean X̄ is neither unbiased nor consistent for p
Setting Up the Problem — What Does b(1, p) Mean?
The notation b(1, p) refers to the Bernoulli distribution with parameter p. Each Xi takes the value 1 with probability p (call it "success") and 0 with probability (1 − p) ("failure"). This is the simplest possible discrete distribution, but it is the building block for the Binomial distribution, and it shows up constantly in estimation theory questions precisely because the algebra stays simple while the concepts stay important.
Before touching the options, we need two basic facts about a single Bernoulli observation Xi:
Mean: E(Xi) = p
Variance: Var(Xi) = p(1 − p)
If you haven't derived these before, here is the quick derivation so nothing is taken on faith. Since Xi only takes values 0 and 1:
E(Xi) = (1)(p) + (0)(1 − p) = p
E(Xi2) = (12)(p) + (02)(1 − p) = p
Var(Xi) = E(Xi2) − [E(Xi)]2 = p − p2 = p(1 − p)
These two results are the only raw material we need. Everything about the sample mean X̄ = (1/n) ∑ Xi follows from them.
Step 1 — Checking Whether X̄ Is Unbiased
An estimator T is called unbiased for a parameter θ if E(T) = θ exactly, for every sample size n. So we need to compute E(X̄) and see if it equals p.
By definition, X̄ = (1/n)(X1 + X2 + ... + Xn). Taking expectation on both sides, and using the fact that expectation is linear (the expectation of a sum is the sum of the expectations):
E(X̄) = (1/n) × [E(X1) + E(X2) + ... + E(Xn)]
Since every Xi is drawn from the same b(1, p) distribution, each E(Xi) = p. There are n such terms, so the sum inside the brackets is n × p:
E(X̄) = (1/n) × (n × p) = p
The n in the numerator and the n in the denominator cancel exactly, leaving E(X̄) = p. This holds for every value of n — n = 1, n = 10, n = 10,000, it does not matter. So by definition,X̄ is an unbiased estimator of p. This immediately rules out options (b) and (d), both of which claim X̄ is not unbiased.
Step 2 — Checking Whether X̄ Is Consistent
Consistency is a different property from unbiasedness, and it is about what happens as the sample size n grows very large. Formally, an estimator Tn is consistent for θ if, for every small number ε > 0,
P( |Tn − θ| ≥ ε ) → 0 as n → ∞
In plain words: as you collect more and more data, the estimator should get arbitrarily close to the true parameter with probability approaching 1. This is also called convergence in probability.
Proving this directly from the definition every time would be tedious, so in practice we use a simple and very useful sufficient condition, which follows from Chebyshev's inequality:
If E(Tn) → θ and Var(Tn) → 0 as n → ∞, then Tn is consistent for θ.
We have already shown E(X̄) = p for every n, so the first condition (E(Tn) → p) is trivially satisfied — it does not just approach p, it equals p exactly, always. Now we just need to check whether Var(X̄) goes to 0 as n grows.
Since X1, X2, ..., Xn are independent (that is what "random sample" means), the variance of their sum is the sum of their variances — there are no covariance cross-terms to worry about:
Var(X1 + X2 + ... + Xn) = Var(X1) + Var(X2) + ... + Var(Xn) = n × p(1 − p)
Now apply the rule Var(aY) = a2Var(Y) with a = 1/n and Y = the sum above:
Var(X̄) = (1/n2) × Var(X1 + X2 + ... + Xn) = (1/n2) × n × p(1 − p) = p(1 − p) / n
Look carefully at this result: p(1 − p) is a fixed number once p is fixed (it does not depend on n), while the denominator n grows without bound as the sample size increases. So as n → ∞,
Var(X̄) = p(1 − p) / n → 0
Both conditions of the sufficient criterion are now satisfied: E(X̄) = p for all n, and Var(X̄) → 0 as n → ∞. Therefore X̄ is a consistent estimator of p as well.
Pro Tip: The trap in this question is assuming unbiasedness and consistency are "the same kind of good property," so students sometimes pick an option that treats them as a package deal in the wrong direction. They are actually independent properties — you can have one without the other (a classic example is the sample variance divided by n for a normal population, which is consistent but slightly biased). The way this is usually taught is to always check the two properties through two completely separate calculations: first find E(Tn) and compare it to θ, then separately find Var(Tn) and see if it vanishes as n → ∞. Students who drill this two-step check until it becomes automatic rarely lose marks on these "which property holds" questions, whichever distribution the examiner dresses it up with.
Step 3 — Matching With the Given Options
Let us now go through all four options with our two results in hand (X̄ is unbiased for p, and X̄ is consistent for p):
(a) "Unbiased but not consistent" — wrong, because we showed Var(X̄) → 0, so it is consistent.
(b) "Consistent but not unbiased" — wrong, because E(X̄) = p exactly, so it is unbiased.
(c) "Unbiased and consistent for p" — matches both of our derivations exactly.
(d) "Neither unbiased nor consistent" — wrong on both counts.
So the correct answer is option (c). Checking against the official ISS 2016 Paper-2 answer key (Series A), the recorded answer for Question 44 is also C — our independent derivation agrees fully with the official key, with no discrepancy.
Why This Question Matters
This question is a textbook check on whether a candidate actually understands the definitions of unbiasedness and consistency, rather than just memorizing "the sample mean is always a good estimator." The Bernoulli/binomial proportion setting is also directly relevant to real applications — estimating a defect rate, a response rate, or a vote share is literally estimating a Bernoulli parameter p from sample proportions. Questions built on this exact template (same logic, different distribution) appear repeatedly across ISS, IIT JAM, and GATE Statistics papers, so mastering the two-step E( ) and Var( ) check here pays off well beyond this one question.
Frequently Asked Questions
What exactly does it mean for an estimator to be "unbiased"?
An estimator T is unbiased for a parameter θ if its expected value equals θ exactly, for every sample size, that is E(T) = θ. It means that if you repeated the sampling process infinitely many times and averaged the estimator's values, you would get exactly the true parameter with no systematic over- or under-estimation.
What does "consistency" mean for an estimator?
Consistency is a large-sample property: it says that as the sample size n grows to infinity, the estimator gets closer and closer to the true parameter value with probability approaching 1. Unlike unbiasedness, which must hold for every n, consistency is only a statement about the limiting behaviour as n becomes very large.
Does an unbiased estimator always have to be consistent?
No. An estimator can be unbiased for every sample size but still fail to be consistent if its variance does not shrink to zero as n increases. Unbiasedness and consistency are checked by two separate conditions — one on the mean, one on the variance — and neither one automatically implies the other.
Can a biased estimator still be consistent?
Yes, this happens often. A classic example is using n in the denominator instead of (n − 1) when estimating the variance of a normal population: this estimator is slightly biased for any finite n, but as n → ∞ the bias shrinks to zero and the variance also shrinks to zero, so the estimator is consistent despite being biased.
Why is Chebyshev's inequality the tool used to prove consistency here?
Chebyshev's inequality bounds the probability that an estimator strays far from its mean, in terms of its variance. If an estimator's mean converges to the true parameter and its variance converges to zero, Chebyshev's inequality guarantees that the probability of straying away from the true parameter also goes to zero — which is exactly the definition of convergence in probability, so it gives us a clean sufficient condition without having to work with limiting distributions directly.
Is the sample mean always a good estimator for a Bernoulli parameter p?
For the plain Bernoulli/binomial proportion problem, yes — the sample mean is unbiased, consistent, and also happens to be the maximum likelihood estimator and a sufficient statistic for p, which is why it is the standard choice in practice. This makes Bernoulli proportion estimation one of the cleanest and most frequently tested settings in estimation theory.
If any step above felt unclear, or you solved this differently, drop a comment below with your doubt — and if this helped, consider sharing it with a fellow ISS aspirant who is also working through the 2016 paper.

Comments