top of page

ISS 2016 Statistics Paper-2 Solution: Question 43 (Unbiased Estimator for Geometric Distribution)

1 day ago
6 min read

Continuing our question-by-question walk through the ISS Statistics Paper-2 archive, today we land on Question 43 of the 2016 paper. This one asks you to spot an unbiased estimator for a reciprocal parameter of the geometric distribution, which trips up a lot of students who have only ever seen the Poisson or exponential versions of this type of question.


Quick Summary


  • Topic: Unbiased estimation, geometric distribution

  • Question Reference: ISS 2016 Statistics Paper-2, Question 43

  • Correct Answer: Option (d), the sample mean x̄

The Question, As Asked

Question 43. An unbiased estimator of 1/&theta; for the distribution f(x, &theta;) = &theta;(1&minus;&theta;)x&minus;1, x = 1, 2, 3, ..., &infin;; 0 < &theta; < 1 is:


  • (a) 1/x

  • (b) 1/(n x&#772;)

  • (c) n x&#772;

  • (d) x&#772;

Step 1: Recognize the Distribution

Before touching any algebra, always name the distribution you are looking at. Here f(x, &theta;) = &theta;(1&minus;&theta;)x&minus;1 for x = 1, 2, 3, ... is the geometric distribution &mdash; specifically, the version that counts the trial number on which the first success occurs in a sequence of independent Bernoulli(&theta;) trials, with support starting at x = 1 rather than x = 0.


This matters because there is a second, equally common version of the geometric distribution with support x = 0, 1, 2, ... (counting the number of failures before the first success), and its mean formula is different ((1&minus;&theta;)/&theta; instead of 1/&theta;). Mixing the two up is the single most common way students lose marks on this exact type of question.

Step 2: Find E(X) Without Calculus

We need to know what 1/&theta; actually equals in terms of the random variable X, so that we can guess which statistic estimates it. The cleanest way to find E(X) for this &ldquo;waiting time to first success&rdquo; interpretation uses the memoryless property of Bernoulli trials &mdash; no differentiation of infinite series required.


Think of X as: &ldquo;the trial number of the first success.&rdquo; Condition on what happens on the very first trial:


  • With probability &theta;, the first trial is itself a success, so X = 1.

  • With probability (1&minus;&theta;), the first trial is a failure. At that point, the process &ldquo;restarts&rdquo; &mdash; the number of remaining trials needed is, by the memoryless property, distributed exactly like X all over again. So in this branch, the total count is 1 (for the failed trial already used) plus a fresh copy of X.


Writing this as an equation for the expected value:


E(X) = &theta;&middot;(1) + (1&minus;&theta;)&middot;(1 + E(X))


Expand the right-hand side carefully, term by term:


E(X) = &theta; + (1&minus;&theta;) + (1&minus;&theta;)E(X)


Since &theta; + (1&minus;&theta;) = 1, this simplifies to:


E(X) = 1 + (1&minus;&theta;)E(X)


Now collect every E(X) term on one side. Subtract (1&minus;&theta;)E(X) from both sides:


E(X) &minus; (1&minus;&theta;)E(X) = 1


E(X)&middot;[1 &minus; (1&minus;&theta;)] = 1


E(X)&middot;&theta; = 1


E(X) = 1/&theta;


Pro Tip: The trickiest part of this question is not the algebra &mdash; it is spotting that you don't need to sum an infinite series at all. The &ldquo;condition on the first trial, then use memorylessness&rdquo; argument turns a calculus problem into two lines of algebra. This is exactly the kind of shortcut that is usually taught by drilling the same restart-argument on several distributions (geometric, negative binomial) until it becomes automatic, rather than memorizing E(X) = 1/&theta; as an isolated formula.

Step 3: Double-Check with the Direct Summation

For completeness, here is the same result obtained the &ldquo;long way,&rdquo; which is useful if you ever need to find E(X&sup2;) or the variance and the restart trick doesn't directly apply.


Start from the geometric series formula, valid for any q with |q| < 1:


&sum;x=0&infin; qx = 1/(1&minus;q)


Differentiate both sides with respect to q. On the left, the derivative of qx is x&middot;qx&minus;1; on the right, the derivative of (1&minus;q)&minus;1 is (1&minus;q)&minus;2 (chain rule, with an extra minus sign from differentiating (1&minus;q) itself that cancels the minus sign already present):


&sum;x=1&infin; x&middot;qx&minus;1 = 1/(1&minus;q)2


Now set q = 1&minus;&theta; (so that 1&minus;q = &theta;) and multiply through by &theta;, since f(x,&theta;) = &theta;&middot;qx&minus;1:


E(X) = &sum;x=1&infin; x&middot;&theta;&middot;(1&minus;&theta;)x&minus;1 = &theta;&middot;&sum;x=1&infin; x&middot;qx&minus;1 = &theta;&middot;1/&theta;2 = 1/&theta;


Both methods agree: E(X) = 1/&theta;.

Step 4: From E(X) to an Unbiased Estimator

Now bring in the sample. We have X1, X2, ..., Xn, an i.i.d. random sample from this distribution, and the sample mean is x&#772; = (1/n)&sum;Xi.


An estimator T is called unbiased for a parameter g(&theta;) if E(T) = g(&theta;) for every valid &theta;. Take expectations of the sample mean:


E(x&#772;) = E[(1/n)&sum;i=1n Xi] = (1/n)&sum;i=1n E(Xi)


Since every Xi has the same distribution, E(Xi) = 1/&theta; for each i, so the sum inside has n identical terms:


E(x&#772;) = (1/n)&middot;n&middot;(1/&theta;) = 1/&theta;


This is exactly the target parameter, 1/&theta;, with no extra bias term left over. So x&#772; is an unbiased estimator of 1/&theta;.

Step 5: Rule Out the Other Options

It is worth seeing briefly why the distractors fail:


  • (a) 1/x uses a single observation and, worse, because of Jensen's inequality for the convex function 1/x, E(1/X) is generally not equal to 1/E(X) &mdash; it is biased.

  • (b) 1/(n x&#772;) scales down by an extra factor of n for no reason; its expectation does not reduce to 1/&theta;.

  • (c) n x&#772; has expectation n/&theta;, which is n times too large &mdash; it estimates n/&theta;, not 1/&theta;.


Final Answer: (d) x&#772;, and this matches the official answer key (Series A, Item 43 = D) exactly, so there is no discrepancy to flag here.

Why This Question Matters

This question is a good checkpoint for two separate skills at once: correctly identifying which &ldquo;flavor&rdquo; of a named distribution you're dealing with (support starting at 0 vs. 1 changes the mean formula), and knowing that for i.i.d. samples, the sample mean is always an unbiased estimator of the population mean, whatever that population mean happens to be in terms of &theta;. Both ideas resurface constantly across ISS Paper-2 &mdash; in questions on the Poisson mean, the exponential rate, and uniform-distribution parameters that we've already covered earlier in this series.

Frequently Asked Questions

Why does the geometric distribution here start at x = 1 instead of x = 0?

Because f(x, &theta;) = &theta;(1&minus;&theta;)x&minus;1 models the trial number on which the first success occurs, and the earliest possible trial is the first one, so x = 1 is the smallest value X can take. If instead you were counting failures before the first success, the support would start at 0 and the pmf and mean formula would both look slightly different.

Is the sample mean always an unbiased estimator of the population mean?

Yes. For any i.i.d. sample X1, ..., Xn with finite mean &mu;, E(x&#772;) = &mu; always holds, regardless of what distribution the Xi come from. The only thing that changes from question to question is what &mu; equals in terms of the parameter, which is why you still have to derive E(X) first.

What is the memoryless property, and why does it apply here?

The memoryless property says that, given a trial has already failed, the number of additional trials needed for the first success has the same distribution as starting completely fresh. It applies to the geometric distribution because each Bernoulli trial is independent and identically distributed, so a failed trial carries no information about how many more trials are needed.

Could I have used the moment generating function instead?

Yes, the MGF of this geometric distribution is M(t) = &theta;et/[1&minus;(1&minus;&theta;)et], and differentiating it once and setting t = 0 also gives E(X) = 1/&theta;. It is a valid alternative, but it involves more algebra than the restart argument used above for this particular question.

Why is 1/x not an unbiased estimator of 1/&theta;, even though X estimates 1/&theta; &ldquo;on average&rdquo;?

Taking the reciprocal of a random variable does not commute with taking an expectation except in special cases. Concretely, E(1/X) &ne; 1/E(X) in general whenever X is non-degenerate, which is a direct consequence of Jensen's inequality applied to the convex function g(x) = 1/x.

How many such unbiased-estimator questions typically appear in ISS Paper-2?

Quite a few &mdash; this ISS 2016 paper alone has several variants on unbiasedness, consistency, and sufficiency across different named distributions (Poisson, uniform, exponential, Bernoulli, and now geometric), so it is worth being comfortable deriving E(X) for every standard distribution rather than memorizing a table of formulas.


If anything in this derivation felt unclear, or you solved it using a different method, drop a comment below &mdash; and if you know a fellow ISS aspirant working through the same paper, consider sharing this post with them.

Recent Posts

See All

Comments


bottom of page