top of page

ISS 2016 Statistics Paper-1 Solution: Question 48 (Degrees of Freedom in a Chi-Square Test)

5 hours ago
5 min read

Continuing our question-by-question walk through the ISS Statistics Paper-1 archive, today's post picks up right where Question 47 left off. This one looks almost too short to be an exam question, but that's exactly why candidates lose marks on it — the trap is in rushing past a one-line formula without actually understanding where it comes from.


Quick Summary


  • Topic:Degrees of freedom in a Chi-square test of independence

  • Question Reference:ISS 2016, Statistics Paper-1, Question 48

  • Correct Answer:(d) 9

The Question, As Asked

In a Chi-square test, the contingency table has 4 rows and 4 columns. What is the number of degrees of freedom?


  • (a) 3

  • (b) 4

  • (c) 8

  • (d) 9

Step 1: Know What "Degrees of Freedom" Actually Means Here

Degrees of freedom (df) is just a count of how many cell values in the table are free to be chosen before every other cell is forced into a fixed value. In a Chi-square test of independence on a contingency table, we are not testing single numbers against a fixed target — we are comparing the observed cell frequencies against expected cell frequencies that we calculate ourselves from the row totals and column totals of the same table.


Because the expected frequencies are built directly from the row totals and column totals, those totals must match the observed row and column totals exactly. That matching requirement is what eats away at the number of cells that are genuinely free to vary.

Step 2: Set Up the Table and Count the Constraints

Say the table has r rows and c columns, so there are r × c cells in total. Now ask: once we fix all the row totals and all the column totals, how many individual cell entries can we still choose freely, with the rest being forced by subtraction?


Fill in the first (r − 1) rows and first (c − 1) columns of the table with any numbers you like. That's (r − 1)(c − 1) cells chosen freely. Once those are fixed:


  • The last cell in each of the first (r − 1) rows is forced, because that row's total is already fixed.

  • The last cell in each of the first (c − 1) columns is forced, because that column's total is already fixed.

  • The very last cell (bottom-right corner) is forced by both the last row total and the last column total simultaneously — and it will agree with both automatically, because the grand total ties everything together.


So out of the full r × c grid, exactly (r − 1)(c − 1) cells are genuinely free. Everything else is determined once the margins (row and column totals) are fixed. This gives the standard result:


df = (r − 1)(c − 1)

Step 3: Plug In the Numbers From This Question

Here we are told r = 4 rows and c = 4 columns. Substituting directly:


df = (r − 1)(c − 1) = (4 − 1)(4 − 1) = 3 × 3 =9


Pro Tip: The single most common slip on this type of question is using just (r − 1) or just (c − 1), or worse, using r × c as the degrees of freedom — treating every cell as independent. Remember that

Final Answer and Verification

Our independently derived answer is (d) 9, matching option (d) exactly. Cross-checking against the official ISS 2016 Statistics Paper-1 answer key (Series A) confirms that the key also lists the answer to Question 48 as (d). Our derivation and the official key are in full agreement — no discrepancy here.

Why This Question Matters

Degrees of freedom for contingency tables shows up constantly — not just as a standalone question like this one, but buried inside larger Chi-square goodness-of-fit and test-of-independence problems where you first have to compute the test statistic and only then look up the critical value using the correct df. If the df is wrong, every subsequent step and conclusion collapses, even if your arithmetic for the test statistic itself was flawless. Examiners know this, which is why a "just plug into a formula" question like this one quietly checks whether you understand the underlying logic of constrained cells rather than having simply memorized a one-line rule.

Frequently Asked Questions

Why isn't the degrees of freedom simply r × c, the total number of cells?

Because not all cells are free to vary once the row totals and column totals are fixed. The expected frequencies in a Chi-square test of independence are computed from these margins, so the margins act as constraints on the table, reducing the count of independently choosable cells from r × c down to (r − 1)(c − 1).

Does this formula change for a goodness-of-fit test instead of a test of independence?

Yes. For a simple goodness-of-fit test with k categories and no parameters estimated from the data, df = k − 1, since only the overall total is fixed. If you estimate m parameters from the sample to build the expected frequencies, it becomes df = k − 1 − m. The (r − 1)(c − 1) formula specifically applies to two-way contingency tables used for independence or homogeneity testing.

What would the degrees of freedom be for a 2×3 contingency table?

Using df = (r − 1)(c − 1) with r = 2 and c = 3, you get (2 − 1)(3 − 1) = 1 × 2 = 2. This kind of quick substitution is worth practicing with several different row/column combinations until it becomes automatic.

Is the degrees of freedom formula the same whether the table is square or rectangular?

Yes, df = (r − 1)(c − 1) holds regardless of whether r equals c. It doesn't matter if the table is square (like the 4×4 case in this question) or rectangular — the same margin-constraint logic applies either way.

How does degrees of freedom affect the final conclusion of the Chi-square test?

The degrees of freedom determines which row of the Chi-square distribution table you use to find the critical value at your chosen significance level. A wrong df value means you're comparing your test statistic against the wrong critical value, which can flip your conclusion about whether to reject the null hypothesis of independence.

Is there an easy way to remember the formula during the exam?

Think of it as "subtract one from each dimension, then multiply": take the number of rows, subtract 1; take the number of columns, subtract 1; multiply the two results. Practicing this on a few quick examples by hand, rather than just reading the formula, is usually what makes it stick under exam pressure.


If anything in this derivation felt unclear, or if you spotted a different way to arrive at the same answer, drop a comment below — working through the "why" together is often more useful than the answer alone. And if you know a fellow ISS aspirant who's grinding through this same paper, feel free to share this post with them.

Recent Posts

See All

Comments


bottom of page