ISS 2016 Statistics Paper-1 Solution: Question 41 (Chi-Square Test for Independence Explained)
Continuing our question-by-question walk through the ISS 2016 Statistics Paper-1, today we tackle Question 41 — a question that doesn't ask you to crunch a single number at first glance, but actually tests whether you understand why you'd reach for one statistical test over another. Get the labels right here and three other options eliminate themselves almost automatically.
Quick Summary
Topic: Choosing the correct hypothesis test for data arranged in a contingency table (chi-square test of independence vs. goodness-of-fit vs. sign test vs. ANOVA)
Question Reference: ISS 2016, Statistics Paper-1, Question 41
Correct Answer: (c) Test for independence
The Question as It Appeared in the Paper
A statistician mailed a survey questionnaire to a random sample of 100 teachers from each of 4 types of schools. The number of responses received is summarized in the following table:
Government Primary — Number of responses: 23, Number of non-responses: 77
Government Secondary — Number of responses: 30, Number of non-responses: 70
Private Primary — Number of responses: 43, Number of non-responses: 57
Private Secondary — Number of responses: 29, Number of non-responses: 71
Before analysing the data collected, the statistician wants to test whether the probability of non-response was the same for all types of schools. What is the appropriate test for this purpose?
Sign test
Goodness-of-fit test
Test for independence
ANOVA F-test
Step 1: Figure Out What Shape the Data Is In
Before you can pick a test, you need to see what kind of table you're actually sitting on. Here we have 4 independent groups (the four types of schools), and for each group we're counting outcomes that fall into 2 categories (responded, did not respond). If you lay this out, you get a 4-row by 2-column table of counts — exactly the structure of a contingency table. Whenever you see counts cross-classified by two categorical variables like this, chi-square-based testing should be the first thing that comes to mind, not a t-test, not ANOVA, and not a sign test.
Step 2: Rule Out the Options That Don't Fit the Data Type
Let's eliminate the wrong answers first, because understanding why they fail is more useful than just memorizing the right one.
(a) Sign test is a test for the median of a single sample of paired or single observations, based on the signs of differences. There's no pairing of observations and no median being tested here, so it doesn't apply.
(b) Goodness-of-fit test is used when you have one sample and you want to check whether its distribution across categories matches some pre-specified theoretical distribution (for example, testing whether a die is fair). Here we don't have one sample being compared to a theoretical model — we have 4 separate samples (one per school type) being compared to each other. That rules out (b).
(d) ANOVA F-test compares means of a continuous variable across groups. Our data here is categorical (responded / did not respond), not a continuous measurement, so ANOVA has no role to play.
Step 3: Why "Test for Independence" Is the Right Label
That leaves (c). Here's the subtlety that trips up a lot of students: strictly speaking, what's being asked — "is the probability of non-response the same across the 4 school types?" — is technically called a test of homogeneity of proportions, not a test of independence. These are conceptually different questions: a test of independence asks whether two attributes of the same population are related, while a test of homogeneity asks whether several independent populations share the same distribution across categories. But here's the part that matters for exam purposes: both tests are carried out using the exact same chi-square statistic, the exact same expected-frequency formula, and the exact same degrees-of-freedom rule, applied to the identical 4×2 table of counts. Since "test of homogeneity" isn't offered as a choice, "test for independence" is the mechanically correct match among the four given options, because the computation you'd run is identical either way.
Pro Tip: When an MCQ gives you only "goodness-of-fit" and "test for independence" as your chi-square options but the actual setup is a homogeneity question, don't get stuck arguing semantics under exam pressure — check which one shares the same computational engine as what the question describes. This is exactly the kind of fine distinction that's usually drilled into students through repeated contingency-table practice until spotting the table shape becomes instant, which is how it's taught in structured test-series environments.
Step 4: Actually Running the Test (So You Can See Why the Label Fits)
Let's go all the way through the calculation, since seeing the mechanics confirms why this is a chi-square procedure.
First, find the row totals, column totals, and grand total. Each school type was surveyed with a sample of 100, so every row total is 100, and with 4 rows the grand total is 400.
Column totals: total responses = 23 + 30 + 43 + 29 = 125. Total non-responses = 77 + 70 + 57 + 71 = 275. Check: 125 + 275 = 400, which matches the grand total — always verify this before moving on.
Next, compute the expected count for each cell under the null hypothesis that response rate is the same everywhere, using the standard contingency-table formula:
Expected count = (row total × column total) ÷ grand total
Since every row total is 100, the expected count in the "responses" column is the same for all four rows: (100 × 125) ÷ 400 = 31.25. Similarly, the expected count in the "non-responses" column is (100 × 275) ÷ 400 = 68.75.
Now apply the chi-square formula, summing (Observed − Expected)2 ÷ Expected over all 8 cells:
Government Primary responses: (23 − 31.25)2 ÷ 31.25 = 2.178
Government Primary non-responses: (77 − 68.75)2 ÷ 68.75 = 0.990
Government Secondary responses: (30 − 31.25)2 ÷ 31.25 = 0.050
Government Secondary non-responses: (70 − 68.75)2 ÷ 68.75 = 0.023
Private Primary responses: (43 − 31.25)2 ÷ 31.25 = 4.418
Private Primary non-responses: (57 − 68.75)2 ÷ 68.75 = 2.008
Private Secondary responses: (29 − 31.25)2 ÷ 31.25 = 0.162
Private Secondary non-responses: (71 − 68.75)2 ÷ 68.75 = 0.074
Adding all 8 contributions: 2.178 + 0.990 + 0.050 + 0.023 + 4.418 + 2.008 + 0.162 + 0.074 ≈ 9.90.
Degrees of freedom for an r × c contingency table is (r − 1)(c − 1). With r = 4 rows and c = 2 columns, df = (4 − 1)(2 − 1) = 3. The tabulated chi-square critical value at 3 degrees of freedom and 5% significance is 7.815. Since our calculated value 9.90 exceeds 7.815, we'd actually reject the null hypothesis — interesting to note, since it means this particular dataset shows that non-response probability was not the same across all four school types (Private Primary schools responded noticeably more than expected).
Final Answer
The correct response is (c) Test for independence, which matches the official answer key for Series A of the ISS 2016 Statistics Paper-1. The test is carried out exactly as shown above: build the contingency table, compute expected frequencies from row and column totals, sum the squared standardized deviations, and compare against the chi-square critical value at (r−1)(c−1) degrees of freedom.
Why This Question Matters
This question isn't really testing your ability to multiply and divide — it's testing whether you can correctly identify what kind of hypothesis-testing problem you're looking at before you touch a formula. In the ISS exam, and in applied survey statistics generally, recognizing "this is a contingency-table situation" versus "this is a goodness-of-fit situation" versus "this needs a test of means" is often harder than the arithmetic that follows. Mastering that classification step is what actually separates candidates who can apply statistics to real data from those who can only solve a textbook problem once it's already been labeled for them.
What is the difference between a chi-square test of independence and a test of homogeneity?
A test of independence uses a single sample drawn once from a population and checks whether two categorical attributes measured on that sample are related. A test of homogeneity draws separate, independent samples from several populations (here, the four school types) and checks whether those populations share the same distribution across categories. Despite this conceptual difference, both are computed with the identical chi-square formula on the same contingency table shape.
Why isn't this a goodness-of-fit test?
A goodness-of-fit test compares one sample's observed category frequencies against a single pre-specified theoretical distribution, such as testing whether dice outcomes are uniform. Here we are comparing four separate samples (one per school type) against one another rather than against a fixed theoretical proportion, which is the hallmark of a contingency-table test rather than a goodness-of-fit test.
How do you calculate degrees of freedom for this kind of table?
For any r-row by c-column contingency table, degrees of freedom equal (r − 1) multiplied by (c − 1). In this question, there are 4 rows (school types) and 2 columns (response, non-response), giving (4 − 1)(2 − 1) = 3 degrees of freedom, which is the value you'd look up in the chi-square table.
What does it mean practically that the chi-square value came out significant here?
A calculated chi-square of 9.90 against a critical value of 7.815 at 3 degrees of freedom means the differences in non-response rates across the four school types are too large to attribute to random sampling variation alone. In practical terms, the statistician would conclude that non-response behavior genuinely differs by school type, which matters for how the survey results get interpreted and weighted afterward.
Can the sign test or ANOVA F-test be used for this kind of data?
No. The sign test is designed for testing a median using the signs of paired differences, and ANOVA F-test compares means of a continuous variable across groups. Response/non-response here is categorical count data, not continuous measurements or paired differences, so neither test's assumptions or mechanics fit this situation.
What assumptions must hold for the chi-square test to be valid?
The observations within each sample must be independent of each other, the categories must be mutually exclusive and exhaustive, and the expected frequency in each cell should generally be 5 or more for the chi-square approximation to the true sampling distribution to hold reasonably well. In this question all expected counts (31.25 and 68.75) comfortably clear that threshold.
If any step above felt rushed or you'd solve this differently, drop a comment below — working through the reasoning out loud is often what makes it stick. And if you know a fellow ISS aspirant still mixing up goodness-of-fit with test of independence, send this one their way.

Comments