AP Statistics Inference for Categorical Data: Proportions — Worked Answer Explanations
Unit 6 · 12 questions explained
Below is a complete answer key for our AP Statistics Inference for Categorical Data: Proportions practice questions. For each question you'll find the correct choice, a full written explanation of how to get there, and — for every wrong answer — a short note on exactly why it's tempting and where it goes wrong. Reading these straight through is one of the fastest ways to find the gaps in a unit before exam day.
Prefer to test yourself first? Take the timed Inference for Categorical Data: Proportions practice test and come back here to review, or head back to the Inference for Categorical Data: Proportions unit overview.
- Question 1 · Easy
A researcher wants to estimate the proportion of adults who exercise regularly. She samples 400 adults and finds 160 exercise regularly. What is the 95% confidence interval for the true proportion?
- AWhy not A: Using instead of for 95% confidence.
- BCorrect
- CWhy not C: Rounding the standard error to 0.02 instead of computing .
- DWhy not D: Reporting only without a margin of error.
Explanation. . 95% CI: , giving .
Key takeawayOne-proportion z-interval: $\hat{p} \pm z^* \sqrt{\hat{p}(1-\hat{p})/n}$; use $z^* = 1.96$ for 95% confidence.
- A
- Question 2 · Easy
A poll finds that 54 of 200 surveyed voters support a ballot measure. A reporter states: "We are 95% confident that between 20.5% and 33.5% of all voters support the measure." How should this interval be correctly interpreted?
- AThere is a 95% probability that the true population proportion is between 20.5% and 33.5%.Why not A: A frequentist confidence interval does not assign probability to a fixed (but unknown) parameter.
- BWe are 95% confident that the interval (20.5%, 33.5%) captures the true proportion of all voters who support the measure.Correct
- C95% of voters support the measure at a level between 20.5% and 33.5%.Why not C: This misinterprets the interval as describing individual voter behavior.
- DThe interval was constructed by a method that is correct 95% of the time in any given poll.Why not D: The 95% confidence level means 95% of intervals constructed by this method will capture in the long run, not that this specific interval is correct 95% of the time.
ExplanationA confidence interval is interpreted as: we are 95% confident this specific interval captures the true population parameter. The "95%" refers to the long-run success rate of the method — if we built many such intervals, about 95% would contain . We do not assign probability to whether a specific interval contains (it either does or doesn't).
Key takeawayConfidence interval interpretation: 'We are C% confident the interval captures the true parameter.' Do not say 'there is a C% probability the parameter is in the interval.'
- A
- Question 3 · Easy
A company claims that 80% of its customers are satisfied. A consumer group surveys 100 customers and finds 72 are satisfied. Which hypothesis should be used to test whether the satisfaction rate is lower than claimed?
- A,Why not A: Using the sample proportion as the null hypothesis value instead of the claimed population value.
- B,Why not B: A two-sided alternative is used when testing whether the proportion differs in either direction; here we only care about lower.
- C, Correct
- D,Why not D: The null hypothesis always states equality; the alternative is the claim being tested.
ExplanationThe null hypothesis states what is claimed (). The consumer group is testing whether satisfaction is lower than claimed, so the alternative is one-sided: . Always state with equality.
Key takeaway$H_0$ states equality to the claimed value; $H_a$ reflects the direction of the research question (less than, greater than, or not equal).
- A
- Question 4 · Easy
For the test vs. using the data from the previous question (, ), the -test statistic is approximately and the -value is 0.023. At , what is the correct conclusion?
- AFail to reject ; there is insufficient evidence that fewer than 80% of customers are satisfied.Why not A: -value = 0.023 < = 0.05, so we reject .
- BReject ; there is convincing evidence that fewer than 80% of customers are satisfied.Correct
- CReject ; we have proved the true satisfaction rate is exactly 0.72.Why not C: Rejecting does not prove the specific value of the parameter; it only provides evidence against .
- DAccept ; the satisfaction rate is 80%.Why not D: We never 'accept' the null hypothesis; failing to reject is different from proving it true.
ExplanationSince -value = 0.023 < , we reject . There is statistically significant evidence that the true proportion satisfied is less than 80%. We conclude there is convincing evidence the claim is overstated.
Key takeawayReject $H_0$ when $p$-value < $\alpha$. State conclusion in context and direction; never 'accept' $H_0$.
- A
- Question 5 · Easy
A researcher increases the confidence level from 90% to 99% while keeping the sample size constant. What happens to the confidence interval?
- AThe interval becomes narrower because higher confidence means more precision.Why not A: Higher confidence requires a wider interval to capture the parameter with greater certainty.
- BThe interval becomes wider because a higher is used.Correct
- CThe interval stays the same because the sample data do not change.Why not C: The critical value changes with the confidence level, so the interval changes.
- DThe interval becomes wider and the point estimate changes.Why not D: The point estimate depends only on the sample data, not on the confidence level.
ExplanationA higher confidence level requires a larger (1.645 for 90%, 2.576 for 99%), which multiplies the standard error to give a larger margin of error. To be more confident of capturing the parameter, the interval must be wider.
Key takeawayHigher confidence level → larger $z^*$ → wider interval. There is a trade-off between confidence and precision.
- A
- Question 6 · Easy
A researcher wants to estimate a population proportion to within ±0.04 with 95% confidence, but has no prior estimate of . What minimum sample size is needed?
- A385Why not A: Using (90% confidence) instead of (95%).
- B601Correct
- C25Why not C: Computing without the and factors.
- D2401Why not D: Using instead of in the formula.
ExplanationWith no prior estimate, use (maximizes required ). . Round up to .
Key takeawaySample size formula: $n = z^{*2} p^*(1-p^*) / ME^2$. When $p$ is unknown, use $p^* = 0.5$ to get the most conservative (largest) $n$.
- A
- Question 7 · Easy
A random sample of 150 teenagers found 90 who own a smartphone. A random sample of 120 adults found 60 who own a smartphone. Researchers test vs. . What is the pooled proportion ?
- AWhy not A: Averaging the two sample proportions (0.60 + 0.50)/2 = 0.55 is not the pooled estimate when sample sizes differ.
- BCorrect
- CWhy not C: This is only the teen sample proportion, not the combined pooled proportion.
- DWhy not D: Simple average ignores that the sample sizes differ; pooled proportion weights by sample size.
ExplanationThe pooled proportion combines all successes and all observations: . This is used as the best estimate of the common under .
Key takeawayPooled proportion for two-sample test: $\hat{p}_c = (x_1 + x_2)/(n_1 + n_2)$; do not average the two sample proportions.
- A
- Question 8 · Easy
A 95% confidence interval for the proportion of adults who read daily is . A school administrator says, "95% of adults read daily between 31% and 47% of the time." What is wrong with this interpretation?
- AThe interval should be centered at the mean, not the proportion.Why not A: For proportions, the center is , which is the sample proportion. This is correct.
- BThe interval estimates the population proportion, not the frequency of individual behavior; 'between 31% and 47%' refers to confidence about a fixed parameter.Correct
- CThe interval is too wide; a 95% interval should be narrower.Why not C: Width depends on sample size and ; there is no inherent standard for what is 'too wide'.
- DThe confidence level should be stated as a range, not a single value.Why not D: Confidence levels are always stated as a single value (e.g., 95%), not a range.
ExplanationThe confidence interval estimates the true population proportion of adults who read daily. It does not mean individuals read 31%–47% of the time. The correct statement: we are 95% confident the true proportion of adults who read daily is between 31% and 47%.
Key takeawayA CI estimates a population proportion, not an individual's frequency of behavior; always restate the parameter in context.
- A
- Question 9 · Medium
The -value for a one-sided test is 0.03. Which of the following is a correct interpretation of this -value?
- AThere is a 3% probability that is true.Why not A: is either true or false (fixed); the -value is not the probability is true.
- BIf were true, there is a 3% probability of obtaining a test statistic at least as extreme as the one observed.Correct
- CThere is a 97% probability that the alternative hypothesis is correct.Why not C: The -value says nothing about the probability that is true.
- DThe data support with 97% confidence.Why not D: A small -value is evidence against , not for it.
ExplanationThe -value is the probability of observing a test statistic as extreme as or more extreme than the observed one, assuming is true. It measures the compatibility of the data with . A small -value means the observed data would be unlikely if were true.
Key takeaway$p$-value = $P(\text{data this extreme or more extreme} \mid H_0 \text{ true})$. It is not the probability $H_0$ is true.
- A
- Question 10 · Medium
Researchers conduct a two-sample -test for proportions comparing recovery rates between two treatments. They obtain -value = 0.08 with . A skeptic says the treatment difference is statistically significant. Is the skeptic correct?
- AYes; -value = 0.08 is close to , which is nearly significant.Why not A: 'Nearly significant' is not a formal statistical conclusion; the standard is whether -value < .
- BNo; since -value = 0.08 > , we fail to reject and cannot conclude a significant difference.Correct
- CYes; any -value below 0.10 indicates significance.Why not C: Significance is determined by comparison to the pre-set , not to 0.10 as a universal threshold.
- DNo; a two-sided test should be used, which would double the -value to 0.16.Why not D: We cannot retroactively change the test direction; the -value is compared to as stated.
ExplanationSince -value = 0.08 > , we fail to reject . The result is not statistically significant at the 5% level. We cannot conclude the proportions differ. 'Borderline' -values do not justify claiming significance.
Key takeawayFail to reject $H_0$ when $p$-value ≥ $\alpha$; 'nearly significant' is not a valid conclusion.
- A
- Question 11 · Hard
A 90% confidence interval for the proportion of defective items is . A quality engineer claims the defect rate is 15%. Is the claim consistent with the interval at the 10% significance level?
- AYes; since 15% is close to 12%, it is consistent.Why not A: 15% falls outside the interval; 'close to' is not a valid statistical criterion.
- BNo; since 15% falls outside the 90% CI, a two-sided test at would reject .Correct
- CCannot determine without computing the -value.Why not C: For a two-sided test, there is a direct duality: a value outside a CI corresponds to rejecting at significance level .
- DYes; a 90% CI at can only reject proportions below 0.04.Why not D: A two-sided interval rejects values outside both bounds (either tail).
ExplanationThere is a duality between confidence intervals and two-sided hypothesis tests: if a value falls outside a CI, the corresponding two-sided test at significance level would reject. falls outside , so a two-sided test at would reject .
Key takeawayDuality: values outside a $(1-\alpha)$ CI would be rejected by a two-sided hypothesis test at level $\alpha$.
- A
- Question 12 · Hard
A large school district conducts a two-sample -test and finds a statistically significant difference in graduation rates between two schools (-value = 0.002). The school with the lower rate has vs. . What is the most important caveat when reporting these results?
- AThe test should not have been performed because both rates are above 0.90.Why not A: There is no restriction on performing tests when rates are high.
- BStatistical significance does not imply practical significance; a 2-percentage-point difference may not be educationally meaningful.Correct
- CA -value of 0.002 proves the schools are different in graduation policy.Why not C: Statistical significance in rates does not reveal policy causes, and 'proves' is too strong.
- DWith a large enough sample, any small difference will become statistically significant, so no conclusion is possible.Why not D: Large samples do enable detecting small differences, but conclusions remain valid — they just need to distinguish statistical from practical significance.
ExplanationWith very large sample sizes, even tiny differences become statistically significant. Here, vs. is a 2-percentage-point gap. A small -value confirms this is not due to chance, but does not tell us whether a 2% difference in graduation rates is meaningful for policy or practice. Always distinguish statistical and practical significance.
Key takeawayStatistical significance ≠ practical importance. Large samples detect tiny differences; always evaluate effect size alongside the $p$-value.
- A