AP Statistics Exploring Two-Variable Data — Worked Answer Explanations
Unit 2 · 12 questions explained
Below is a complete answer key for our AP Statistics Exploring Two-Variable Data practice questions. For each question you'll find the correct choice, a full written explanation of how to get there, and — for every wrong answer — a short note on exactly why it's tempting and where it goes wrong. Reading these straight through is one of the fastest ways to find the gaps in a unit before exam day.
Prefer to test yourself first? Take the timed Exploring Two-Variable Data practice test and come back here to review, or head back to the Exploring Two-Variable Data unit overview.
- Question 1 · Easy
A scatterplot of hours studied () versus exam score () for 20 students shows a positive linear association with no obvious outliers. The correlation coefficient is . What does this value tell us?
- A83% of the variation in exam scores is explained by hours studied.Why not A: This describes , not .
- BThere is a strong positive linear association between hours studied and exam score.Correct
- CStudying more hours causes higher exam scores.Why not C: Correlation does not imply causation; association is not the same as a causal relationship.
- DEach additional hour of studying increases the score by 0.83 points.Why not D: This describes the slope, not the correlation coefficient.
Explanationdescribes the strength and direction of a linear association. Values of near 1 indicate a strong linear relationship; indicates a positive direction. does not measure causation or slope.
Key takeaway$r$ measures the strength and direction of a linear association; it does not imply causation or represent the slope.
- A
- Question 2 · Easy
Computer output for a regression of weight (, in lbs) on height (, in inches) gives: , . Interpret the slope.
- AFor each additional pound, height is predicted to increase by 4.5 inches.Why not A: This reverses the roles of and .
- BFor each additional inch of height, weight is predicted to increase by 4.5 lbs.Correct
- CThe average weight is 4.5 lbs.Why not C: 4.5 is the slope (rate of change), not an average value of .
- D71% of the increase in weight is caused by height.Why not D: This misinterprets and incorrectly claims causation.
ExplanationThe slope is interpreted as: for each 1-unit increase in (height), the predicted value of (weight) increases by 4.5 units (lbs). The slope is a rate of change of the predicted response per unit of the explanatory variable.
Key takeawaySlope interpretation: for each 1-unit increase in $x$, the predicted $y$ changes by $b$ units.
- A
- Question 3 · Easy
A least-squares regression line is fit to data on advertising spending (, thousands of dollars) and sales (, thousands of units). The residual for one observation is . What does this mean?
- AThe actual sales were 12 thousand units less than predicted.Why not A: A positive residual means actual exceeds predicted, not the reverse.
- BThe actual sales were 12 thousand units more than predicted.Correct
- CAdvertising spending was overestimated by 12 thousand dollars.Why not C: Residuals refer to the response variable , not the explanatory variable .
- DThe regression line passes 12 units above the data point.Why not D: A positive residual means the data point is above the line, not below.
ExplanationResidual . A residual of means the actual sales exceeded the predicted sales by 12 thousand units. The point lies above the regression line.
Key takeawayResidual $= y - \hat{y}$. Positive residual: point above the line; negative residual: point below the line.
- A
- Question 4 · Easy
A residual plot for a linear regression model shows a clear curved (U-shaped) pattern. What does this indicate?
- AThe regression line fits the data well.Why not A: A well-fitting linear model produces a residual plot with no pattern.
- BThere is likely an outlier in the data.Why not B: A U-shaped pattern across the whole plot indicates a systematic non-linear trend, not just an outlier.
- CA linear model is not appropriate; the relationship is non-linear.Correct
- DThe standard deviation of the residuals is too large.Why not D: The size of residuals relates to variability, not the pattern shape, which indicates non-linearity.
ExplanationA residual plot for a well-fitting linear model should show random scatter with no pattern. A systematic curve in the residual plot reveals that the linear model does not capture the true relationship — a non-linear model would be more appropriate.
Key takeawayPatterns in a residual plot (curves, fans) signal that the linear model is inadequate.
- A
- Question 5 · Easy
For a regression of final exam score on midterm score, the regression equation is , . A student scored 70 on the midterm. What is the student's predicted final exam score?
- A63Why not A: Using as the coefficient instead of the slope: .
- B66Correct
- C80Why not C: Using the slope as the predicted value without computing .
- D70Why not D: Using the midterm score directly without applying the regression equation.
ExplanationSubstitute into the regression equation: . The correct answer is 66.
Key takeawayTo predict, substitute the given $x$-value into $\hat{y} = a + bx$ and compute.
- A
- Question 6 · Easy
A regression of salary (, thousands of dollars) on years of experience () yields . Which interpretation is correct?
- A, indicating a moderately positive linear association.Why not A: The value given is , not ; .
- B64% of the variation in salary is explained by the linear relationship with years of experience.Correct
- C64% of employees have salaries predicted correctly by the model.Why not C: measures explained variation, not prediction accuracy rate for individuals.
- DYears of experience causes 64% of the increase in salary.Why not D: describes association strength, not causal attribution.
Explanationmeans 64% of the variability in salary is accounted for by the least-squares regression on years of experience. Note that , not 0.64.
Key takeaway$r^2$ = proportion of variability in $y$ explained by the linear model; it is always between 0 and 1.
- A
- Question 7 · Easy
A researcher removes one influential point from a scatterplot. After removing it, increases from 0.60 to 0.91. What type of point was removed?
- AA point with a large positive residualWhy not A: Large residuals indicate poor fit but don't necessarily reduce correlation if the point follows the linear trend direction.
- BAn outlier in the -direction only (high-leverage point)Why not B: High leverage is about extreme -values, not -values, and pulling down suggests the point did not fit the linear pattern.
- CAn influential outlier that did not follow the linear pattern of the other pointsCorrect
- DA data point with -value equal to the mean ofWhy not D: Points at have zero leverage and minimal influence on the regression line.
ExplanationWhen removing a point dramatically increases (0.60 → 0.91), that point weakened the linear association. It was an influential outlier — a point that did not fit the overall linear pattern and pulled downward. Influential points are often far from the bulk of the data in the -direction.
Key takeawayAn influential point can dramatically change the regression line and $r$ when removed; identify them with leverage and residual analysis.
- A
- Question 8 · Easy
The regression equation for predicting a city's high temperature in July (, °F) from its latitude (, degrees north) is . Interpret the -intercept in context.
- AAt 0° latitude (equator), the predicted July high is 115°F.Correct
- BFor each degree increase in latitude, temperature decreases by 115°F.Why not B: This describes the slope, not the -intercept.
- CThe average July high temperature across all cities is 115°F.Why not C: The -intercept is the predicted when , not the average .
- DThe -intercept has no meaningful interpretation because latitude cannot equal zero.Why not D: Latitude can equal 0 (the equator), so the intercept has a technically valid interpretation even if extrapolation beyond the data is risky.
ExplanationThe -intercept is the predicted value of when . In context, when (equator, 0° latitude), the model predicts a July high of 115°F. Whether this is a practically meaningful prediction depends on whether 0° is within the range of the data.
Key takeaway$y$-intercept is the predicted $y$ when $x = 0$; always interpret it in the context of the data.
- A
- Question 9 · Medium
A scatterplot of two variables shows a strong curved relationship. A student computes and concludes the association is weak. What is the error in the student's reasoning?
- AThe student should use instead of to assess strength.Why not A: also measures strength of linear fit; it would similarly underestimate a curved relationship.
- Bmeasures only the strength of the linear association; a low with a curved pattern means the relationship is strong but non-linear.Correct
- CThe student should use Spearman's rank correlation for curved data.Why not C: While rank correlation is more robust, the AP Stats curriculum focuses on and its limitations with non-linear data.
- DA value of can still indicate a strong relationship.Why not D: Without additional context, 0.45 is not considered a strong linear association; the issue is that is the wrong tool here.
ExplanationThe correlation coefficient measures only the strength of a linear association. If the scatterplot shows a curved relationship, may be low even though the association is actually very strong. The student should use the scatterplot itself, not alone, to characterize the association.
Key takeaway$r$ only measures linear association; always look at the scatterplot — a low $r$ with a clear curve means a strong non-linear relationship.
- A
- Question 10 · Medium
The regression of on produces a better fit than the regression of on . What does this suggest about the relationship between and ?
- AThe relationship between and is exponential.Correct
- BThe relationship between and is linear.Why not B: If the relationship were linear, the untransformed regression would fit as well as the transformed one.
- CThe relationship between and is a power function.Why not C: A power function would be linearized by taking of both and , not just .
- Dmust be transformed, not .Why not D: Taking (not ) linearizes exponential growth in ; the choice of transformation depends on the pattern.
ExplanationAn exponential model can be linearized by taking . If the regression of on is linear, the original relationship is exponential. A power model is linearized by taking of both variables.
Key takeawayRegressing $\ln(y)$ on $x$ linearizes exponential relationships; regressing $\ln(y)$ on $\ln(x)$ linearizes power relationships.
- A
- Question 11 · Hard
A statistician notes that cities with more churches tend to have more crime. Before concluding that religion causes crime, what is the most important alternative explanation?
- AThe correlation between churches and crime is not strong enough to draw any conclusion.Why not A: The issue is not the strength of the correlation but its causal interpretation.
- BPopulation size is a confounding variable — larger cities have both more churches and more crime.Correct
- CThe data should be analyzed with a test, not correlation.Why not C: The type of test is not the issue; the issue is the causal interpretation of correlation.
- DCrime causes people to build more churches for moral guidance.Why not D: This still claims a causal relationship without evidence and in the wrong direction.
ExplanationThis is a classic example of a confounding variable (lurking variable). Population is the confound: larger cities have more of everything, including churches and crime. Both variables are associated with population, creating a spurious correlation. Correlation alone cannot establish causation.
Key takeawayA confounding variable is associated with both the explanatory and response variables, creating a spurious association; correlation does not imply causation.
- A
- Question 12 · Hard
Regression output for predicting plant height (, cm) from weekly fertilizer dose (, grams) is shown below:
Predictor Coef SE Coef T P Constant 5.2 1.1 4.73 0.000 Dose 3.8 0.6 6.33 0.000 A plant received 4 grams of fertilizer per week. What is its predicted height, and what does tell us?
- APredicted height = 20.4 cm; 77% of the variation in height is explained by fertilizer dose.Correct
- BPredicted height = 20.4 cm; indicates a moderately strong positive association.Why not B: means ; calling 0.77 itself a correlation confuses and .
- CPredicted height = 20.4 cm; 23% of the variation in height is unexplained.Why not C: While 23% is unexplained, the answer misses the main interpretation of as the explained portion.
- DPredicted height = 17.4 cm; 77% of plants are predicted correctly.Why not D: Omits the constant (5.2 + 3.8×4 = 20.4, not 3.8×4 = 15.2) and misinterprets .
Explanationcm. The means 77% of the variation in plant height is explained by the linear regression on fertilizer dose. The remaining 23% is due to other factors.
Key takeawayPredicted value: substitute $x$ into $\hat{y} = a + bx$. $r^2$: proportion of $y$-variation explained by the model.
- A