☰ SAT · Math

Two-variable data: models and scatterplots

SAT Math · Problem-Solving and Data Analysis · Week 8 of the 12-week plan
15

Two-variable data: models and scatterplots

College Board skill: Two-variable data: models and scatterplots
GoalRead a scatterplot, describe the association, use a line of best fit to predict, explain what its slope and intercept mean, find residuals, and choose between a linear and an exponential model.
On the test

Problem-Solving and Data Analysis is about 15% of SAT Math (5–7 of 44 questions). Two-variable data questions give a scatterplot, a table or a model equation and ask you to read it, predict from it, interpret its numbers or choose the best model.

Key words
scatterplot · a graph in which each dot is one case, with its x-value and y-valueline of best fit · the straight line that follows the trend of the dots as closely as possibleresidual · actual y minus predicted y: how far a dot is above (+) or below (−) the lineextrapolation · using a model far outside the x-values of the data; it is unreliable
Explanation

Reading a scatterplot

Each dot is one case: for example one person, with a value of x and a value of y. If the dots rise from left to right, the association is positive: larger x goes with larger y. If they fall, it is negative. If they form a cloud with no direction, there is no association. The closer the dots are to a line (or other curve), the stronger the association. A line of best fit runs through the middle of the dots, with about as many dots above it as below it.

hours of practicescore123456789102030400

What the slope and the intercept mean

A line of best fit is written ŷ = mx + b, where ŷ (“y-hat”) is the predicted value of y. The slope m is the predicted change in y when x increases by 1. The y-intercept b is the predicted value of y when x = 0; it only has a real meaning if x = 0 makes sense in the story. In the graph above the line is ŷ = 3x + 10: each extra hour of practice goes with 3 more points of predicted score, and the model predicts a score of 10 for no practice.

ƒFormula
ŷ = m·x + b
m = predicted change in y for 1 more unit of x; b = predicted y when x = 0

Predictions and residuals

To predict, put the x-value into the equation. A residual compares a real dot with the line: residual = actual y − predicted y. A dot above the line has a positive residual, and a dot below the line has a negative residual. In the graph, the dot at x = 3 has actual y = 21 and predicted y = 3(3) + 10 = 19, so its residual is 21 − 19 = +2.

xy01234568121620242832actual 21predicted 19

Inside or outside the data?

A prediction for an x-value between the smallest and largest x in the data is an interpolation and is usually reasonable. A prediction far outside that range is an extrapolation and can be nonsense: a line of best fit for a child’s height at ages 2 to 10 would predict a height of several meters at age 40. Also remember that association is not causation. A scatterplot can show that two quantities move together, but it cannot prove that one causes the other.

Linear, exponential or quadratic?

Choose the model that matches the pattern. Dots along a straight band suggest a linear model. Dots that curve upward faster and faster, or whose outputs are multiplied by the same factor, suggest an exponential model. Dots that rise and then fall (or fall and then rise) suggest a quadratic model. With a table, check differences for a linear model and ratios for an exponential one.

Pattern in the dataBest model
a straight band of dotslinear: y = mx + b
outputs multiplied by the same factorexponential: y = a·b^x
dots rise then fall (or the reverse)quadratic: y = ax² + bx + c
Worked examples
Example 1.
A scientist measures the height y, in centimeters, of a seedling x weeks after planting. The line of best fit is y = 4x + 6. Which statement is the best interpretation of the slope, 4?
weekscm12345678102030400
  1. The seedling was 4 centimeters tall when it was planted.
  2. The model predicts that the seedling grows 4 centimeters each week.
  3. The model predicts that the seedling is 4 weeks old when it is 6 centimeters tall.
  4. The seedling was measured for 4 weeks.
  1. The slope is the predicted change in y for 1 more unit of x.
  2. Here x is weeks and y is centimeters, so the slope 4 means 4 centimeters more for each extra week.
  3. The number 6 is the intercept: the predicted height at planting.
Trap: The first choice describes the intercept 6 (and gets the number wrong).
Example 2.
The line of best fit for the price y, in thousands of dollars, of a used car that is x years old is y = −1.6x + 24. A particular 5-year-old car is sold for 17.5 thousand dollars. What price does the model predict for a 5-year-old car, and what is the residual for this car?
age, yearsprice12345678910111248121620240y = −1.6x + 24actual 17.5predicted 16
  1. Predicted price: y = −1.6(5) + 24 = −8 + 24 = 16 thousand dollars.
  2. Residual = actual − predicted = 17.5 − 16 = +1.5 thousand dollars.
  3. The car was sold for more than the model predicts, so its dot is above the line.
Example 3.
The table shows the number y of a plant’s leaves that are healthy after x weeks of a long drought. Which equation best models the data?
x0123
y100806451.2
  1. y = 100 − 20x
  2. y = 100(0.8)^x
  3. y = 100(1.25)^x
  4. y = 80x + 100
  1. The differences are −20, −16 and −12.8: not equal, so the model is not linear.
  2. The ratios are 80/100 = 64/80 = 51.2/64 = 0.8: equal, so the model is exponential decay.
  3. With a = 100 and b = 0.8: y = 100(0.8)^x. Check x = 3: 100 × 0.512 = 51.2 ✓
Trap: y = 100 − 20x matches the first two columns (100 and 80) but gives 60 at x = 2, not 64.
Example 4.
The model y = −1.6x + 24 in the car example was built from cars that are 1 to 10 years old. What does it predict for a 20-year-old car? Is this prediction reasonable?
  1. y = −1.6(20) + 24 = −32 + 24 = −8 thousand dollars.
  2. A price cannot be negative, so the prediction is not reasonable.
  3. The age 20 is far outside the data (1 to 10 years). This is an extrapolation, and the line does not describe such old cars.
Common traps
  • Reading the slope as the total changeThe slope is the change in y per 1 unit of x, not the change over the whole data set.
  • Explaining the intercept when x = 0 makes no senseThe intercept is the predicted y at x = 0. If x = 0 is outside the data or impossible, say only that it is a starting value of the model.
  • Subtracting residuals the wrong wayResidual = actual − predicted. A dot below the line has a negative residual.
  • Trusting predictions far outside the dataPredictions for x-values inside the range of the data are more trustworthy. Far outside, the line can give impossible values.
  • Claiming causationA strong association in a scatterplot does not prove that x causes y.
  • Calling every increasing pattern linearCheck whether the differences or the ratios are constant before choosing the model.
The Desmos way

Desmos can plot the data and the model together. Type the table by clicking + and choosing table, then enter x₁ and y₁. To see a model against the dots, type its equation, for example y = 4x + 6, and use sliders for m and b: y = mx + b. Desmos can also compute the line of best fit for you with a regression command such as y₁ ~ m x₁ + b, and an exponential fit with y₁ ~ a b^x₁; this is an advanced tool and you will not need it on most questions.

  1. Enter the points in a table
  2. Type y = mx + b and accept the sliders
  3. Move m and b until the line follows the dots
  4. Type y = 4x + 6 and click on the line to read predicted values

Use Desmos to graph a model against data or to evaluate it at ugly numbers. Reading slope and intercept in words is a skill you do by hand.

Open Desmos ↗
Quick check
1
A line of best fit is ŷ = 5x + 20. What does the slope mean?
2
For the line ŷ = 0.5x + 3, what y does the model predict for x = 14?
3
A data point has actual y = 30 and predicted y = 26. What is its residual?
4
The dots curve upward and each y is about 1.5 times the one before. Which model fits: linear or exponential?
Practice set: 10 SAT-style questionsEasy → hard, with typed answers like the real test. Your score is saved in your cabinet.
Start practice →

More official practice