Two-variable data: models and scatterplots
Problem-Solving and Data Analysis is about 15% of SAT Math (5–7 of 44 questions). Two-variable data questions give a scatterplot, a table or a model equation and ask you to read it, predict from it, interpret its numbers or choose the best model.
Reading a scatterplot
Each dot is one case: for example one person, with a value of x and a value of y. If the dots rise from left to right, the association is positive: larger x goes with larger y. If they fall, it is negative. If they form a cloud with no direction, there is no association. The closer the dots are to a line (or other curve), the stronger the association. A line of best fit runs through the middle of the dots, with about as many dots above it as below it.
What the slope and the intercept mean
A line of best fit is written ŷ = mx + b, where ŷ (“y-hat”) is the predicted value of y. The slope m is the predicted change in y when x increases by 1. The y-intercept b is the predicted value of y when x = 0; it only has a real meaning if x = 0 makes sense in the story. In the graph above the line is ŷ = 3x + 10: each extra hour of practice goes with 3 more points of predicted score, and the model predicts a score of 10 for no practice.
Predictions and residuals
To predict, put the x-value into the equation. A residual compares a real dot with the line: residual = actual y − predicted y. A dot above the line has a positive residual, and a dot below the line has a negative residual. In the graph, the dot at x = 3 has actual y = 21 and predicted y = 3(3) + 10 = 19, so its residual is 21 − 19 = +2.
Inside or outside the data?
A prediction for an x-value between the smallest and largest x in the data is an interpolation and is usually reasonable. A prediction far outside that range is an extrapolation and can be nonsense: a line of best fit for a child’s height at ages 2 to 10 would predict a height of several meters at age 40. Also remember that association is not causation. A scatterplot can show that two quantities move together, but it cannot prove that one causes the other.
Linear, exponential or quadratic?
Choose the model that matches the pattern. Dots along a straight band suggest a linear model. Dots that curve upward faster and faster, or whose outputs are multiplied by the same factor, suggest an exponential model. Dots that rise and then fall (or fall and then rise) suggest a quadratic model. With a table, check differences for a linear model and ratios for an exponential one.
| Pattern in the data | Best model |
|---|---|
| a straight band of dots | linear: y = mx + b |
| outputs multiplied by the same factor | exponential: y = a·b^x |
| dots rise then fall (or the reverse) | quadratic: y = ax² + bx + c |
- The seedling was 4 centimeters tall when it was planted.
- The model predicts that the seedling grows 4 centimeters each week.
- The model predicts that the seedling is 4 weeks old when it is 6 centimeters tall.
- The seedling was measured for 4 weeks.
- The slope is the predicted change in y for 1 more unit of x.
- Here x is weeks and y is centimeters, so the slope 4 means 4 centimeters more for each extra week.
- The number 6 is the intercept: the predicted height at planting.
- Predicted price: y = −1.6(5) + 24 = −8 + 24 = 16 thousand dollars.
- Residual = actual − predicted = 17.5 − 16 = +1.5 thousand dollars.
- The car was sold for more than the model predicts, so its dot is above the line.
| x | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| y | 100 | 80 | 64 | 51.2 |
- y = 100 − 20x
- y = 100(0.8)^x
- y = 100(1.25)^x
- y = 80x + 100
- The differences are −20, −16 and −12.8: not equal, so the model is not linear.
- The ratios are 80/100 = 64/80 = 51.2/64 = 0.8: equal, so the model is exponential decay.
- With a = 100 and b = 0.8: y = 100(0.8)^x. Check x = 3: 100 × 0.512 = 51.2 ✓
- y = −1.6(20) + 24 = −32 + 24 = −8 thousand dollars.
- A price cannot be negative, so the prediction is not reasonable.
- The age 20 is far outside the data (1 to 10 years). This is an extrapolation, and the line does not describe such old cars.
- Reading the slope as the total changeThe slope is the change in y per 1 unit of x, not the change over the whole data set.
- Explaining the intercept when x = 0 makes no senseThe intercept is the predicted y at x = 0. If x = 0 is outside the data or impossible, say only that it is a starting value of the model.
- Subtracting residuals the wrong wayResidual = actual − predicted. A dot below the line has a negative residual.
- Trusting predictions far outside the dataPredictions for x-values inside the range of the data are more trustworthy. Far outside, the line can give impossible values.
- Claiming causationA strong association in a scatterplot does not prove that x causes y.
- Calling every increasing pattern linearCheck whether the differences or the ratios are constant before choosing the model.
Desmos can plot the data and the model together. Type the table by clicking + and choosing table, then enter x₁ and y₁. To see a model against the dots, type its equation, for example y = 4x + 6, and use sliders for m and b: y = mx + b. Desmos can also compute the line of best fit for you with a regression command such as y₁ ~ m x₁ + b, and an exponential fit with y₁ ~ a b^x₁; this is an advanced tool and you will not need it on most questions.
- Enter the points in a table
- Type y = mx + b and accept the sliders
- Move m and b until the line follows the dots
- Type y = 4x + 6 and click on the line to read predicted values
Use Desmos to graph a model against data or to evaluate it at ugly numbers. Reading slope and intercept in words is a skill you do by hand.
Open Desmos ↗