Inference from sample statistics and margin of error
Problem-Solving and Data Analysis is about 15% of SAT Math (5–7 of 44 questions). Inference questions describe a survey or an experiment and ask which conclusion is supported, how to estimate a population value, or what a margin of error means.
Population, sample and bias
We usually cannot ask everyone, so we study a sample and use it to say something about the population. This works well only if the sample is random, so that every member has an equal chance of being chosen, for example by a lottery from a full list. A sample made of volunteers, of people who are easy to reach, or of one special group is biased. Asking only the people at a sports center about exercise habits tells you about people at the sports center, not about the whole town.
What each kind of study allows
Two different kinds of randomness give two different kinds of conclusions. A random sample from a population lets you generalize the result to that population (and only that one). Random assignment of subjects to treatment groups in an experiment lets you say that the treatment caused the difference. A study can have one, both or neither. If the subjects were volunteers, the result applies to the volunteers; if no one was randomly assigned, you can only say that two things are associated.
| Study has | Generalize to the population? | Cause and effect? |
|---|---|---|
| random sample only | yes | no |
| random assignment only | no | yes |
| both | yes | yes |
| neither | no | no |
From a sample to a population
If a random sample has a certain percent with a property, the best estimate is that the population has about the same percent. To estimate a count, multiply that percent by the size of the population, not the size of the sample. Suppose 60 of 200 randomly chosen students own a bicycle: that is 30%. For a school of 6,000 students the estimate is 0.30 × 6,000 = 1,800 students. It is an estimate, not an exact count.
Margin of error
A sample result is never exactly the population value, because a different random sample would give a slightly different result. The margin of error says how far off the estimate could plausibly be. If a poll gives 52% with a margin of error of 4 percentage points, the plausible values for the population percent are 52 − 4 = 48% to 52 + 4 = 56%. Every value in this interval is plausible; values outside it are not likely. Notice that 50% is inside the interval, so this poll cannot show that the percent is above 50%.
What a larger sample changes
A larger random sample gives a smaller margin of error: the estimate is more precise. The relation is not proportional: to cut the margin of error in half you need about four times as many people. The table is only a rough illustration: it uses the rule margin ≈ 1/√n (as a percent), which works for percents near 50%; you do not need this rule on the SAT. But a bigger sample does not cure bias. If the sample was chosen badly, a huge sample only gives a precise answer to the wrong question. Also, the margin of error covers only the luck of random sampling, not wrong answers or a badly chosen population.
| Sample size n | Rough margin of error (≈ 1/√n) |
|---|---|
| 100 | 1/10 = 10%: about ±10 points |
| 400 | 1/20 = 5%: about ±5 points |
| 1,600 | 1/40 = 2.5%: about ±2.5 points |
- Teachers across the whole country use online tools at about the same rate.
- About 63% of all 4,200 teachers in the region use online tools in class.
- Using online tools makes teachers more effective.
- Exactly 63% of the 4,200 teachers in the region use online tools.
- The sample is random, so the result can be applied to the population it came from: the 4,200 teachers in the region.
- A sample gives an estimate, so the word “about” is right and “exactly” is wrong.
- Nothing was tested about effectiveness, and the sample does not cover the rest of the country.
- Sample percent: 90 ÷ 250 = 0.36, or 36%.
- Apply it to the whole town: 0.36 × 18,000 = 6,480.
- 40%
- 41%
- 47%
- 52%
- Lower end: 46 − 3 = 43%. Upper end: 46 + 3 = 49%.
- The plausible values are from 43% to 49%.
- Only 47% is inside this interval.
- Yes, because the sample is so large.
- No, because the commenters may not represent all travelers, and the margin of error does not cover that.
- Yes, because 1 percentage point is a small margin.
- No, because a margin of error that small cannot be correct.
- The 5,000 people chose themselves by writing comments. They are not a random sample of all travelers.
- A margin of error measures only random sampling variation. It cannot repair a biased sample.
- A large sample does not fix bias.
- Generalizing to a bigger population than the one sampledA random sample of teachers in one region says something about teachers in that region, not about the whole country.
- Thinking a bigger sample removes biasOnly random selection removes bias. A large volunteer sample is still a volunteer sample.
- Claiming cause and effect from a surveySurveys and other studies without random assignment show association only.
- Treating the estimate as the exact population valueUse words such as “about” and “plausible”. The margin of error gives an interval of plausible values.
- Scaling with the wrong sizeTo estimate a count in the population, multiply the sample percent by the population size, not by the sample size.
- Reversing sample size and margin of errorMore people in a random sample means a smaller margin of error, and about four times the people halves it.
Desmos is a calculator here. Use it for scaling a sample result, for example 90/250 * 18000, and for interval ends such as 46 − 3 and 46 + 3. To see how the margin of error shrinks with sample size, you can try the rough rule 100/sqrt(n), which gives an approximate margin in percentage points for a sample of size n when the percent is near 50%. Four times the sample gives half the margin.
- Type 90/250 * 18000 to see 6480
- Type 100/sqrt(400) to see 5
- Type 100/sqrt(1600) to see 2.5
The hard part of these questions is deciding what the sample allows you to conclude, and Desmos cannot do that. Use it only for the arithmetic.
Open Desmos ↗