
About this chapter
Measures of location and spread is the calculation-heavy start of Statistics 1: means, medians, quartiles and standard deviations from lists, frequency tables and grouped data. The arithmetic is done on a calculator, so the marks go to setting it up correctly, and in recent examiner reports to coding, where students regularly scaled a variance by the wrong factor.
This chapter covers class boundaries and midpoints, quartiles for discrete data, linear interpolation for grouped data, variance and standard deviation, coding, combining two data sets, correcting a wrongly recorded value, and choosing between the mean and the median. The longer questions work backwards from summary statistics and ask how the statistics change when a value is added or corrected.
The ten questions
- Q1 (6 marks): median, interquartile range, mean and standard deviation of goals from a frequency table.
- Q2 (5 marks): class boundaries for heights to the nearest cm, an estimated mean, and a median by interpolation.
- Q3 (5 marks): mean and standard deviation from Σx and Σx², then the corrected mean after one time was misrecorded.
- Q4 (6 marks): apple masses coded with y = (x − 200)/5: decode the mean, standard deviation and variance.
- Q5 (5 marks): combine two classes' means and standard deviations into one mean and one standard deviation.
- Q6 (12 marks): grouped potato masses: median and quartiles by interpolation, mean and standard deviation, and more.
- Q7 (11 marks): coded parcel weights: show the mean, find the standard deviation, recover Σw and Σw², then add a parcel equal to the mean.
- Q8 (10 marks): waiting times at a bank: estimates from grouped data, then test the bank's claim about the percentage who wait.
- Q9 (11 marks): a table with an unknown frequency p fixed by a mean of exactly 3, then the median, quartiles and spread.
- Q10 (12 marks): two filling machines combined into one data set of 50 bags, then a bag found to be misrecorded.
Key skills tested
Types of data and classes. Discrete data is counted and continuous data is measured. For heights of 150 to 159 cm to the nearest cm, the class boundaries are 149.5 and 159.5.
Averages from tables. The mean is Σfx ÷ Σf, using class midpoints for grouped data, which makes it an estimate.
Quartiles for discrete data. If n/4 is not a whole number, round up to find the position; if it is whole, go halfway between that value and the next. The same rule works for n/2 and 3n/4.
Interpolation. For grouped data, the median is the (n/2)th value: the lower class boundary plus the fraction of the class needed, times the class width.
Variance and standard deviation. The variance is Σx²/n minus the mean squared, and for a table it is Σfx²/Σf minus the mean squared. The standard deviation is its square root.
Coding. If y = (x − a)/b, then the mean of x is b times the mean of y, plus a, and the standard deviation of x is b times that of y. Variances scale by b².
Combining data sets. Work with totals: Σx is n times the mean, and Σx² is n times (σ² plus the mean squared). Add the totals, then recalculate.
Changing the data. Adding a value equal to the mean leaves the mean unchanged but reduces the standard deviation. To correct a value, adjust Σx and Σx².
Choosing a measure. The median and interquartile range resist extreme values, while the mean uses all the data. Always interpret in context.

Worked example
Question 1 from this chapter. The goals scored in 40 matches are: 0 goals 7 times, 1 goal 12 times, 2 goals 10 times, 3 goals 6 times, 4 goals 3 times and 5 goals twice. Find (a) the median and interquartile range and (b) the mean and standard deviation.
(a) The cumulative frequencies are 7, 19, 29, 35, 38, 40. With n = 40, the median is halfway between the 20th and 21st values, which are both 2, so the median is 2. Q1 uses the 10th and 11th values, both 1, and Q3 the 30th and 31st, both 3. So the interquartile range is 3 − 1 = 2.
(b) Σfx = 72, so the mean is 72 ÷ 40 = 1.8. Σfx² = 204, so the variance is 204 ÷ 40 − 1.8² = 5.1 − 3.24 = 1.86, and the standard deviation is √1.86 = 1.36.
The mark scheme names the two common slips: forgetting to subtract the mean squared, and giving the variance as the standard deviation.
Seen on real papers
Every question in our booklets is original. These are the recent papers where each question type has appeared.
- June 2024, Q3(b): the median of grouped data by linear interpolation, as in Q2 and Q6.
- June 2024, Q3(c) to (e): coded data decoded into a mean and standard deviation, then the effect of adding new values, as in Q4 and Q7.
- June 2022, Q3(e): decoding a mean and a variance from a coding with an assumed mean, as in Q4 and Q10.
Where marks are lost
- Variance scaling. In June 2024 many did not realise that multiplying data by a scales the variance by a², and in June 2022 some scaled by 0.5 instead of 0.5², or multiplied when they should have divided.
- Decoding a total. In June 2024, adding the coding constant only once, instead of once for every value, when turning Σy back into Σw cost both marks.
- Adding a value at the mean. In June 2024 too many said the mean would rise or fall. It stays the same, while the standard deviation decreases because the new value adds no spread.
- Too few figures. The front of the paper asks for 3 significant figures, and rounding intermediate values earlier than that loses accuracy marks.
Common questions
Should I use n or n − 1 for the variance?
Use n. Statistics 1 defines the variance as Σx²/n minus the mean squared.
Why is the mean from grouped data only an estimate?
Because every value in a class is replaced by the class midpoint.
When is the median better than the mean?
When the data has extreme values or is skewed, because the median is not pulled by them. The mean uses all the data.
Where this chapter leads
- Statistics 1 Chapter 3: Representations of Data: quartiles and outliers on box plots and histograms.
- Statistics 1 Chapter 5: Correlation and Regression: coding and summary statistics such as Sxx.

