Pepelen
← Statistics and Probability from Scratch

Lesson

Lesson 3: Outliers, resistance, and sample vs. population

Distinguish a sample (n) from a population (N), understand Bessel's correction (n−1), and choose measures that are resistant to outliers.

1 / 6

Sample, population, and resistance to outliers

Sample, population, and resistance to outliers

Statistics draws a clear line between two concepts. A population is everything we want to draw conclusions about (for example, all residents of a country). A sample is a subset of n observations that we actually measure. We usually work with a sample and use it to judge the population. Bessel's correction: why n−1? When you compute the variance from a sample, dividing by n underestimates it: the sample is “pulled in” toward its own mean and misses part of the population's spread. Dividing by n−1 corrects this bias and gives an unbiased estimate of the population variance. Most calculators and sample formulas use n−1. For [2, 4, 4, 4, 5, 5, 7, 9], the sample variance = 32 / (8−1) = 32 / 7 ≈ 4.57, while the population variance = 32 / 8 = 4. Outliers and the 1.5·IQR rule. A value is an outlier if it lies outside the interval [Q1 − 1.5·IQR, Q3 + 1.5·IQR]. For our data set, IQR = 2, Q1 = 4, Q3 = 6: lower fence = 4 − 3 = 1, upper fence = 6 + 3 = 9. Every value falls inside, so there are no outliers. But in the salaries [20, 22, 24, 25, 26, 28, 300], 300 is clearly outside: it's an outlier. Resistance (robustness): the median and the IQR barely change when an outlier is added — they're resistant. The mean, range, and SD react strongly — they're not resistant. For noisy or skewed data, prefer resistant measures: the median for center, the IQR for spread.
Lesson notes
Sample, population, and resistance to outliers
Statistics draws a clear line between two concepts. A population is everything we want to draw conclusions about (for example, all residents of a country). A sample is a subset of n observations that we actually measure. We usually work with a sample and use it to judge the population. Bessel's correction: why n−1? When you compute the variance from a sample, dividing by n underestimates it: the sample is “pulled in” toward its own mean and misses part of the population's spread. Dividing by n−1 corrects this bias and gives an unbiased estimate of the population variance. Most calculators and sample formulas use n−1. For [2, 4, 4, 4, 5, 5, 7, 9], the sample variance = 32 / (8−1) = 32 / 7 ≈ 4.57, while the population variance = 32 / 8 = 4. Outliers and the 1.5·IQR rule. A value is an outlier if it lies outside the interval [Q1 − 1.5·IQR, Q3 + 1.5·IQR]. For our data set, IQR = 2, Q1 = 4, Q3 = 6: lower fence = 4 − 3 = 1, upper fence = 6 + 3 = 9. Every value falls inside, so there are no outliers. But in the salaries [20, 22, 24, 25, 26, 28, 300], 300 is clearly outside: it's an outlier. Resistance (robustness): the median and the IQR barely change when an outlier is added — they're resistant. The mean, range, and SD react strongly — they're not resistant. For noisy or skewed data, prefer resistant measures: the median for center, the IQR for spread.
Lesson 3: Outliers, resistance, and sample vs. population — Statistics and Probability from Scratch