Tell a random sample from a biased one and recognize sources of bias in how data are collected.
Random sampling and bias
Random sampling and bias
To draw conclusions about a large group (the population), we study a small part of it — a sample. The key requirement: the sample must be representative, that is, reflect the properties of the whole population. The simplest way to achieve this is a random sample, in which every member of the population has the same chance of being included.
Sampling bias occurs when the selection process systematically favors some members and ignores others. Three classic sources of bias: self-selection (a volunteer survey — those who respond are the ones with something to say), survivorship bias (we analyze only the “survivors” — successful companies, planes that returned from missions — and miss those that are gone), and surveying only active users (e.g., an app's rating from those who left a review, not from everyone who downloaded it).
Bias distorts the conclusion: if only people who were annoyed by something take an online survey about service quality, the average rating will come out lower than the real one. Removing bias after the fact is almost impossible, so it's important to recognize it while the data are still being collected.
Lesson notes
Random sampling and bias
To draw conclusions about a large group (the population), we study a small part of it — a sample. The key requirement: the sample must be representative, that is, reflect the properties of the whole population. The simplest way to achieve this is a random sample, in which every member of the population has the same chance of being included.
Sampling bias occurs when the selection process systematically favors some members and ignores others. Three classic sources of bias: self-selection (a volunteer survey — those who respond are the ones with something to say), survivorship bias (we analyze only the “survivors” — successful companies, planes that returned from missions — and miss those that are gone), and surveying only active users (e.g., an app's rating from those who left a review, not from everyone who downloaded it).
Bias distorts the conclusion: if only people who were annoyed by something take an online survey about service quality, the average rating will come out lower than the real one. Removing bias after the fact is almost impossible, so it's important to recognize it while the data are still being collected.