๐Ÿ“Š Statistics Steppe ยท Statistics

Samples & Inference

Use a sample to say something about a population you cannot count, and be honest about how much the answer could be out by.

In short

  • A sample stands in for a population only if every member had a fair chance of being chosen.
  • Scale a sample up with estimate = sample share x population size, and report it as "about".
  • Capture-recapture sets the share marked in the second catch equal to the share marked in the whole population.
  • The margin of error shrinks with the SQUARE ROOT of the sample size, so quadrupling the sample only halves the wobble.
  • Only a randomised comparative experiment supports a claim about cause; a survey and an observational study can report that two things go together and no more.

Asking a few to learn about many

Nobody can count every frog in a marsh or ask every traveller on a steppe. So you take a sample: a small group you *can* measure, chosen to stand in for the whole population.

The whole method rests on one assumption โ€” that the sample is a fair miniature of the population. When that holds, a share found in the sample can be scaled up:

estimate = (sample share) x (population size)

Sample 40 frogs, find 14 spotted, and the share is 14/40 = 0.35. For a population of 600, estimate 0.35 x 600 = 210 spotted frogs.

Say "about". An estimate carries the sample's uncertainty with it, and reporting it as an exact number claims more than the evidence supports.

What makes a sample fair

A sample is random when every member of the population has the same chance of being in it. That is a property of the *method*, not of the result.

Numbering every household and drawing forty numbers from a sealed jar is random. Taking every tenth name from an alphabetical roll, starting from a random point, is fine too, because alphabetical order has nothing to do with the question.

Two things go wrong, and they are different faults:

  • Bias. The method reaches a particular group. Asking only the people queuing for a ferry whether another ferry is needed, or inviting only the unhappy to come and complain, produces a sample that leaned before a single answer was given. A bigger sample does not fix bias; it just makes the wrong answer more precise.
  • Too small. The method is fair but three people cannot speak for a steppe.

Diagnose which fault you are looking at before you suggest a remedy.

Surveys, experiments and observational studies

A sample tells you what *is*. To claim what causes what, the design has to do more. Three designs cover almost every study, and they differ in one thing only: who decided the condition.

  • A sample survey draws a sample from a population and asks it something. Nobody is set against anybody, and the result describes the population.
  • An observational study sets two groups that were already there beside each other โ€” the growers who *already* water twice a day against the rest. The surveyors only record.
  • An experiment hands out the condition itself. The subjects are sorted into groups, one group receives the treatment, another is kept at the usual condition as the control group, and the response is measured at the end.

The word that separates the last two is assignment. In an observational study the subjects chose their own group, and whatever led them to choose it walks into the result with them. In an experiment the surveyors sort them, and if they sort them at random then every other difference โ€” age, habit, weather, keenness โ€” is scattered evenly between the groups. That is what leaves the treatment as the only thing available to explain a difference.

So the two designs are entitled to different sentences:

  • Observational study: "these two things go together." Never "this one produced that one."
  • Randomised experiment: "this treatment produced this difference."

Random sampling and random assignment are different jobs, and it is worth keeping them apart. Random sampling decides who is in the study at all, and it is what lets the result speak for a wider population. Random assignment decides which group each subject lands in, and it is what lets the result speak of cause. A study can have one, both or neither.

Three things break a design, and a bigger study mends none of them:

  • Self-selection. The subjects volunteered, so only the keenest are measured.
  • No control group. Everyone was treated, so an ordinary good spell looks exactly like an effect.
  • Confounding. Two things were changed at once, so no result can say which of them did the work.

Capture-recapture

To estimate a population you cannot line up, mark some of it and see how many marks come back.

Catch 60 carp, tag them, release them. Later catch 50 carp and find 6 already tagged. If the tagged fish mixed back in evenly, the share tagged in the second catch should match the share tagged in the whole lake:

6/50 = 60/N

Cross-multiply: 6N = 60 x 50 = 3000, so N = 500 carp.

Sense-check the shape of the answer. The estimate must be larger than the number you tagged, and the smaller the fraction that came back marked, the larger the population must be. If your answer came out below 60, the proportion was set up upside down.

How far out could it be?

Two fair samples from the same population will not agree exactly, and neither will agree exactly with the truth. The margin of error puts a size on that wobble.

A classroom rule of thumb: for a sample of n people, the margin of error is roughly 1/sqrt(n), expressed as a percent.

  • n = 100: margin about 1/10 = 10%
  • n = 400: margin about 1/20 = 5%
  • n = 2500: margin about 1/50 = 2%

If 52% of a sample of 400 use the north road, the plausible interval is 52% +/- 5%, that is 47% to 57%.

Notice the square root. Four times the sample halves the margin; to halve it again you need four times as many again. That is why serious surveys stop at a few thousand people โ€” the next improvement costs far more than it is worth.

And pooling beats averaging. To combine several samples of different sizes, add the counts and divide once; averaging the separate percentages gives the small samples the same say as the large ones.

Worked examples

Example 1

A ranger tags 80 hares and releases them. Later she catches 60 hares and finds 5 already tagged. Estimate the hare population.

  1. The share tagged in the second catch is 5/60.
  2. If the tagged hares mixed back in evenly, that should match the share tagged in the whole population: 80/N.
  3. So 5/60 = 80/N.
  4. Cross-multiply: 5N = 80 x 60 = 4800, so N = 4800 / 5 = 960.
  5. About 960 hares. Check the shape: that is much larger than the 80 tagged, as it must be, since only one hare in twelve came back marked.

Example 2

Four surveyors sampled 25, 40, 50 and 100 travellers and found 10, 18, 22 and 30 carrying a map. What is the best single estimate of the percent carrying a map?

  1. Pool the samples rather than averaging the four percentages, because the samples are different sizes.
  2. Total sampled = 25 + 40 + 50 + 100 = 215.
  3. Total carrying a map = 10 + 18 + 22 + 30 = 80.
  4. Pooled share = 80 / 215 = 0.3721..., so about 37.2%.
  5. The plain average of the four percentages (40%, 45%, 44%, 30%) is 39.75%, which is higher because it gives the 25-person sample as much say as the 100-person one.

Example 3

One field is sown with the new seed and watered twice a day. The field beside it keeps the old seed and its usual watering, and it yields less. Can the growers claim the new seed raised the yield?

  1. Ask first who decided the condition. The growers did, so this is not an observational study โ€” something was handed out.
  2. Ask next whether the groups were sorted at random. They were not: one whole field received both changes and the other received neither.
  3. Two things differ between the fields, the seed and the watering, so they are confounded โ€” no result can say which of them did the work.
  4. So the claim is not justified. The honest report is that the new field yielded more, and that two changes were made to it at once.
  5. A design that would support the claim: sort many plots into two groups at random, sow the new seed in one group only, and water every plot alike.

Practice problems, with solutions

Three problems of increasing difficulty, each with the full working. In the game these are generated fresh every time; these three are fixed so this page always shows the same ones.

Problem 1

Difficulty 1 of 5

Nadia caught a random sample of 20 bog carp and found 5 of them silver-scaled. The marsh holds about 120 bog carp in all. Estimate how many are silver-scaled.

Answer: 30 bog carp

  1. Sample proportion = 5/20 = 0.25 (25%).
  2. Estimate = 5/20 x 120
  3. = 5 x 6 = 30
  4. So about 30 bog carp โ€” an estimate carrying the sample's uncertainty with it.

Problem 2

Difficulty 3 of 5

A surveyor picks every tenth name from an alphabetical roll of all the marsh guilds, starting from a name drawn at random. Is this sample likely to represent the whole population?

  1. yes โ€” the method gives everyone a fair chance of being chosen
  2. no โ€” the method reaches a particular group and leans that way
  3. no โ€” the sample is far too small to speak for the population

Answer: A. yes โ€” the method gives everyone a fair chance of being chosen

  1. Consider how the sample was chosen: A surveyor picks every tenth name from an alphabetical roll of all the marsh guilds, starting from a name drawn at random.
  2. Here an alphabetical roll has nothing to do with the question, so taking every tenth name spreads the sample evenly across the whole population.
  3. So the honest answer is: yes โ€” the method gives everyone a fair chance of being chosen.

Problem 3

Difficulty 4 of 5

Kofi caught 40 bog carp, tagged them all, and let them go. Later Kofi caught 80 bog carp and found 8 of them already tagged. Estimate how many bog carp live in the marsh.

Answer: 400 bog carp

  1. Share marked in the second catch = 8/80.
  2. That should equal the share marked in the whole marsh: 40/N.
  3. 8/80 = 40/N, so N = 40 x 80 / 8
  4. N = 400 bog carp.

Common mistakes

  • Reporting the count found in the sample as if it were the population figure.
  • Calling a tidy but biased method fair โ€” a self-selected or convenient group is not a random one.
  • Setting the capture-recapture proportion upside down, giving a population smaller than the number tagged.
  • Using n rather than the square root of n in the margin of error, which makes the survey look far more precise than it is.
  • Averaging the percentages of several differently sized samples instead of pooling the raw counts.
  • Reading an observational study as though it were an experiment: the groups chose the condition for themselves, so whatever led them to choose it explains the difference just as well.
  • Confusing random sampling with random assignment โ€” drawing the sample fairly says who the result describes, while sorting the groups at random is what lets the result speak of cause.

What you should be able to do

  • Estimate a population total from the proportion found in a random sample.
  • Judge whether a sampling method is likely to be representative.
  • Estimate a population size with the capture-recapture method.
  • State a margin of error and the interval of plausible values around an estimate.

Where this fits in the curriculum

Common Core

  • HSS-IC.B.3

    High school โ€” Recognise the purposes of and differences among sample surveys, experiments and observational studies; explain how randomisation relates to each.

  • HSS-IC.A.1

    High school โ€” Understand statistics as a process for making inferences about population parameters based on a random sample from that population.

  • 7.SP.A.1

    Grade 7 โ€” Understand that statistics can gain information about a population by examining a sample, and that a random sample tends to be representative.

  • 7.SP.A.2

    Grade 7 โ€” Use data from a random sample to draw inferences about a population with an unknown characteristic of interest.

  • HSS-IC.B.4

    High school โ€” Use data from a sample survey to estimate a population mean or proportion, and develop a margin of error through the use of simulation models.

    The Common Core develops the margin of error through simulation; the 1/sqrt(n) figure this skill uses is the classroom rule of thumb that approximates it.

Ontario

  • MTH1W.D1

    Grade 9 de-streamed โ€” Describe the collection and use of data, and represent and analyse data involving one and two variables.

    Ontario numbers its Data expectations per grade document and the specific numbering could not be verified line by line here, so this tag names the strand's OVERALL expectation rather than guessing at a specific one.

SAT

Learn these first