MathBored

Virginia SOL Mathematics Textbook

Workbook pagesAnswer key

Chapter 16 — The Data Cycle and Boxplots

Standard: 8.PS.2 — The student will apply the data cycle (formulate questions; collect or acquire data; organize and represent data; and analyze data and communicate results) with a focus on boxplots.

By the end of this chapter you will be able to:

Lessons: 16.1 The Data Cycle, Pointed at Boxplots · 16.2 Asking the Question, Getting the Data, and Statistical Bias · 16.3 The Five-Number Summary · 16.4 Building and Reading a Boxplot · 16.5 Extreme Data Points · 16.6 Comparing Boxplots, Choosing a Display, and Spotting a Misleading One

Numbering note. Item numbers run straight through the chapter, from 1 in Lesson 16.1 to 150 at the end of the review. They do not restart at each lesson.

Quartile convention used in this book. Different textbooks compute quartiles differently, so this one states its rule once and uses it everywhere. Put the data in order. The median splits the set into a lower half and an upper half. When the count is odd, the median itself belongs to neither half. The lower quartile is the median of the lower half, and the upper quartile is the median of the upper half. Every boxplot, answer, and figure in this chapter follows that rule. A graphing calculator or spreadsheet may use a different interpolation rule and report a slightly different quartile; that is a difference in convention, not an error, but on Virginia assessments and in this book use the rule stated here.

What a boxplot does not show. A boxplot shows five numbers and the gaps between them. It does not show the individual data values, it does not show how many values there are, and it does not show the mean. Two very different data sets can have the same boxplot. Keep this in mind every time you are tempted to say "the middle line is the average" — it is the median, and that is a different statistic.

Size limit. Every data set in this chapter has no more than 20 items, which is the limit the standard places on Grade 8 work.


Lesson 16.1 — The Data Cycle, Pointed at Boxplots

The same four stages as last year

In Grade 6 and Grade 7 you learned the data cycle — the four-stage process statisticians repeat whenever they want to answer a question using data (the measurements, counts, or responses gathered to answer that question). The four stages have the same names this year:

The four stages of the data cycle arranged in a loop with arrows returning to the start

The cycle is drawn as a loop because stage 4 usually hands you a new question, which starts stage 1 again.

What is different this year

Only the display in stage 3 changes. Grade 6 used tables, dot plots, and circle graphs. Grade 7 used histograms, which group values into intervals and show the shape of the distribution. Grade 8 uses boxplots, which throw away every individual value and show exactly five numbers: the smallest value, the lower quartile, the median, the upper quartile, and the largest value.

That sounds like a loss, and in one sense it is — you cannot recover the data from a boxplot. What you get in exchange is a display so compact that two or three of them stack on one axis, which makes comparing groups easy in a way no histogram is. Most of the reason Grade 8 studies boxplots is that comparison.

A second change: the questions you formulate this year should be ones a boxplot can answer. A boxplot answers questions about center ("what is a typical value?"), spread ("how much do the values vary?"), and position ("is this value unusually high?"). It cannot answer "how many students said 20 minutes?" — that is a dot-plot or histogram question.

One project, four stages

Here is the whole cycle on one question, the way you will run it yourself.

Stage 1 — Formulate. Nadia wonders, "How long do eighth graders at my school spend reading for pleasure on a school night?" This is a statistical question: it expects a set of numbers that vary, not one fixed answer.

Stage 2 — Collect. The data she needs is one number per student: minutes spent reading last night. She surveys 15 eighth graders chosen at random from the school roster and records:

12, 15, 18, 20, 22, 25, 28, 30, 32, 35, 38, 40, 45, 48, 5512,\ 15,\ 18,\ 20,\ 22,\ 25,\ 28,\ 30,\ 32,\ 35,\ 38,\ 40,\ 45,\ 48,\ 55

Stage 3 — Organize and represent. She orders the values (already done above), finds the five-number summary, and draws a boxplot.

Stage 4 — Analyze and communicate. She reports that a typical eighth grader reads about 3030 minutes, that the middle half of students read between 2020 and 4040 minutes, and that the whole group ranges from 1212 to 5555 minutes. Then she notices the upper whisker is long, which raises a new question: are the students at the high end in a book club? Stage 1 begins again.

Worked examples

Example 1 — Naming the stage

"A coach downloads last season's game scores from the league website." Which stage is this?

Downloading data someone else gathered is acquiring data, which is part of stage 2.

Answer: Stage 2, collect or acquire data

Example 2 — A question a boxplot can answer

Rewrite "Do students like the new lunch menu?" as a question a boxplot could answer.

The original question produces opinions, not numbers, so no boxplot is possible. Ask for something measured.

Answer: "How many minutes do students spend in the lunch line?" or "How many of the 5 lunch items do students eat?" — either produces a numeric data set that a boxplot can summarize.

Example 3 — Choosing the display in stage 3

Two eighth-grade classes each recorded how many push-ups they could do. The question is whether one class is stronger overall. Which display?

The question compares two groups on center and spread, which is exactly what a pair of boxplots on one axis shows.

Answer: Two boxplots drawn on the same scale, one for each class

Example 4 — The loop

Nadia's boxplot shows a long upper whisker. Describe the next turn of the cycle.

Stage 4 produced a new question — why do a few students read so much longer? — so the cycle returns to stage 1: formulate "Do students in a book club read longer on a school night?" Then collect data from book-club and non-book-club students, represent both groups as boxplots, and compare.

Answer: The new question restarts the cycle at stage 1

Guided practice

  1. Name the four stages of the data cycle in order.
  2. Which stage includes computing the five-number summary?
  3. Which stage includes deciding whom to survey?
  4. A student says the cycle is finished the moment the boxplot is drawn. Correct the statement in one sentence.
  5. Name two things a boxplot shows and two things it does not show.
  6. Which stage does "acquiring existing data from a website" belong to?

Independent practice

  1. Classify each activity by stage: a) writing the survey question; b) sorting 14 numbers from least to greatest; c) emailing your conclusion to the principal; d) recording 18 measurements in a notebook.
  2. Explain in one sentence why the data cycle is drawn as a loop.
  3. Write a statistical question about your class that a boxplot could summarize.
  4. Write a question about your class that a boxplot could not answer, and say why not.
  5. For your question in item 9, state exactly what one piece of data you would record from each person.
  6. A boxplot is drawn from 16 test scores. Can you tell from the boxplot how many students scored exactly 8585? Explain.
  7. Can you find the mean of a data set from its boxplot? Explain.
  8. Application. A school nurse wants to know whether eighth graders get enough sleep. Run stages 1 and 2 on paper: state the question, state the data needed, and state how you would collect it from no more than 20 students.
  9. Reasoning. Grade 7 used histograms and Grade 8 uses boxplots for the same four-stage cycle. Name one thing a boxplot does better than a histogram, and one thing a histogram does better.

Exit ticket 16.1

  1. Name the four stages of the data cycle in order.
  2. In which stage is the boxplot actually drawn?
  3. State one question about center and one question about spread that a boxplot can answer.
  4. Give one reason a finished boxplot might send you back to stage 1.

Lesson 16.2 — Asking the Question, Getting the Data, and Statistical Bias

A question worth collecting data for

A statistical question anticipates variability: the answers are expected to differ from person to person or day to day. "How tall is Marcus?" is not statistical; it has one answer. "How tall are the eighth graders on the team?" is statistical.

For a boxplot you need one more thing: the answers must be numeric. "What is your favorite sport?" produces categories, which a boxplot cannot display. "How many hours a week do you practice a sport?" produces numbers, which it can.

A well-formulated question says four things:

Compare a vague question with a sharpened one:

Vague Sharpened
Do students use their phones a lot? How many minutes of screen time did each of 20 randomly chosen eighth graders at Lee Middle School log yesterday?
Is the bus slow? How many minutes did Route 7 take on each of its last 15 morning runs?

Deciding what data you need, and getting it

Once the question is sharp, the data it needs is usually obvious: one number per member of the group, in the units named. Then choose a method:

Each method has a trade-off. Surveys depend on honest answers. Measurement is accurate but slow. Acquired data is fast and often large, but you did not control how it was gathered, so you must ask who gathered it and how.

Because this chapter caps data sets at 20 items, most collection plans here take one sample of at most 20 members from a larger population.

Statistical bias

The population is the whole group you want to describe. A sample is the part of it you actually collect data from. A sample is useful only when it is representative — when its values look like the population's values would.

Statistical bias is any feature of how the data was collected that pushes the sample away from the population in a predictable direction. The word does not mean the researcher was unfair on purpose. A biased plan produces a boxplot whose center, spread, or both are shifted, and no amount of careful arithmetic afterward can repair it.

Four kinds of bias are worth naming:

The fix for most bias is random selection: every member of the population has an equal chance of being chosen, so no group is systematically favored. Random selection does not guarantee a perfect sample — a random sample can still be unlucky — but it removes the systematic push.

Bias is about the method, not the numbers. You cannot look at a data set and see bias in it. You detect bias by asking how the data was collected, and who could not possibly have ended up in it.

Worked examples

Example 1 — Sharpening a question

Sharpen "Do eighth graders exercise?" into a question a boxplot can answer.

Name who, what with units, and when.

Answer: "How many minutes did each of 20 randomly chosen eighth graders at our school exercise yesterday?"

Example 2 — Naming the population and the sample

A principal wants to describe study time for all 340340 eighth graders and surveys 2020 of them.

Answer: Population: all 340340 eighth graders. Sample: the 2020 surveyed.

Example 3 — Identifying bias

To find typical daily reading time for the school, a student surveys 20 people leaving the library. Is the sample representative?

People leaving a library read more than the school average, and students who never enter the library have no chance of being chosen. The center of the boxplot would sit too high.

Answer: No — selection bias. The sample overrepresents heavy readers, so both the median and the quartiles would be pulled upward.

Example 4 — Fixing a biased plan

Repair the plan in Example 3.

Give every student a chance of selection: number the whole roster and use a random number generator to pick 20 students, then survey those 20 wherever they are.

Answer: Choose the 20 students at random from the full roster instead of from library traffic.

Example 5 — Response bias in wording

A survey asks, "How many minutes did you waste on your phone yesterday?" Name the problem and rewrite the question.

The word waste judges the behavior, so students will underreport.

Answer: Response bias. Rewrite as "How many minutes of screen time did your phone report yesterday?"

Guided practice

  1. Write the definition of population and sample in your own words.
  2. State whether each is statistical and numeric: a) "How many pets does each student have?" b) "What is the school mascot?" c) "How many minutes did each bus run take?"
  3. Name the population and the sample: a coach measures the vertical jump of 15 of the 60 students who tried out.
  4. Name the bias: a survey about school lunch is given only to students who buy school lunch.
  5. Name the bias: "Wouldn't you agree that our library needs more computers?"
  6. Explain how random selection reduces bias, in one sentence.

Independent practice

  1. Sharpen each question so a boxplot could answer it. a) Do students sleep enough? b) Is the walk to school long? c) Do eighth graders read?
  2. For each question in item 26, name the exact quantity you would record and its units.
  3. For each question in item 26, name the collection method you would use and why.
  4. Name the bias in each plan and say which direction it pushes the results. a) Asking about exercise habits only at a gym. b) Emailing a survey about internet speed. c) Asking students their grades out loud in class. d) Letting anyone who wants to fill out a form about school spirit.
  5. A student says a sample of 20 is biased just because 20 is small. Correct the statement.
  6. A student wants to describe typical bus-ride times for all riders and records the times for the 12 riders on her own bus. Is her sample representative? Explain.
  7. Describe a way to choose 20 students at random from a roster of 340.
  8. Give an example of acquiring existing data for a question about temperatures, and name one thing you should check about the source.
  9. Application. You want to know how many minutes eighth graders spend on homework on a school night. Write the sharpened question, name the population, describe a random sampling plan for at most 20 students, and name one bias your plan avoids.
  10. Error analysis. A student collects reading times from 20 friends and writes, "Since I asked 20 people, my sample represents the school." Explain what is wrong.

Exit ticket 16.2

  1. Define statistical bias in one sentence.
  2. Name the population and the sample: 18 of the 200 band members are timed on a scale exercise.
  3. Name the bias and suggest a fix: a survey about how far students travel to school is handed out in the car-rider line.
  4. Explain why bias cannot be detected by looking only at the list of numbers collected.

Lesson 16.3 — The Five-Number Summary

Five numbers, in order

A boxplot is drawn from exactly five statistics, and every one of them is a position in the ordered data. Order the data first, every single time. Almost every mistake in this chapter comes from computing a quartile on an unsorted list.

The five statistics, using the standard's own names:

Statistic What it is
lower extreme (minimum) the smallest value
lower quartile (Q1Q_1) the median of the lower half
median the middle value of the whole set
upper quartile (Q3Q_3) the median of the upper half
upper extreme (maximum) the largest value

Two more statistics are computed from those five and describe spread:

range=upper extremelower extreme\text{range} = \text{upper extreme} - \text{lower extreme}

interquartile range (IQR)=Q3Q1\text{interquartile range (IQR)} = Q_3 - Q_1

The range measures the whole spread, edge to edge. The IQR measures the spread of the middle half of the data, which is the part a single unusual value cannot easily disturb.

Finding the median

The median is the middle value once the data is in order. With an odd count nn, it is the single value in position n+12\frac{n+1}{2}. With an even count, it is the mean of the two middle values, and it may not be a member of the data set at all.

Finding the quartiles — the rule this book uses

Split at the median. When the count is odd, the median belongs to neither half. Then Q1Q_1 is the median of the lower half and Q3Q_3 is the median of the upper half.

Here is that rule carried out on Nadia's 15 reading times.

Fifteen ordered values with the median highlighted and brackets showing the lower and upper halves

The count is 1515, which is odd, so the median is the 88th value, 3030. Removing it leaves seven values on each side. The lower half is 12,15,18,20,22,25,2812, 15, 18, 20, 22, 25, 28, whose median is 2020, so Q1=20Q_1 = 20. The upper half is 32,35,38,40,45,48,5532, 35, 38, 40, 45, 48, 55, whose median is 4040, so Q3=40Q_3 = 40. The five-number summary is

12,20,30,40,5512,\quad 20,\quad 30,\quad 40,\quad 55

range=5512=43IQR=4020=20\text{range} = 55 - 12 = 43 \qquad \text{IQR} = 40 - 20 = 20

With an even count the split is cleaner: the two halves are simply the bottom n2\frac{n}{2} values and the top n2\frac{n}{2} values, and nothing is left out.

What each quartile means

Roughly one quarter of the data lies below Q1Q_1, one quarter between Q1Q_1 and the median, one quarter between the median and Q3Q_3, and one quarter above Q3Q_3. So half the data lies between Q1Q_1 and Q3Q_3 — that is the middle half the IQR measures. Say "about," not "exactly": with 15 values you cannot split into four groups of equal size.

Worked examples

Example 1 — Odd count

Find the five-number summary, range, and IQR of 21,24,25,27,30,33,35,38,4021, 24, 25, 27, 30, 33, 35, 38, 40.

The nine values are already in order, so the median is the 55th, 3030. The lower half is 21,24,25,2721, 24, 25, 27, whose median is 24+252=24.5\frac{24+25}{2} = 24.5. The upper half is 33,35,38,4033, 35, 38, 40, whose median is 35+382=36.5\frac{35+38}{2} = 36.5.

Answer: 21, 24.5, 30, 36.5, 4021,\ 24.5,\ 30,\ 36.5,\ 40; range =19= 19; IQR =12= 12

Example 2 — Even count

Find the five-number summary of 6,8,10,12,14,15,18,20,22,26,30,366, 8, 10, 12, 14, 15, 18, 20, 22, 26, 30, 36 (minutes riding the bus).

Twelve values, so the median is the mean of the 66th and 77th: 15+182=16.5\frac{15+18}{2} = 16.5. The lower half is the first six values, 6,8,10,12,14,156, 8, 10, 12, 14, 15, whose median is 10+122=11\frac{10+12}{2} = 11. The upper half is 18,20,22,26,30,3618, 20, 22, 26, 30, 36, whose median is 22+262=24\frac{22+26}{2} = 24.

Answer: 6, 11, 16.5, 24, 366,\ 11,\ 16.5,\ 24,\ 36; range =30= 30; IQR =13= 13

Example 3 — Data given out of order

Find the median and quartiles of 19,11,28,14,23,16,13,2119, 11, 28, 14, 23, 16, 13, 21.

Sort first: 11,13,14,16,19,21,23,2811, 13, 14, 16, 19, 21, 23, 28. Eight values, so the median is 16+192=17.5\frac{16+19}{2} = 17.5. Lower half 11,13,14,1611, 13, 14, 16 gives Q1=13+142=13.5Q_1 = \frac{13+14}{2} = 13.5. Upper half 19,21,23,2819, 21, 23, 28 gives Q3=21+232=22Q_3 = \frac{21+23}{2} = 22.

Answer: 11, 13.5, 17.5, 22, 2811,\ 13.5,\ 17.5,\ 22,\ 28; range =17= 17; IQR =8.5= 8.5

Example 4 — Range and IQR compared

For the data in Example 3, explain what the range and the IQR each tell you.

The range, 1717, says the whole data set spans 17 units from smallest to largest. The IQR, 8.58.5, says the middle half of the values is packed into a span of 8.58.5 units — half the width of the whole spread.

Answer: Range describes the total spread; IQR describes the spread of the middle half

Example 5 — Working backward

A data set of 8 values has Q1=20Q_1 = 20, median =26= 26, and Q3=35Q_3 = 35. What is the IQR, and what fraction of the data lies between 2020 and 3535?

IQR=3520=15\text{IQR} = 35 - 20 = 15

Answer: IQR =15= 15; about half the data lies between Q1Q_1 and Q3Q_3

Guided practice

  1. Put in order and find the median: 9,4,12,7,159, 4, 12, 7, 15.
  2. For the set in item 40, find Q1Q_1 and Q3Q_3.
  3. For the set in item 40, find the range and the IQR.
  4. Find the five-number summary of 2,4,6,8,10,122, 4, 6, 8, 10, 12.
  5. Find the range and the IQR for item 43.
  6. Find the five-number summary of 3,5,8,10,12,15,183, 5, 8, 10, 12, 15, 18.
  7. In item 45, explain why the value 1010 is in neither half when you find the quartiles.
  8. State the quartile rule used in this book in your own words.

Independent practice

  1. Find the five-number summary, range, and IQR of 45,50,52,55,58,60,62,65,70,7545, 50, 52, 55, 58, 60, 62, 65, 70, 75.
  2. Find the five-number summary, range, and IQR of 11,13,14,16,19,21,23,2811, 13, 14, 16, 19, 21, 23, 28.
  3. Find the five-number summary, range, and IQR of 58,60,63,65,66,68,70,72,74,77,80,8458, 60, 63, 65, 66, 68, 70, 72, 74, 77, 80, 84.
  4. Find the five-number summary of 2,3,3,5,7,8,9,9,10,142, 3, 3, 5, 7, 8, 9, 9, 10, 14, and explain how repeated values are handled.
  5. Find the five-number summary of 55,57,60,60,63,64,66,7055, 57, 60, 60, 63, 64, 66, 70.
  6. Find the median of 1,3,4,4,6,8,91, 3, 4, 4, 6, 8, 9 and of 10,12,13,15,16,18,19,20,2210, 12, 13, 15, 16, 18, 19, 20, 22.
  7. A data set has range 3030 and lower extreme 1212. Find the upper extreme.
  8. A data set has Q1=18Q_1 = 18 and IQR =11= 11. Find Q3Q_3.
  9. Explain why the IQR can never be larger than the range.
  10. Give two different data sets of 5 values each with the same median but different IQRs.
  11. Application. A trainer records seconds for a 40-meter dash: 6.2,5.8,7.1,6.5,6.0,6.9,7.4,6.3,5.96.2, 5.8, 7.1, 6.5, 6.0, 6.9, 7.4, 6.3, 5.9. Find the five-number summary, the range, and the IQR.
  12. Error analysis. A student finds the quartiles of 8,3,15,6,118, 3, 15, 6, 11 by taking Q1=3Q_1 = 3 and Q3=11Q_3 = 11 from the list as written. Identify the error and give the correct summary.

Exit ticket 16.3

  1. Find the five-number summary of 6,9,11,14,206, 9, 11, 14, 20.
  2. Find the range and the IQR for item 60.
  3. Find the five-number summary of 22,25,28,30,31,3422, 25, 28, 30, 31, 34.
  4. State in one sentence what the IQR measures that the range does not.

Lesson 16.4 — Building and Reading a Boxplot

From five numbers to a picture

A boxplot (also called a box-and-whisker plot) draws the five-number summary above a number line, so every horizontal position on the picture is a real value on that scale.

To build one:

A labeled boxplot of fifteen reading times showing the five-number summary, range, and interquartile range

Every feature of this picture is a length you can read off the axis:

Reading the density of the data

Each of the four sections — lower whisker, left part of the box, right part of the box, upper whisker — holds about a quarter of the values. So a long section means about a quarter of the data is spread thinly across a wide interval, and a short section means about a quarter of the data is crowded into a narrow interval.

In the figure above, the lower whisker runs from 1212 to 2020, a span of 88, while the upper whisker runs from 4040 to 5555, a span of 1515. Each holds about four students. So the four shortest readers are packed tightly and the four longest readers are spread widely — the distribution is stretched toward the high end.

What the picture leaves out

The same fifteen values as a dot plot with the mean marked, and as a boxplot

The dot plot and the boxplot above show the same 15 numbers. The dot plot shows every value and lets you mark the mean at about 30.930.9. The boxplot shows neither. You cannot read from a boxplot:

That is why a boxplot is always labeled with what it displays and how many values it summarizes — that count must be written in words, because the picture cannot carry it.

An even count

A boxplot of twelve bus-ride times with the five-number summary labeled

This boxplot shows the 12 bus-ride times from Lesson 16.3. The median line sits at 16.516.5, which is not one of the recorded times — with an even count the median is the mean of the two middle values, and a boxplot happily draws a line at a value nobody recorded.

Making observations and drawing conclusions

An observation is something you read directly off the plot. A conclusion is what that means for the question you asked. Analysis needs both, in that order.

Observation Conclusion
Median =16.5= 16.5 A typical bus ride takes about 161216\tfrac{1}{2} minutes.
Q3=24Q_3 = 24 About a quarter of rides take longer than 2424 minutes.
IQR =13= 13, range =30= 30 The middle half is far more consistent than the full set; a few long rides stretch the top.
Upper whisker (2424 to 3636) is longer than the lower whisker (66 to 1111) The long rides vary much more than the short ones, so a rider planning around the bus should budget for the top of the range.

Worked examples

Example 1 — Building a boxplot

Describe the boxplot for 4,7,9,12,154, 7, 9, 12, 15.

Five values, so the median is 99. The lower half is 4,74, 7 giving Q1=5.5Q_1 = 5.5; the upper half is 12,1512, 15 giving Q3=13.5Q_3 = 13.5.

Answer: Box from 5.55.5 to 13.513.5, median line at 99, whiskers to 44 and 1515; range =11= 11, IQR =8= 8

Example 2 — Reading statistics off a plot

A boxplot has its left whisker end at 33, box edges at 77 and 1717, median line at 1212, and right whisker end at 2424. List all seven statistics.

Answer: Lower extreme 33, Q1=7Q_1 = 7, median 1212, Q3=17Q_3 = 17, upper extreme 2424, range =243=21= 24 - 3 = 21, IQR =177=10= 17 - 7 = 10

Example 3 — Interpreting a section

In Example 2, what fraction of the days had more than 1717 books checked out, and over what span?

Above Q3Q_3 lies about a quarter of the data, spread from 1717 to 2424.

Answer: About one quarter of the days, spread across a span of 77 books

Example 4 — An observation and a conclusion

For the bus data, write one observation and the conclusion that follows.

Answer: Observation: the box runs from 1111 to 2424. Conclusion: about half of all rides take between 1111 and 2424 minutes, so a rider who leaves 25 minutes before the bell will usually arrive on time.

Example 5 — A claim the plot cannot support

A student looks at the bus boxplot and says, "Two students rode for exactly 16.516.5 minutes." Evaluate the claim.

The median of an even-count set is the mean of the two middle values, and no boxplot reports individual values at all.

Answer: Unsupported. The median 16.516.5 is a computed midpoint; the two middle recorded times were 1515 and 1818, and neither the count at any value nor any individual value is visible in a boxplot.

Guided practice

Use the labeled boxplot figures from this lesson where a figure is named.

  1. In the labeled reading-times boxplot, name all five plotted statistics.
  2. In that same plot, find the range and the IQR.
  3. In that same plot, which whisker is longer, and what does that tell you?
  4. In the bus-ride boxplot, name all five plotted statistics.
  5. In the bus-ride boxplot, about what fraction of rides took longer than 2424 minutes?
  6. Describe the boxplot you would draw for 3,5,8,10,12,15,183, 5, 8, 10, 12, 15, 18: give the box edges, the median line, and the whisker ends.
  7. Describe the boxplot you would draw for 2,4,6,8,10,122, 4, 6, 8, 10, 12.

Independent practice

  1. Draw a boxplot for 21,24,25,27,30,33,35,38,4021, 24, 25, 27, 30, 33, 35, 38, 40 on a scale from 2020 to 4040 with ticks every 22.
  2. Draw a boxplot for 45,50,52,55,58,60,62,65,70,7545, 50, 52, 55, 58, 60, 62, 65, 70, 75 on a scale from 4040 to 8080 with ticks every 55.
  3. Draw a boxplot for 58,60,63,65,66,68,70,72,74,77,80,8458, 60, 63, 65, 66, 68, 70, 72, 74, 77, 80, 84 on a scale from 5555 to 8585 with ticks every 55.
  4. A boxplot has whisker ends at 1010 and 4646 and box edges at 1818 and 3434, with the median line at 2222. List all seven statistics.
  5. In item 74, which is longer, the part of the box below the median or above it? What does that say about the data?
  6. In item 74, about what fraction of the data lies between 1818 and 3434?
  7. Two data sets have the same range but different IQRs. Sketch what that difference looks like in two boxplots.
  8. Explain why a boxplot alone cannot tell you how many values are in the data set.
  9. A boxplot is drawn with the median line exactly in the middle of the box. What does that say about the data between Q1Q_1 and Q3Q_3?
  10. Application. A café records the number of customers in the first hour on 8 days: 19,11,28,14,23,16,13,2119, 11, 28, 14, 23, 16, 13, 21. Find the five-number summary, describe the boxplot, then write one observation and one conclusion for the manager.
  11. Application. Daily high temperatures for 12 days are 58,60,63,65,66,68,70,72,74,77,80,8458, 60, 63, 65, 66, 68, 70, 72, 74, 77, 80, 84 degrees Fahrenheit. Write two observations and one conclusion a gardener could use.
  12. Error analysis. A student says, "The line inside the box is always exactly halfway between the box edges, because it is the middle." Explain the error using a data set of your own.

Exit ticket 16.4

  1. Give the box edges, the median line, and the whisker ends for a boxplot of 10,12,13,15,16,18,19,20,2210, 12, 13, 15, 16, 18, 19, 20, 22.
  2. A boxplot has whisker ends at 55 and 2929 and box edges at 1212 and 2020. Find the range and the IQR.
  3. About what fraction of the data lies between the two box edges?
  4. Name two statistics you cannot read from a boxplot.

Lesson 16.5 — Extreme Data Points

What an extreme data point is

An extreme data point, or outlier, is a value far away from the rest of the data. A basketball team that usually scores in the twenties and thirties scores 7676 in one wild game; a class of readers who all read 10 to 40 minutes has one student who read 200200.

Outliers are not mistakes to delete. Sometimes an outlier is a recording error, and sometimes it is the most interesting value in the set. Either way, you must notice it and say what it did to your summary — which is exactly what bullet (f) of this standard asks.

How an outlier changes the five statistics

Look at 11 games where a team scored 20,22,24,25,26,28,30,31,32,34,3620, 22, 24, 25, 26, 28, 30, 31, 32, 34, 36, then a 1212th game where it scored 7676.

Three boxplots of the same team, before the outlier, after it, and with the outlier plotted separately

Statistic 11 games With the 7676-point game Effect
lower extreme 2020 2020 unchanged
Q1Q_1 2424 24.524.5 barely moved
median 2828 2929 barely moved
Q3Q_3 3232 3333 barely moved
upper extreme 3636 7676 jumped by 4040
range 1616 5656 more than tripled
IQR 88 8.58.5 barely moved

Two lessons live in that table.

Shape. A single extreme value stretches the picture toward itself. The box barely moves, but the whisker on that side becomes very long, so the plot looks lopsided, or skewed, toward the outlier. The three-quarters of the plot occupied by that whisker holds a single game.

Spread. The range is extremely sensitive to an outlier, because it is computed from the two most extreme values and an outlier is an extreme value. The IQR is barely affected, because it is computed from Q1Q_1 and Q3Q_3, which sit deep inside the ordered data where one distant value cannot reach. This is why the IQR is called a resistant measure of spread and why statisticians reach for it when a data set has an outlier.

The median is also resistant, for the same reason: adding one enormous value shifts the middle position by half a step, not to the outlier. The mean is not resistant at all — the mean of the 11 games is about 28.028.0, and adding the 7676 raises it to about 32.032.0. A boxplot does not show the mean, but this is worth knowing when you decide which measure of center to report.

Drawing an outlier separately

The third panel of the figure shows the common alternative: plot the outlier as its own point and stop the whisker at the largest ordinary value, 3636. The box and median are computed from the other eleven games. This version tells the reader two things at once — the shape of the ordinary data, and the existence of one far-away value — so it is usually the more honest picture. Both versions are correct as long as you say which you drew.

Deciding what to do about it

Ask, in this order:

Worked examples

Example 1 — Effect on range and IQR

For 5,6,7,8,9,10,11,12,135, 6, 7, 8, 9, 10, 11, 12, 13, find the range and IQR. Then add the value 4040 and recompute.

Nine values: median 99; lower half 5,6,7,85, 6, 7, 8 gives Q1=6.5Q_1 = 6.5; upper half 10,11,12,1310, 11, 12, 13 gives Q3=11.5Q_3 = 11.5. So range =135=8= 13 - 5 = 8 and IQR =11.56.5=5= 11.5 - 6.5 = 5.

With 4040 added there are ten values: median =9+102=9.5= \frac{9+10}{2} = 9.5; lower half 5,6,7,8,95, 6, 7, 8, 9 gives Q1=7Q_1 = 7; upper half 10,11,12,13,4010, 11, 12, 13, 40 gives Q3=12Q_3 = 12. So range =405=35= 40 - 5 = 35 and IQR =127=5= 12 - 7 = 5.

Answer: Range jumps from 88 to 3535; IQR stays at 55

Example 2 — Describing the shape change

Describe how the boxplot in Example 1 changes.

The box shifts a little to the right and keeps almost exactly its width, and the median line moves from 99 to 9.59.5. The upper whisker stretches from ending at 1313 to ending at 4040, so most of the width of the picture is now one whisker holding a single value.

Answer: The box is nearly unchanged; the upper whisker becomes very long and the plot looks strongly stretched to the right

Example 3 — Which measure to report

A reporter asks for one number describing a typical game for the 12-game season including the 7676. Median or mean?

The mean, about 32.032.0, is higher than 99 of the 1212 games — it is not typical of anything. The median, 2929, sits in the middle of the ordinary games.

Answer: The median, because it is resistant to the one extreme game

Example 4 — An outlier that is an error

A data set of daily temperatures in degrees Fahrenheit reads 58,60,63,650,6658, 60, 63, 650, 66. What should you do?

650F650^\circ\text{F} is impossible; it is almost certainly 6565 typed with an extra zero.

Answer: Treat it as a recording error — correct it to 6565 if the original record confirms it, or drop it and note the removal. Do not silently keep it.

Example 5 — Outlier at the low end

A set of 10 quiz scores is 71,74,75,77,78,80,82,84,86,1271, 74, 75, 77, 78, 80, 82, 84, 86, 12. Describe the effect of the 1212.

Sorted: 12,71,74,75,77,78,80,82,84,8612, 71, 74, 75, 77, 78, 80, 82, 84, 86. The lower whisker now stretches from 7171 all the way down to 1212, so the plot is stretched to the left, the range is 8612=7486 - 12 = 74, and the lower extreme no longer resembles any other score. The box, running from Q1=74Q_1 = 74 to Q3=82Q_3 = 82 with median 77+782=77.5\frac{77+78}{2} = 77.5, still describes the ordinary scores well.

Answer: The low outlier stretches the plot leftward and inflates the range to 7474, while the box and median still describe the other nine scores

Guided practice

  1. Define extreme data point (outlier) in your own words.
  2. In the three-panel figure of this lesson, what is the upper extreme in the first panel, and in the second?
  3. In that figure, compare the IQR before and after the 7676-point game, and explain why it barely changed.
  4. In that figure, compare the range before and after, and explain why it changed so much.
  5. In the third panel, why does the whisker stop at 3636?
  6. Which is more resistant to an outlier, the range or the IQR? Why?

Independent practice

  1. For 5,6,7,8,9,10,11,12,135, 6, 7, 8, 9, 10, 11, 12, 13, find the range and IQR. Then add 4040 and find them again.
  2. Describe in words how the boxplot in item 93 changes shape.
  3. For 71,74,75,77,78,80,82,84,8671, 74, 75, 77, 78, 80, 82, 84, 86, find the median and IQR. Then add a score of 1212 and find them again.
  4. In item 95, say which of the seven statistics changed the most and which did not change at all.
  5. A data set has an outlier at the high end. Which will be larger, the upper whisker or the lower whisker? Explain.
  6. Explain why the median moves so little when one very large value is added to a data set.
  7. Compute the mean of 20,22,24,25,26,28,30,31,32,34,3620, 22, 24, 25, 26, 28, 30, 31, 32, 34, 36 and of that set plus 7676. Compare the change in the mean to the change in the median, which is from 2828 to 2929.
  8. Application. A pizza shop's delivery times in minutes over 10 orders are 18,20,21,22,24,25,26,28,30,6218, 20, 21, 22, 24, 25, 26, 28, 30, 62. Find the five-number summary. Then say what happened on the 6262-minute delivery, what the range says, what the IQR says, and which one you would quote to a customer.
  9. Error analysis. A student sees an outlier and deletes it, saying, "It ruins the graph." Explain what is wrong with that reasoning and what to do instead.

Exit ticket 16.5

  1. Adding one very large value to a data set: does the range change a lot or a little? Does the IQR? Explain each.
  2. Describe in one sentence how a high outlier changes the shape of a boxplot.
  3. Give one situation in which you would remove an outlier and one in which you would keep it.

Lesson 16.6 — Comparing Boxplots, Choosing a Display, and Spotting a Misleading One

Two boxplots, one axis

The reason boxplots earn their place is comparison. Stack two of them on the same scale and differences in center and spread are visible instantly.

Boxplots of spelling test scores for Class A and Class B on the same scale

Class A Class B
lower extreme 5555 6868
Q1Q_1 6565 7373
median 7575 7878
Q3Q_3 8585 8282
upper extreme 9595 8888
range 4040 2020
IQR 2020 99

Compare in three passes, always in this order:

A useful sentence pattern for the conclusion: Class B scored higher on average and much more consistently, but the highest individual scores were in Class A.

Equal medians do not mean equal data

Boxplots of delivery times on two routes with the same median and very different spread

Both routes have a median of 5050 minutes. If you compared only centers you would call them identical. But Route 1 has range 8080 and IQR 4040, while Route 2 has range 2020 and IQR 88. A driver who must arrive on time should choose Route 2 every day: half its trips fall between 4646 and 5454 minutes, while half of Route 1's fall anywhere between 3030 and 7070. Always compare spread, not just center.

Choosing the right graphical representation

Given a situation, which display should you make? Match the display to the question.

The same twenty sit-up counts shown as a dot plot, a histogram, and a boxplot

Display Shows Best when the question is
dot plot every individual value, and repeats "What exact values occurred, and which are most common?" — small data sets
histogram counts within intervals; the shape of the distribution "What is the shape? How many fall between 20 and 30?"
boxplot the five-number summary and spread "What is typical, how spread out is it, and how do two groups compare?"
circle graph parts of a whole "What share of the total is each category?" — categorical data
line graph change over time "How did this quantity change from month to month?"

Some rules that decide most cases:

When you justify a choice, name the feature of the question that forces it: "The question asks which of two teams is more consistent, and consistency is spread, so I chose two boxplots on one axis."

Misleading components of a graphical display

A graph can be technically accurate and still leave a false impression. The features to check:

The same fifteen scores drawn on a 50-to-100 axis and on a 0-to-150 axis

How to check a display in ten seconds. Read the axis and its scale. Read the units. Read how many values are summarized and how they were collected. Only then look at the boxes.

Worked examples

Example 1 — Comparing centers and spreads

Compare P=12,14,15,17,18,20,22,24,26,30P = 12, 14, 15, 17, 18, 20, 22, 24, 26, 30 with Q=16,17,18,19,20,21,22,23,24,25Q = 16, 17, 18, 19, 20, 21, 22, 23, 24, 25.

PP: median 18+202=19\frac{18+20}{2} = 19, Q1=15Q_1 = 15, Q3=24Q_3 = 24, range 1818, IQR 99. QQ: median 20+212=20.5\frac{20+21}{2} = 20.5, Q1=18Q_1 = 18, Q3=23Q_3 = 23, range 99, IQR 55.

Answer: QQ has the higher median (20.520.5 vs 1919) and about half the spread by either measure, so QQ is both higher and more consistent; PP contains both the smallest value (1212) and the largest (3030).

Example 2 — A conclusion that does not follow

From the Class A / Class B plots, a student concludes, "Class B has more students who scored above 8080." Evaluate.

Boxplots do not show counts. Both classes have 15 students, and about a quarter of each class scored above its own Q3Q_3, but the plot cannot tell you how many exceeded a particular score like 8080.

Answer: Not supported — a boxplot shows positions, not counts

Example 3 — Justifying a representation

A student wants to show how her town's monthly rainfall changed across last year. Which display, and why?

The question is about change across ordered months, and a boxplot discards order entirely.

Answer: A line graph, because the question is about change over time and a boxplot would lose the month-by-month order

Example 4 — Justifying a representation

A coach wants to know whether the varsity or junior varsity team is more consistent in points scored. Which display, and why?

"More consistent" is a question about spread, comparing two groups.

Answer: Two boxplots on one shared axis, because the question compares two groups on spread, which the box width and whisker lengths show directly

Example 5 — Finding what is misleading

A club's poster shows one boxplot labeled "Our members read a lot!" with a box from 3030 to 9090, no axis numbers, and no note of how many members were surveyed. Name three problems.

Answer: (1) No labeled axis or scale, so the numbers cannot be verified or compared; (2) no units — minutes? pages? per day or per week?; (3) no sample size or sampling method, and members who chose to respond are a voluntary-response sample likely to read more than average.

Guided practice

Use the figures in this lesson where a figure is named.

  1. From the two-class figure, list the median and IQR of each class.
  2. From that figure, which class is more consistent, and how do you know?
  3. From that figure, which class contains the single highest score?
  4. From the two-routes figure, state the medians and explain why the routes are not equivalent.
  5. From the two-routes figure, which route would you choose to guarantee arriving within an hour? Justify with two statistics.
  6. From the three-display figure, name one thing the dot plot shows that the boxplot does not.
  7. From the two-scale figure, explain how the same data can look so different.

Independent practice

  1. Compare P=12,14,15,17,18,20,22,24,26,30P = 12, 14, 15, 17, 18, 20, 22, 24, 26, 30 and Q=16,17,18,19,20,21,22,23,24,25Q = 16, 17, 18, 19, 20, 21, 22, 23, 24, 25: give both five-number summaries, then compare center, spread, and overlap in three sentences.
  2. Two boxplots have the same median but very different IQRs. Write one sentence describing what that means in context, choosing your own context.
  3. Two boxplots have the same range but different medians. What can you conclude, and what can you not?
  4. Justify the best display for each situation, naming the feature of the question that decides it. a) The share of eighth graders choosing each of four electives. b) A city's population each year from 2015 to 2024. c) Whether morning or afternoon bus routes take longer, and which is more reliable. d) Which of 1818 recorded quiz scores occurred most often.
  5. A newspaper prints two boxplots of house prices side by side, one on a 00300300 axis and one on a 150150250250 axis. Explain why the comparison is invalid.
  6. Name three components you would check on any boxplot before believing a claim made from it.
  7. A boxplot's box is very wide and its whiskers are very short. Describe the data.
  8. Explain why "the wider section of the box contains more data values" is a misreading.
  9. A display shows a boxplot with no sample size given. Give two different data sets, one of 6 values and one of 20, and explain why the omission matters.
  10. Application. Team X scored 4,6,7,9,10,12,13,15,16,18,20,264, 6, 7, 9, 10, 12, 13, 15, 16, 18, 20, 26 points per game and Team Y scored 8,9,10,11,12,13,14,15,16,17,19,228, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 22. Give both five-number summaries and decide which team a coach should call more reliable, justifying with the IQR and the range.
  11. Application. You must present to the school board whether eighth graders or seventh graders spend more time on homework, and how much the two groups vary. Say which display you would use, why, and what you would label on it.
  12. Error analysis. A student compares two boxplots and writes, "Group 1's box is wider, so Group 1 has more students." Explain the error and write a correct statement.

Exit ticket 16.6

  1. Two boxplots share an axis. Group 1 has median 4040 and IQR 66; Group 2 has median 4040 and IQR 2222. Write one sentence comparing them.
  2. Name the display you would choose to compare the spread of two groups, and why.
  3. Name two components of a graphical display that can mislead a reader.
  4. Explain why two boxplots must share one axis to be compared.

Chapter 16 Review

Vocabulary. data cycle · data · statistical question · population · sample · representative · statistical bias · selection bias · response bias · nonresponse bias · voluntary response bias · random selection · five-number summary · lower extreme (minimum) · upper extreme (maximum) · median · lower quartile (Q1Q_1) · upper quartile (Q3Q_3) · range · interquartile range (IQR) · boxplot · box · whisker · extreme data point (outlier) · resistant · skewed

Part A — The data cycle, collecting data, and bias (8.PS.2a, b, c)

  1. Name the four stages of the data cycle in order, and say what happens in each in one phrase.
  2. Sharpen this into a question a boxplot can answer, then name the population and the data you would record: "Do eighth graders carry heavy backpacks?"
  3. For item 129, describe a collection method for at most 20 students and say why you chose it.
  4. Name the bias in each plan and the direction it pushes results. a) Measuring backpack weight only on Fridays, when lockers are cleaned out. b) Asking for volunteers to have their backpacks weighed. c) Weighing only backpacks in the athletics hallway.
  5. Explain what "representative sample" means and why bias prevents it, in two sentences.

Part B — Organizing, representing, and describing (8.PS.2d, e)

  1. Find the five-number summary, range, and IQR of 1,2,2,3,4,4,5,5,6,6,7,7,8,8,9,10,11,12,14,181, 2, 2, 3, 4, 4, 5, 5, 6, 6, 7, 7, 8, 8, 9, 10, 11, 12, 14, 18.
  2. Find the five-number summary, range, and IQR of 100,102,105,108,110,115,120,125100, 102, 105, 108, 110, 115, 120, 125.
  3. State the quartile rule this book uses, and apply it to 3,5,8,10,12,15,183, 5, 8, 10, 12, 15, 18.
  4. Use the blank axes below. On the middle axis, draw a boxplot for 11,13,14,16,19,21,23,2811, 13, 14, 16, 19, 21, 23, 28.

Three blank number-line axes for drawing boxplots by hand

  1. On the top axis of that figure, draw a boxplot for 6,8,10,12,14,15,18,20,22,26,30,366, 8, 10, 12, 14, 15, 18, 20, 22, 26, 30, 36.
  2. Use the boxplot below. Identify the lower extreme, Q1Q_1, median, Q3Q_3, and upper extreme.

A boxplot of books checked out per day over 13 school days, with no labels on the statistics

  1. From that same boxplot, find the range and the IQR, and say what each measures.

Part C — Outliers, analysis, and comparison (8.PS.2f, g, h)

  1. From the books-checked-out boxplot in item 138, write two observations and one conclusion a librarian could act on.
  2. From that same boxplot, explain what the length of the upper whisker tells you compared with the lower whisker.
  3. A data set of 99 values has range 88 and IQR 55. One value of 4040 is added, making the range 3535 while the IQR stays 55. Explain, in terms of how each is computed, why one changed and the other did not.
  4. Describe in two sentences how a single very large value changes the shape of a boxplot.
  5. Use the two boxplots below. Give the five-number summary of each team from the plot, then compare the two teams on center and on spread.

Boxplots of points per game for Team X and Team Y on one shared axis

  1. From that same figure, which team would a coach call more reliable, and which team produced the single best game? Justify both answers with statistics.

Part D — Choosing a display and spotting a misleading one (8.PS.2i, j)

  1. Justify the best display for each. a) Comparing the spread of scores in two classes. b) Showing what fraction of students chose each of three field trips. c) Showing every one of 15 recorded times, including repeats. d) Showing a store's daily sales across two weeks.
  2. Explain why categorical data can never be shown in a boxplot.
  3. A poster shows one boxplot with no axis labels and no sample size. Name three questions you would ask before believing it.
  4. Two boxplots of the same quantity are printed on different scales. Explain the false impression this creates and how to fix it.
  5. Reasoning. A club claims, "Our members read more than the average student," and supports it with a boxplot of reading times from 12 members who volunteered. Name the bias in the collection, name one misleading feature of presenting a boxplot alone here, and describe a collection and display plan that would actually settle the claim.

Standards coverage check — Chapter 16

Knowledge and Skill Where it is taught Where it is practiced
8.PS.2a — formulate questions that require the collection or acquisition of data with a focus on boxplots 16.1 (statistical questions a boxplot can answer, stage 1), 16.2 (who / what / when / why it varies; sharpening a vague question) Items 1, 3, 9, 10, 14, 18; 21, 26, 34, 36; Review 128, 129
8.PS.2b — determine the data needed to answer a formulated question and collect the data (or acquire existing data) using various methods 16.1 (stage 2, acquiring versus collecting), 16.2 (survey, measurement, observation, acquisition, and the trade-offs) Items 6, 7, 11, 14; 22, 27, 28, 32, 33, 34, 37; Review 129, 130
8.PS.2c — determine how statistical bias might affect whether the data collected from the sample is representative of the larger population 16.2 (population and sample; selection, response, nonresponse, and voluntary response bias; random selection as the fix) Items 20, 23, 24, 25, 29, 30, 31, 32, 35, 36, 38, 39; Review 131, 132, 150
8.PS.2d — organize and represent a numeric data set of no more than 20 items, using boxplots, with and without technology 16.3 (ordering the data), 16.4 (the five construction steps, by hand and with technology) Items 40, 43, 45; 69, 70, 71, 72, 73, 77, 80, 81, 83; Review 133, 136, 137
8.PS.2e — identify and describe the lower extreme (minimum), upper extreme (maximum), median, upper quartile, lower quartile, range, and interquartile range given a data set, represented by a boxplot 16.3 (all seven statistics and the quartile convention), 16.4 (reading each one off the drawn plot) Items 40–63; 64, 65, 67, 68, 74, 75, 76, 79, 84, 85, 86; Review 133, 134, 135, 138, 139
8.PS.2f — describe how the presence of an extreme data point (outlier) affects the shape and spread of the data distribution of a boxplot 16.5 (effect on each statistic; range versus IQR resistance; drawing the outlier separately; what to do about it) Items 87–104; Review 142, 143
8.PS.2g — analyze data represented in a boxplot by making observations and drawing conclusions 16.4 (observation versus conclusion, section-by-section reading), 16.5 (what the outlier means for the question), 16.6 (claims a boxplot cannot support) Items 12, 13, 66, 68, 75, 76, 78, 79, 80, 81, 82; 100; 119, 123; Review 140, 141
8.PS.2h — compare and analyze two data sets represented in boxplots 16.6 (compare center, then spread, then overlap; equal medians with unequal spread) Items 105–109, 112, 113, 114, 121, 124; Review 144, 145
8.PS.2i — given a contextual situation, justify which graphical representation best represents the data 16.1 (choosing the display in stage 3), 16.6 (dot plot, histogram, boxplot, circle graph, line graph, and the question each answers) Items 5, 15; 110, 115, 118, 122, 125; Review 146, 147
8.PS.2j — identify components of graphical displays that can be misleading 16.4 (what a boxplot cannot show), 16.6 (scale, mismatched axes, missing labels, missing sample size, hidden outliers, misreading section width as count) Items 12, 13, 78, 86; 111, 116, 117, 119, 120, 123, 126, 127; Review 148, 149, 150

Every data set in this chapter has at most 20 items, and every five-number summary quoted in the text, the figures, and the answer key is computed with the single quartile rule stated in the chapter opening: quartiles are the medians of the halves, with the overall median excluded from both halves when the count is odd.

Answer keys for every set in this chapter are in Appendix A.