Appendix A — Answer Key, Chapter 16: The Data Cycle and Boxplots
SOL 8.PS.2 · Covers textbook Chapter 16 and the companion workbook. Item numbers match the textbook; workbook items are the same problems, so this key serves both. Item numbers run continuously from 1 to 150 across the chapter. Reasoning answers show an acceptable response, not the only wording.
Every five-number summary in this key was computed with the chapter's quartile rule — order the data, split at the median, exclude the median from both halves when the count is odd, and take the median of each half — and every range and IQR was recomputed from the summary it follows. A calculator using a different interpolation rule may report a slightly different quartile; on this chapter's work, use the rule above.
Several items ask students to write a question, choose a sample, or state a conclusion. For those, the key gives one acceptable response and names what any acceptable response must contain.
Lesson 16.1 — The Data Cycle, Pointed at Boxplots
Guided practice
- Formulate questions; collect or acquire data; organize and represent data; analyze data and communicate results.
- Stage 3, organize and represent data. The summary is what the boxplot is drawn from.
- Stage 2, collect or acquire data. (Accept: the sampling plan is decided as part of collection.)
- The cycle is not finished; the boxplot is only stage 3. You still have to analyze it, report what you found, and usually formulate a new question from what you saw.
- Shows: the median (center) and the spread, through the box width and whisker lengths — also the extremes and both quartiles. Does not show: individual data values, the number of values, repeats, or the mean.
- Stage 2. Acquiring existing data is a collection method, not a separate stage.
Independent practice
- a) stage 1 b) stage 3 c) stage 4 d) stage 2
- Because stage 4 almost always raises a new question, which sends you back to stage 1 with a sharper version of what you wanted to know.
- Any statistical question with a numeric answer that varies, for example "How many minutes did each student in our class spend on homework last night?"
- Any question that is not statistical ("How tall is our teacher?") or not numeric ("What is your favorite sport?"). A boxplot needs a set of numbers that vary; one fixed answer or a list of categories cannot be displayed as five positions on a number line.
- One number per student, in stated units — for item 9's example, the minutes of homework for one student on one named night.
- No. A boxplot reports five positions, not counts, so it cannot say how many students hit any particular score. It cannot even confirm that occurred at all unless is one of the extremes.
- No. The mean is not one of the five numbers plotted, and it cannot be recovered from them. Two data sets with the same five-number summary can have different means.
- Question: "How many hours of sleep did each of 20 randomly chosen eighth graders get last night?" Data needed: one number of hours per student. Collection: number the eighth-grade roster, use a random number generator to pick 20, and survey those students. Any acceptable answer names a numeric quantity with units and a way of choosing at most 20 students that does not favor a subgroup.
- A boxplot compares two or more groups on one axis far better, because two boxplots stack on the same scale and their centers and spreads line up visually. A histogram shows the shape of the distribution and how many values fall in each interval, which a boxplot cannot show at all.
Exit ticket 16.1
- Formulate questions; collect or acquire data; organize and represent data; analyze data and communicate results.
- Stage 3, organize and represent data.
- Center: "What is a typical value?" Spread: "How much do the values vary?" (Accept any equivalent pair.)
- The plot may reveal something unexpected — a very long whisker, a surprisingly wide box, or two groups that differ — and explaining it needs new data, which is a new question.
Lesson 16.2 — Asking the Question, Getting the Data, and Statistical Bias
Guided practice
- The population is the whole group you want to describe. The sample is the part of the population you actually collect data from.
- a) yes — counts vary from student to student b) no — one fixed answer, and not numeric c) yes — times vary from run to run
- Population: all 60 students who tried out. Sample: the 15 whose jumps were measured.
- Selection bias. Students who bring lunch from home have no chance of being included, and they are exactly the students most likely to dislike the school lunch, so the results would look too favorable.
- Response bias. The wording tells the respondent which answer is expected, so agreement is inflated.
- Random selection gives every member of the population an equal chance of being chosen, so no subgroup is systematically over- or underrepresented.
Independent practice
- a) "How many hours of sleep did each of 20 randomly chosen eighth graders at our school get last night?" b) "How many minutes did each of 20 randomly chosen students spend walking to school this morning?" c) "How many minutes did each of 20 randomly chosen eighth graders read for pleasure yesterday?" Any acceptable rewrite names who, what with units, and when.
- a) hours of sleep, in hours b) walking time, in minutes c) reading time, in minutes
- a) survey — only the student knows b) survey, or direct measurement with a stopwatch if you can observe the walk c) survey — only the student knows. Any answer that matches the method to who has access to the number is acceptable.
- a) Selection bias — people at a gym exercise more, so the results are pushed high. b) Nonresponse bias — households with slow or no internet are least likely to answer, so reported speeds are pushed high. c) Response bias — students may not answer honestly in front of classmates, so reported grades are pushed high. d) Voluntary response bias — students with strong feelings, usually enthusiastic ones, choose to respond, so school spirit looks higher than it is.
- Small and biased are different problems. A sample of 20 chosen at random is unbiased but imprecise; a sample of 2000 chosen only from the library is precise and still biased. Bias comes from how members were chosen, not from how many.
- No. Her own bus serves one neighborhood, so its ride times reflect that route's distance and traffic and not the whole system. That is selection bias, and it could push the median either high or low depending on the route.
- Number the roster from 1 to 340, then use a random number generator (or draw numbered slips from a container) to choose 20 different numbers, and survey exactly those students.
- Download daily high temperatures for your city from the National Weather Service. Check who collected the data, when and where it was measured, and what units it is reported in.
- Question: "How many minutes did each of 20 randomly chosen eighth graders at our school spend on homework last night?" Population: all eighth graders at the school. Plan: number the eighth-grade roster, randomly select 20, and survey them privately. Bias avoided: selection bias, because students in every class and activity have an equal chance, and response bias, because the survey is private rather than read aloud.
- Friends are not a random sample of the school. They tend to share classes, schedules, and habits with the student collecting the data, so the sample is systematically like her and unlike the rest of the school. The size, 20, is fine; the selection method is the problem.
Exit ticket 16.2
- Statistical bias is any feature of how the data was collected that pushes the sample away from the population in a predictable direction.
- Population: all 200 band members. Sample: the 18 who were timed.
- Selection bias — car riders live farther away on average than walkers, so the distances collected are pushed high. Fix: select students at random from the full roster instead of from one arrival group.
- Bias lives in the collection method, not in the values. The numbers themselves look perfectly ordinary; what tells you the sample is biased is knowing who could not possibly have been included.
Lesson 16.3 — The Five-Number Summary
Guided practice
- Ordered: . Median (the third of five values).
- Lower half gives . Upper half gives . The median, , is in neither half.
- Range ; IQR .
- . With six values the median is ; lower half gives ; upper half gives .
- Range ; IQR .
- . Median is the fourth of seven values, ; lower half gives ; upper half gives .
- The count is odd, so the median is an actual data value sitting exactly in the middle. This book's rule excludes it from both halves, which leaves three values on each side and makes the two halves the same size.
- Order the data. Find the median; if the count is odd, set that value aside. is the median of the values below it and is the median of the values above it.
Independent practice
- . Ten values: median ; lower half gives ; upper half gives . Range ; IQR .
- . Eight values: median ; ; . Range ; IQR .
- . Twelve values: median ; ; . Range ; IQR .
- . Repeated values are treated as separate values and stay in the list; nothing is combined. Ten values: median ; lower half gives ; upper half gives . Range ; IQR .
- . Median ; ; .
- First set: seven values, median is the fourth, . Second set: nine values, median is the fifth, .
- .
- .
- is never below the minimum and is never above the maximum, so the interval from to always sits inside the interval from minimum to maximum. A shorter interval cannot have a greater length. They are equal only when every value is the same.
- For example (median , , , IQR ) and (median , , , IQR ). Any pair with equal medians and different IQRs is acceptable.
- Ordered: . Median is the fifth value, ; lower half gives ; upper half gives . Summary ; range s; IQR s.
- The student read positions off the unsorted list. Quartiles are positions in the ordered data. Sorted: . Median ; ; . Summary ; range ; IQR .
Exit ticket 16.3
- . Median is the third of five, ; ; .
- Range ; IQR .
- . Median ; lower half gives ; upper half gives .
- The IQR measures the spread of the middle half only, so it describes how tightly the typical values cluster and ignores the two ends, where a single unusual value can stretch the range.
Lesson 16.4 — Building and Reading a Boxplot
Guided practice
- Lower extreme , , median , , upper extreme minutes.
- Range minutes; IQR minutes.
- The upper whisker, which runs from to (a span of ), is longer than the lower whisker, which runs from to (a span of ). Each whisker holds about a quarter of the students, so the longest readers are spread across a wider range of times than the shortest readers — the distribution is stretched toward the high end.
- Lower extreme , , median , , upper extreme minutes.
- About one quarter, since is .
- Box from to , median line at , whiskers to and .
- Box from to , median line at , whiskers to and .
Independent practice
- Summary . Box from to , median line at , whiskers to and . Range ; IQR .
- Summary . Box from to , median line at , whiskers to and .
- Summary . Box from to , median line at , whiskers to and .
- Lower extreme , , median , , upper extreme , range , IQR .
- The part above the median is longer: against . Both parts hold about a quarter of the data, so the values just above the median are spread out while the values just below it are packed tightly.
- About half, since and are and .
- Any two plots whose whisker tips are the same distance apart but whose boxes differ in width — for example both running from to , one with a box from to and one with a box from to . The second data set has the same total spread but a much more tightly packed middle half.
- A boxplot marks five positions on a number line. Positions do not carry counts, and the same five positions can come from 8 values or from 20. The sample size must be written in the label.
- It says the middle half is symmetric about the median: the quarter of the data just below the median is spread over the same width as the quarter just above it.
- Ordered: . Summary ; range ; IQR . Box from to , median line at , whiskers to and . Observation: the upper whisker ( to ) is longer than the lower whisker ( to ). Conclusion: a typical first hour brings about or customers, and while about a quarter of days run above , those busy days vary a lot, so the manager should staff for the median and have one person on call for the busiest quarter.
- Summary . Observations: the median high is F, and the middle half of days falls between and , an IQR of ; the range is . Conclusion: about a quarter of days topped , so a gardener should plan extra watering for roughly one day in four and should not treat the day as typical.
- The median line sits at the median, which is the middle value, not the midpoint between the box edges. Example: has median , , ; the line at sits far left of the box's midpoint, . The line is centered only when the middle half happens to be symmetric.
Exit ticket 16.4
- Nine values, median is the fifth, ; ; . Box from to , median line at , whiskers to and .
- Range ; IQR .
- About half.
- Any two of: the mean, any individual value other than the extremes, the number of values, whether a value repeats.
Lesson 16.5 — Extreme Data Points
Guided practice
- An extreme data point, or outlier, is a value that sits far away from the rest of the data.
- Panel 1: points. Panel 2: points.
- IQR goes from to , a change of just half a point. The IQR is computed from and , which are positions deep inside the ordered data. Adding one value at the far right shifts each of those positions by half a step, so they move to neighboring data values and the IQR barely changes.
- Range goes from to , more than tripling. The range is computed from the two most extreme values, and the new value is the new extreme, so the whole -point jump in the maximum passes straight into the range.
- Because is plotted as a separate point in that panel, so it is excluded from the whisker. The whisker then reaches the largest ordinary value, .
- The IQR. It is built from and , which sit in the middle of the ordered data where one distant value cannot reach, while the range is built from the extremes themselves.
Independent practice
- Original nine values: median ; lower half gives ; upper half gives . Range ; IQR . With added (ten values): median ; ; . Range ; IQR .
- The box stays almost exactly the same width and shifts right by a small amount, and the median line moves from to . The upper whisker stretches from ending at to ending at , so most of the width of the plot becomes one long whisker holding a single game — the picture looks strongly stretched to the right.
- Original nine scores: median ; ; ; IQR . With added (ten scores, sorted ): median ; ; ; IQR .
- The lower extreme changed the most, from to , which drops the range from to . The upper extreme, , did not change at all. (The IQR changed by only and the median by only .)
- The upper whisker, because it must reach all the way out to the outlier while the box stays with the bulk of the data.
- The median is a position, not a total. Adding one value moves the middle position by half a step, so the median slides to a neighboring data value no matter how enormous the added value is.
- Mean of the eleven games . With added, mean . The mean rose by points while the median rose by only . The mean uses the actual size of every value, so a huge value drags it; the median only counts positions, so it is resistant.
- Ordered: . Summary ; range ; IQR . The -minute delivery is an outlier — something went wrong on that one order. The range, minutes, is almost entirely the story of that single delivery. The IQR, minutes, says half of all orders arrive within a -minute window. Quote the median and the IQR to a customer — "about or minutes, and half of our orders land between and " — while still being honest that one delivery took an hour.
- Deleting a value because it is inconvenient changes the data to fit the picture, which is the wrong direction. Instead, check whether the value is a recording error; if it is, correct or remove it and say so. If it is real, keep it and report both the range (which includes it) and the IQR (which is not disturbed by it), and describe the outlier in words.
Exit ticket 16.5
- The range changes a lot, because it is computed from the extremes and the new value becomes the new extreme. The IQR changes very little, because it is computed from and , positions well inside the ordered data that shift by at most half a step.
- It stretches the plot toward the outlier: one whisker becomes very long while the box stays put, so the display looks lopsided, or skewed, in that direction.
- Remove it when it is a recording error, such as a reading time of minutes in one night — and say that you removed it. Keep it when it is a real event, such as a -point game, and report the IQR alongside the range.
Lesson 16.6 — Comparing Boxplots, Choosing a Display, and Spotting a Misleading One
Guided practice
- Class A: median , IQR . Class B: median , IQR .
- Class B. Its IQR is against Class A's , and its range is against Class A's , so by both measures of spread Class B's scores cluster far more tightly.
- Class A, whose upper extreme is against Class B's .
- Both medians are minutes. They are not equivalent because spread differs enormously: Route 1 has range and IQR , while Route 2 has range and IQR . Equal centers say nothing about reliability.
- Route 2. Its upper extreme is minutes, so no recorded trip took even an hour, and its IQR of means half its trips fall between and minutes. Route 1's upper extreme is minutes and its IQR is .
- The dot plot shows every individual value and lets you see repeats — for instance the two students who did sit-ups — and it allows the mean to be marked. A boxplot shows none of that.
- Only the axis changed. On the -to- axis the same box is drawn across a much larger fraction of the picture than on the -to- axis, so the eye reads the first as wide spread and the second as tight clustering. The five numbers are identical in both.
Independent practice
- : ; range , IQR . : ; range , IQR . Center: 's median of is higher than 's , so a typical value is a little larger. Spread: is about twice as consistent by either measure ( against for range, against for IQR). Overlap: the two boxes overlap heavily between and , but holds both the smallest value, , and the largest, .
- For example: two bus routes with the same median travel time of minutes, one with IQR and one with IQR . A typical trip takes the same time on both, but the second route is unpredictable — on any given day it might be much faster or much slower — so a rider who cannot be late should take the first.
- You can conclude that one group's typical value is higher than the other's, and that both groups span the same total distance from smallest to largest. You cannot conclude anything about how tightly each group clusters — equal ranges are consistent with very different IQRs — and you cannot conclude anything about the number of values in either group.
- a) A circle graph, because the question is about each category's share of one whole. b) A line graph, because the question is about change over time in one quantity, and order matters. c) Two boxplots on one shared axis, because the question compares two groups on both center and spread, and "reliable" is a spread question. d) A dot plot, because the question asks which exact value occurred most often, which requires seeing individual values and repeats.
- Because the two axes are not the same, equal-looking widths represent different amounts. A box drawn the same physical size on a -wide axis and on a -wide axis represents three times the spread on the second. Comparison by eye is impossible unless both plots share one axis and one scale.
- Any three of: the axis and its scale; the units; the number of values summarized; how the sample was chosen; whether both plots in a comparison share the same axis; whether any outliers were dropped.
- The middle half of the values is spread widely while the top quarter and the bottom quarter are packed into narrow intervals just outside the box — so the data clusters at the two ends of the box, with the extremes close by.
- Every one of the four sections holds about a quarter of the values, regardless of its length. A wide section means that quarter of the data is spread thinly across a wide interval; a narrow one means that quarter is crowded together. Width shows spread, not count.
- For example and a 20-value set with the same five-number summary. Both could produce the same picture, but a conclusion drawn from 6 values is far weaker than the same conclusion from 20 — and neither the count nor the sampling method is visible in the plot, so the reader cannot judge how much to trust it.
- Team X: ; range , IQR . Team Y: ; range , IQR . Team Y is more reliable: its IQR of is smaller than X's and its range of is smaller than X's , so its scores vary less from game to game. Y also has the higher median, against , and the higher floor, against . Team X produced the single best game, points.
- Two boxplots on one shared axis, one per grade, because the question compares two groups on both center and spread. Label the axis "minutes of homework per school night" with a numeric scale, label each plot with its grade, and state the sample size and how students were chosen for each grade.
- Box width shows spread, not sample size. A boxplot cannot display how many values it summarizes. A correct statement: "Group 1's box is wider, so the middle half of Group 1's values is spread over a larger interval — Group 1 is less consistent than Group 2."
Exit ticket 16.6
- The two groups have the same typical value, , but Group 2 is far less consistent: the middle half of its values spans units against Group 1's .
- Two boxplots drawn on one shared axis, because box width is the IQR and whisker length shows the reach of each quarter, so both measures of spread can be compared directly by eye.
- Any two of: a stretched or squeezed axis scale; two plots drawn on different scales; a missing or unlabeled axis; missing units; no sample size or sampling method given; an outlier dropped without a note.
- Because position on the picture only means a value when both plots use the same scale. On different axes, equal-looking boxes can represent very different spreads, so any visual comparison is meaningless.
Chapter 16 Review
Part A — The data cycle, collecting data, and bias (8.PS.2a, b, c)
- Formulate questions — decide what you want to know and word it so data can answer it. Collect or acquire data — gather the numbers yourself or find an existing set. Organize and represent data — order it, compute the summary, draw the boxplot. Analyze data and communicate results — read the plot, draw conclusions, report them.
- "What is the mass, in kilograms, of each of 20 randomly chosen eighth graders' backpacks on a normal school day?" Population: all eighth graders at the school. Data recorded: one backpack mass in kilograms per student.
- Direct measurement with a scale: number the eighth-grade roster, randomly select 20 students, and weigh each of those students' backpacks on the same scale on the same ordinary school day. Measurement is chosen because a scale gives an accurate number that students could only estimate.
- a) Selection bias by timing — on locker-cleanout Fridays backpacks are unusually heavy, so the results are pushed high. b) Voluntary response bias — students who volunteer may be those who think their backpack is remarkably heavy, pushing results high. c) Selection bias — athletes carry sports gear, so the results are pushed high.
- A representative sample is one whose values look like the population's values would, so that conclusions about the sample carry over to the population. Bias prevents this because it systematically favors part of the population, shifting the sample's center, spread, or both in a predictable direction that no later arithmetic can undo.
Part B — Organizing, representing, and describing (8.PS.2d, e)
- Twenty values, already in order. Median . Lower half (first ten) gives . Upper half (last ten) gives . Summary ; range ; IQR .
- Eight values. Median ; ; . Summary ; range ; IQR .
- Rule: order the data, split at the median, exclude the median from both halves when the count is odd, and take the median of each half. Applied: seven values, median is the fourth, ; lower half gives ; upper half gives . Summary .
- Summary . On the middle axis (scale to , ticks every ): box from to , median line at , whiskers to and .
- Summary . On the top axis (scale to , ticks every ): box from to , median line at , whiskers to and .
- Lower extreme , , median , , upper extreme books.
- Range books; IQR books. The range measures the total spread from the quietest day to the busiest. The IQR measures the spread of the middle half of the days, which describes an ordinary day without letting the two most unusual days affect it.
Part C — Outliers, analysis, and comparison (8.PS.2f, g, h)
- Observations (any two): the median is books; the middle half of days falls between and books; the busiest day saw and the quietest ; the upper whisker ( to , a span of ) is longer than the lower whisker ( to , a span of ). Conclusion: a typical day needs enough stock and staff for about checkouts, but about a quarter of days exceed , so the librarian should keep reserve capacity for roughly one day in four rather than staffing to the median alone.
- The upper whisker spans books ( to ) and the lower spans ( to ). Each holds about a quarter of the days, so the busiest quarter of days varies much more than the quietest quarter — busy days are unpredictable in how busy they get, while slow days are all similarly slow.
- The range is the difference of the two extremes. The added value of became the new maximum, so its entire distance from the old maximum passed straight into the range. The IQR is the difference of and , which are positions well inside the ordered data. Adding one value at the far end moves each of those positions by half a step, so they land on neighboring data values and the IQR is essentially unchanged.
- One very large value stretches the plot toward itself: the whisker on that side becomes very long while the box and the median line stay almost exactly where they were. The result is a lopsided, right-skewed picture in which most of the width of the display is occupied by a whisker holding a single value.
- Team X: ; range , IQR . Team Y: ; range , IQR . Center: Y's median of is a point higher than X's , so a typical Y game is slightly better. Spread: X is more spread out by both measures, with a range of against and an IQR of against , so X's results swing more from game to game.
- Team Y is more reliable, because its IQR of and range of are both smaller than Team X's and — its scores cluster more tightly, and its lower extreme of means it never had a game as poor as X's . Team X produced the single best game, its upper extreme of points, which is above Y's maximum of .
Part D — Choosing a display and spotting a misleading one (8.PS.2i, j)
- a) Two boxplots on one shared axis — the question compares two groups on spread, which box width and whisker length show directly. b) A circle graph — the question is about each option's share of one whole. c) A dot plot — the question requires every individual value to be visible, including repeats. d) A line graph — the question is about change over time, and order matters.
- A boxplot places values as positions on a number line and computes a median and quartiles from their order. Categories such as "soccer" and "band" have no numeric value and no meaningful order, so there is nothing to order, no middle value, and no distance to measure.
- Any three of: What is the scale on the axis, and what are the units? How many values does the plot summarize? How was the sample chosen, and who could not have been included? Were any extreme values left out? What exactly was measured, and when?
- The false impression is about spread: on a narrow axis a data set looks wildly spread out, and on a wide axis the same data set looks tightly clustered, so the reader concludes one group varies far more than the other when they may not. The fix is to redraw both boxplots on a single axis with one scale covering both data sets.
- Bias: voluntary response, and selection bias too — the 12 who volunteered are club members, a group that already reads more than average, and among them the volunteers are likely the keenest readers. So the sample cannot possibly represent "the average student." Misleading feature: the boxplot has no comparison group, so a claim about "more than average" is made from a display of only one group; it also hides the sample size and shows no individual values, so 12 volunteers look as authoritative as a proper sample. Better plan: define the population as all students at the school, randomly select up to 20 students from the full roster and up to 20 club members, record the same measurement (minutes read yesterday) the same way for both, and display the two groups as boxplots on one shared axis with the sample sizes and the sampling method stated. Then compare medians and IQRs.
Every data set used in this chapter has at most 20 items, as 8.PS.2d requires, and every summary above matches the figures in ../figures/, which are generated from the same data by virginia-math-sol/figures/make_grade8_ch16.py.