MathBored

Virginia SOL Mathematics Textbook

Appendix A — Answer Key, Chapter 16: The Data Cycle and Boxplots

SOL 8.PS.2 · Covers textbook Chapter 16 and the companion workbook. Item numbers match the textbook; workbook items are the same problems, so this key serves both. Item numbers run continuously from 1 to 150 across the chapter. Reasoning answers show an acceptable response, not the only wording.

Every five-number summary in this key was computed with the chapter's quartile rule — order the data, split at the median, exclude the median from both halves when the count is odd, and take the median of each half — and every range and IQR was recomputed from the summary it follows. A calculator using a different interpolation rule may report a slightly different quartile; on this chapter's work, use the rule above.

Several items ask students to write a question, choose a sample, or state a conclusion. For those, the key gives one acceptable response and names what any acceptable response must contain.


Lesson 16.1 — The Data Cycle, Pointed at Boxplots

Guided practice

  1. Formulate questions; collect or acquire data; organize and represent data; analyze data and communicate results.
  2. Stage 3, organize and represent data. The summary is what the boxplot is drawn from.
  3. Stage 2, collect or acquire data. (Accept: the sampling plan is decided as part of collection.)
  4. The cycle is not finished; the boxplot is only stage 3. You still have to analyze it, report what you found, and usually formulate a new question from what you saw.
  5. Shows: the median (center) and the spread, through the box width and whisker lengths — also the extremes and both quartiles. Does not show: individual data values, the number of values, repeats, or the mean.
  6. Stage 2. Acquiring existing data is a collection method, not a separate stage.

Independent practice

  1. a) stage 1 b) stage 3 c) stage 4 d) stage 2
  2. Because stage 4 almost always raises a new question, which sends you back to stage 1 with a sharper version of what you wanted to know.
  3. Any statistical question with a numeric answer that varies, for example "How many minutes did each student in our class spend on homework last night?"
  4. Any question that is not statistical ("How tall is our teacher?") or not numeric ("What is your favorite sport?"). A boxplot needs a set of numbers that vary; one fixed answer or a list of categories cannot be displayed as five positions on a number line.
  5. One number per student, in stated units — for item 9's example, the minutes of homework for one student on one named night.
  6. No. A boxplot reports five positions, not counts, so it cannot say how many students hit any particular score. It cannot even confirm that 8585 occurred at all unless 8585 is one of the extremes.
  7. No. The mean is not one of the five numbers plotted, and it cannot be recovered from them. Two data sets with the same five-number summary can have different means.
  8. Question: "How many hours of sleep did each of 20 randomly chosen eighth graders get last night?" Data needed: one number of hours per student. Collection: number the eighth-grade roster, use a random number generator to pick 20, and survey those students. Any acceptable answer names a numeric quantity with units and a way of choosing at most 20 students that does not favor a subgroup.
  9. A boxplot compares two or more groups on one axis far better, because two boxplots stack on the same scale and their centers and spreads line up visually. A histogram shows the shape of the distribution and how many values fall in each interval, which a boxplot cannot show at all.

Exit ticket 16.1

  1. Formulate questions; collect or acquire data; organize and represent data; analyze data and communicate results.
  2. Stage 3, organize and represent data.
  3. Center: "What is a typical value?" Spread: "How much do the values vary?" (Accept any equivalent pair.)
  4. The plot may reveal something unexpected — a very long whisker, a surprisingly wide box, or two groups that differ — and explaining it needs new data, which is a new question.

Lesson 16.2 — Asking the Question, Getting the Data, and Statistical Bias

Guided practice

  1. The population is the whole group you want to describe. The sample is the part of the population you actually collect data from.
  2. a) yes — counts vary from student to student b) no — one fixed answer, and not numeric c) yes — times vary from run to run
  3. Population: all 60 students who tried out. Sample: the 15 whose jumps were measured.
  4. Selection bias. Students who bring lunch from home have no chance of being included, and they are exactly the students most likely to dislike the school lunch, so the results would look too favorable.
  5. Response bias. The wording tells the respondent which answer is expected, so agreement is inflated.
  6. Random selection gives every member of the population an equal chance of being chosen, so no subgroup is systematically over- or underrepresented.

Independent practice

  1. a) "How many hours of sleep did each of 20 randomly chosen eighth graders at our school get last night?" b) "How many minutes did each of 20 randomly chosen students spend walking to school this morning?" c) "How many minutes did each of 20 randomly chosen eighth graders read for pleasure yesterday?" Any acceptable rewrite names who, what with units, and when.
  2. a) hours of sleep, in hours b) walking time, in minutes c) reading time, in minutes
  3. a) survey — only the student knows b) survey, or direct measurement with a stopwatch if you can observe the walk c) survey — only the student knows. Any answer that matches the method to who has access to the number is acceptable.
  4. a) Selection bias — people at a gym exercise more, so the results are pushed high. b) Nonresponse bias — households with slow or no internet are least likely to answer, so reported speeds are pushed high. c) Response bias — students may not answer honestly in front of classmates, so reported grades are pushed high. d) Voluntary response bias — students with strong feelings, usually enthusiastic ones, choose to respond, so school spirit looks higher than it is.
  5. Small and biased are different problems. A sample of 20 chosen at random is unbiased but imprecise; a sample of 2000 chosen only from the library is precise and still biased. Bias comes from how members were chosen, not from how many.
  6. No. Her own bus serves one neighborhood, so its ride times reflect that route's distance and traffic and not the whole system. That is selection bias, and it could push the median either high or low depending on the route.
  7. Number the roster from 1 to 340, then use a random number generator (or draw numbered slips from a container) to choose 20 different numbers, and survey exactly those students.
  8. Download daily high temperatures for your city from the National Weather Service. Check who collected the data, when and where it was measured, and what units it is reported in.
  9. Question: "How many minutes did each of 20 randomly chosen eighth graders at our school spend on homework last night?" Population: all eighth graders at the school. Plan: number the eighth-grade roster, randomly select 20, and survey them privately. Bias avoided: selection bias, because students in every class and activity have an equal chance, and response bias, because the survey is private rather than read aloud.
  10. Friends are not a random sample of the school. They tend to share classes, schedules, and habits with the student collecting the data, so the sample is systematically like her and unlike the rest of the school. The size, 20, is fine; the selection method is the problem.

Exit ticket 16.2

  1. Statistical bias is any feature of how the data was collected that pushes the sample away from the population in a predictable direction.
  2. Population: all 200 band members. Sample: the 18 who were timed.
  3. Selection bias — car riders live farther away on average than walkers, so the distances collected are pushed high. Fix: select students at random from the full roster instead of from one arrival group.
  4. Bias lives in the collection method, not in the values. The numbers themselves look perfectly ordinary; what tells you the sample is biased is knowing who could not possibly have been included.

Lesson 16.3 — The Five-Number Summary

Guided practice

  1. Ordered: 4,7,9,12,154, 7, 9, 12, 15. Median =9= 9 (the third of five values).
  2. Lower half 4,74, 7 gives Q1=4+72=5.5Q_1 = \frac{4+7}{2} = 5.5. Upper half 12,1512, 15 gives Q3=12+152=13.5Q_3 = \frac{12+15}{2} = 13.5. The median, 99, is in neither half.
  3. Range =154=11= 15 - 4 = 11; IQR =13.55.5=8= 13.5 - 5.5 = 8.
  4. 2, 4, 7, 10, 122,\ 4,\ 7,\ 10,\ 12. With six values the median is 6+82=7\frac{6+8}{2} = 7; lower half 2,4,62, 4, 6 gives Q1=4Q_1 = 4; upper half 8,10,128, 10, 12 gives Q3=10Q_3 = 10.
  5. Range =122=10= 12 - 2 = 10; IQR =104=6= 10 - 4 = 6.
  6. 3, 5, 10, 15, 183,\ 5,\ 10,\ 15,\ 18. Median is the fourth of seven values, 1010; lower half 3,5,83, 5, 8 gives Q1=5Q_1 = 5; upper half 12,15,1812, 15, 18 gives Q3=15Q_3 = 15.
  7. The count is odd, so the median is an actual data value sitting exactly in the middle. This book's rule excludes it from both halves, which leaves three values on each side and makes the two halves the same size.
  8. Order the data. Find the median; if the count is odd, set that value aside. Q1Q_1 is the median of the values below it and Q3Q_3 is the median of the values above it.

Independent practice

  1. 45, 52, 59, 65, 7545,\ 52,\ 59,\ 65,\ 75. Ten values: median =58+602=59= \frac{58+60}{2} = 59; lower half 45,50,52,55,5845, 50, 52, 55, 58 gives Q1=52Q_1 = 52; upper half 60,62,65,70,7560, 62, 65, 70, 75 gives Q3=65Q_3 = 65. Range =30= 30; IQR =13= 13.
  2. 11, 13.5, 17.5, 22, 2811,\ 13.5,\ 17.5,\ 22,\ 28. Eight values: median =16+192=17.5= \frac{16+19}{2} = 17.5; Q1=13+142=13.5Q_1 = \frac{13+14}{2} = 13.5; Q3=21+232=22Q_3 = \frac{21+23}{2} = 22. Range =17= 17; IQR =8.5= 8.5.
  3. 58, 64, 69, 75.5, 8458,\ 64,\ 69,\ 75.5,\ 84. Twelve values: median =68+702=69= \frac{68+70}{2} = 69; Q1=63+652=64Q_1 = \frac{63+65}{2} = 64; Q3=74+772=75.5Q_3 = \frac{74+77}{2} = 75.5. Range =26= 26; IQR =11.5= 11.5.
  4. 2, 3, 7.5, 9, 142,\ 3,\ 7.5,\ 9,\ 14. Repeated values are treated as separate values and stay in the list; nothing is combined. Ten values: median =7+82=7.5= \frac{7+8}{2} = 7.5; lower half 2,3,3,5,72, 3, 3, 5, 7 gives Q1=3Q_1 = 3; upper half 8,9,9,10,148, 9, 9, 10, 14 gives Q3=9Q_3 = 9. Range =12= 12; IQR =6= 6.
  5. 55, 58.5, 61.5, 65, 7055,\ 58.5,\ 61.5,\ 65,\ 70. Median =60+632=61.5= \frac{60+63}{2} = 61.5; Q1=57+602=58.5Q_1 = \frac{57+60}{2} = 58.5; Q3=64+662=65Q_3 = \frac{64+66}{2} = 65.
  6. First set: seven values, median is the fourth, 44. Second set: nine values, median is the fifth, 1616.
  7. 12+30=4212 + 30 = 42.
  8. Q3=18+11=29Q_3 = 18 + 11 = 29.
  9. Q1Q_1 is never below the minimum and Q3Q_3 is never above the maximum, so the interval from Q1Q_1 to Q3Q_3 always sits inside the interval from minimum to maximum. A shorter interval cannot have a greater length. They are equal only when every value is the same.
  10. For example 1,2,5,8,91, 2, 5, 8, 9 (median 55, Q1=1.5Q_1 = 1.5, Q3=8.5Q_3 = 8.5, IQR =7= 7) and 4,4,5,6,64, 4, 5, 6, 6 (median 55, Q1=4Q_1 = 4, Q3=6Q_3 = 6, IQR =2= 2). Any pair with equal medians and different IQRs is acceptable.
  11. Ordered: 5.8,5.9,6.0,6.2,6.3,6.5,6.9,7.1,7.45.8, 5.9, 6.0, 6.2, 6.3, 6.5, 6.9, 7.1, 7.4. Median is the fifth value, 6.36.3; lower half 5.8,5.9,6.0,6.25.8, 5.9, 6.0, 6.2 gives Q1=5.9+6.02=5.95Q_1 = \frac{5.9+6.0}{2} = 5.95; upper half 6.5,6.9,7.1,7.46.5, 6.9, 7.1, 7.4 gives Q3=6.9+7.12=7.0Q_3 = \frac{6.9+7.1}{2} = 7.0. Summary 5.8, 5.95, 6.3, 7.0, 7.45.8,\ 5.95,\ 6.3,\ 7.0,\ 7.4; range =1.6= 1.6 s; IQR =1.05= 1.05 s.
  12. The student read positions off the unsorted list. Quartiles are positions in the ordered data. Sorted: 3,6,8,11,153, 6, 8, 11, 15. Median =8= 8; Q1=3+62=4.5Q_1 = \frac{3+6}{2} = 4.5; Q3=11+152=13Q_3 = \frac{11+15}{2} = 13. Summary 3, 4.5, 8, 13, 153,\ 4.5,\ 8,\ 13,\ 15; range =12= 12; IQR =8.5= 8.5.

Exit ticket 16.3

  1. 6, 7.5, 11, 17, 206,\ 7.5,\ 11,\ 17,\ 20. Median is the third of five, 1111; Q1=6+92=7.5Q_1 = \frac{6+9}{2} = 7.5; Q3=14+202=17Q_3 = \frac{14+20}{2} = 17.
  2. Range =206=14= 20 - 6 = 14; IQR =177.5=9.5= 17 - 7.5 = 9.5.
  3. 22, 25, 29, 31, 3422,\ 25,\ 29,\ 31,\ 34. Median =28+302=29= \frac{28+30}{2} = 29; lower half 22,25,2822, 25, 28 gives Q1=25Q_1 = 25; upper half 30,31,3430, 31, 34 gives Q3=31Q_3 = 31.
  4. The IQR measures the spread of the middle half only, so it describes how tightly the typical values cluster and ignores the two ends, where a single unusual value can stretch the range.

Lesson 16.4 — Building and Reading a Boxplot

Guided practice

  1. Lower extreme 1212, Q1=20Q_1 = 20, median 3030, Q3=40Q_3 = 40, upper extreme 5555 minutes.
  2. Range =5512=43= 55 - 12 = 43 minutes; IQR =4020=20= 40 - 20 = 20 minutes.
  3. The upper whisker, which runs from 4040 to 5555 (a span of 1515), is longer than the lower whisker, which runs from 1212 to 2020 (a span of 88). Each whisker holds about a quarter of the students, so the longest readers are spread across a wider range of times than the shortest readers — the distribution is stretched toward the high end.
  4. Lower extreme 66, Q1=11Q_1 = 11, median 16.516.5, Q3=24Q_3 = 24, upper extreme 3636 minutes.
  5. About one quarter, since 2424 is Q3Q_3.
  6. Box from 55 to 1515, median line at 1010, whiskers to 33 and 1818.
  7. Box from 44 to 1010, median line at 77, whiskers to 22 and 1212.

Independent practice

  1. Summary 21, 24.5, 30, 36.5, 4021,\ 24.5,\ 30,\ 36.5,\ 40. Box from 24.524.5 to 36.536.5, median line at 3030, whiskers to 2121 and 4040. Range =19= 19; IQR =12= 12.
  2. Summary 45, 52, 59, 65, 7545,\ 52,\ 59,\ 65,\ 75. Box from 5252 to 6565, median line at 5959, whiskers to 4545 and 7575.
  3. Summary 58, 64, 69, 75.5, 8458,\ 64,\ 69,\ 75.5,\ 84. Box from 6464 to 75.575.5, median line at 6969, whiskers to 5858 and 8484.
  4. Lower extreme 1010, Q1=18Q_1 = 18, median 2222, Q3=34Q_3 = 34, upper extreme 4646, range =36= 36, IQR =16= 16.
  5. The part above the median is longer: 3422=1234 - 22 = 12 against 2218=422 - 18 = 4. Both parts hold about a quarter of the data, so the values just above the median are spread out while the values just below it are packed tightly.
  6. About half, since 1818 and 3434 are Q1Q_1 and Q3Q_3.
  7. Any two plots whose whisker tips are the same distance apart but whose boxes differ in width — for example both running from 1010 to 5050, one with a box from 1515 to 4545 and one with a box from 2828 to 3232. The second data set has the same total spread but a much more tightly packed middle half.
  8. A boxplot marks five positions on a number line. Positions do not carry counts, and the same five positions can come from 8 values or from 20. The sample size must be written in the label.
  9. It says the middle half is symmetric about the median: the quarter of the data just below the median is spread over the same width as the quarter just above it.
  10. Ordered: 11,13,14,16,19,21,23,2811, 13, 14, 16, 19, 21, 23, 28. Summary 11, 13.5, 17.5, 22, 2811,\ 13.5,\ 17.5,\ 22,\ 28; range =17= 17; IQR =8.5= 8.5. Box from 13.513.5 to 2222, median line at 17.517.5, whiskers to 1111 and 2828. Observation: the upper whisker (2222 to 2828) is longer than the lower whisker (1111 to 13.513.5). Conclusion: a typical first hour brings about 1717 or 1818 customers, and while about a quarter of days run above 2222, those busy days vary a lot, so the manager should staff for the median and have one person on call for the busiest quarter.
  11. Summary 58, 64, 69, 75.5, 8458,\ 64,\ 69,\ 75.5,\ 84. Observations: the median high is 6969^\circF, and the middle half of days falls between 6464^\circ and 75.575.5^\circ, an IQR of 11.511.5^\circ; the range is 2626^\circ. Conclusion: about a quarter of days topped 75.575.5^\circ, so a gardener should plan extra watering for roughly one day in four and should not treat the 8484^\circ day as typical.
  12. The median line sits at the median, which is the middle value, not the midpoint between the box edges. Example: 1,2,3,4,20,21,221, 2, 3, 4, 20, 21, 22 has median 44, Q1=2Q_1 = 2, Q3=21Q_3 = 21; the line at 44 sits far left of the box's midpoint, 11.511.5. The line is centered only when the middle half happens to be symmetric.

Exit ticket 16.4

  1. Nine values, median is the fifth, 1616; Q1=12+132=12.5Q_1 = \frac{12+13}{2} = 12.5; Q3=19+202=19.5Q_3 = \frac{19+20}{2} = 19.5. Box from 12.512.5 to 19.519.5, median line at 1616, whiskers to 1010 and 2222.
  2. Range =295=24= 29 - 5 = 24; IQR =2012=8= 20 - 12 = 8.
  3. About half.
  4. Any two of: the mean, any individual value other than the extremes, the number of values, whether a value repeats.

Lesson 16.5 — Extreme Data Points

Guided practice

  1. An extreme data point, or outlier, is a value that sits far away from the rest of the data.
  2. Panel 1: 3636 points. Panel 2: 7676 points.
  3. IQR goes from 88 to 8.58.5, a change of just half a point. The IQR is computed from Q1Q_1 and Q3Q_3, which are positions deep inside the ordered data. Adding one value at the far right shifts each of those positions by half a step, so they move to neighboring data values and the IQR barely changes.
  4. Range goes from 1616 to 5656, more than tripling. The range is computed from the two most extreme values, and the new value is the new extreme, so the whole 4040-point jump in the maximum passes straight into the range.
  5. Because 7676 is plotted as a separate point in that panel, so it is excluded from the whisker. The whisker then reaches the largest ordinary value, 3636.
  6. The IQR. It is built from Q1Q_1 and Q3Q_3, which sit in the middle of the ordered data where one distant value cannot reach, while the range is built from the extremes themselves.

Independent practice

  1. Original nine values: median 99; lower half 5,6,7,85, 6, 7, 8 gives Q1=6.5Q_1 = 6.5; upper half 10,11,12,1310, 11, 12, 13 gives Q3=11.5Q_3 = 11.5. Range =135=8= 13 - 5 = 8; IQR =11.56.5=5= 11.5 - 6.5 = 5. With 4040 added (ten values): median =9+102=9.5= \frac{9+10}{2} = 9.5; Q1=7Q_1 = 7; Q3=12Q_3 = 12. Range =405=35= 40 - 5 = 35; IQR =127=5= 12 - 7 = 5.
  2. The box stays almost exactly the same width and shifts right by a small amount, and the median line moves from 99 to 9.59.5. The upper whisker stretches from ending at 1313 to ending at 4040, so most of the width of the plot becomes one long whisker holding a single game — the picture looks strongly stretched to the right.
  3. Original nine scores: median 7878; Q1=74+752=74.5Q_1 = \frac{74+75}{2} = 74.5; Q3=82+842=83Q_3 = \frac{82+84}{2} = 83; IQR =8.5= 8.5. With 1212 added (ten scores, sorted 12,71,74,75,77,78,80,82,84,8612, 71, 74, 75, 77, 78, 80, 82, 84, 86): median =77+782=77.5= \frac{77+78}{2} = 77.5; Q1=74Q_1 = 74; Q3=82Q_3 = 82; IQR =8= 8.
  4. The lower extreme changed the most, from 7171 to 1212, which drops the range from 1515 to 7474. The upper extreme, 8686, did not change at all. (The IQR changed by only 0.50.5 and the median by only 0.50.5.)
  5. The upper whisker, because it must reach all the way out to the outlier while the box stays with the bulk of the data.
  6. The median is a position, not a total. Adding one value moves the middle position by half a step, so the median slides to a neighboring data value no matter how enormous the added value is.
  7. Mean of the eleven games =30811=28= \frac{308}{11} = 28. With 7676 added, mean =38412=32= \frac{384}{12} = 32. The mean rose by 44 points while the median rose by only 11. The mean uses the actual size of every value, so a huge value drags it; the median only counts positions, so it is resistant.
  8. Ordered: 18,20,21,22,24,25,26,28,30,6218, 20, 21, 22, 24, 25, 26, 28, 30, 62. Summary 18, 21, 24.5, 28, 6218,\ 21,\ 24.5,\ 28,\ 62; range =44= 44; IQR =7= 7. The 6262-minute delivery is an outlier — something went wrong on that one order. The range, 4444 minutes, is almost entirely the story of that single delivery. The IQR, 77 minutes, says half of all orders arrive within a 77-minute window. Quote the median and the IQR to a customer — "about 2424 or 2525 minutes, and half of our orders land between 2121 and 2828" — while still being honest that one delivery took an hour.
  9. Deleting a value because it is inconvenient changes the data to fit the picture, which is the wrong direction. Instead, check whether the value is a recording error; if it is, correct or remove it and say so. If it is real, keep it and report both the range (which includes it) and the IQR (which is not disturbed by it), and describe the outlier in words.

Exit ticket 16.5

  1. The range changes a lot, because it is computed from the extremes and the new value becomes the new extreme. The IQR changes very little, because it is computed from Q1Q_1 and Q3Q_3, positions well inside the ordered data that shift by at most half a step.
  2. It stretches the plot toward the outlier: one whisker becomes very long while the box stays put, so the display looks lopsided, or skewed, in that direction.
  3. Remove it when it is a recording error, such as a reading time of 20002000 minutes in one night — and say that you removed it. Keep it when it is a real event, such as a 7676-point game, and report the IQR alongside the range.

Lesson 16.6 — Comparing Boxplots, Choosing a Display, and Spotting a Misleading One

Guided practice

  1. Class A: median 7575, IQR =8565=20= 85 - 65 = 20. Class B: median 7878, IQR =8273=9= 82 - 73 = 9.
  2. Class B. Its IQR is 99 against Class A's 2020, and its range is 2020 against Class A's 4040, so by both measures of spread Class B's scores cluster far more tightly.
  3. Class A, whose upper extreme is 9595 against Class B's 8888.
  4. Both medians are 5050 minutes. They are not equivalent because spread differs enormously: Route 1 has range 8080 and IQR 4040, while Route 2 has range 2020 and IQR 88. Equal centers say nothing about reliability.
  5. Route 2. Its upper extreme is 6060 minutes, so no recorded trip took even an hour, and its IQR of 88 means half its trips fall between 4646 and 5454 minutes. Route 1's upper extreme is 9090 minutes and its IQR is 4040.
  6. The dot plot shows every individual value and lets you see repeats — for instance the two students who did 2121 sit-ups — and it allows the mean to be marked. A boxplot shows none of that.
  7. Only the axis changed. On the 5050-to-100100 axis the same box is drawn across a much larger fraction of the picture than on the 00-to-150150 axis, so the eye reads the first as wide spread and the second as tight clustering. The five numbers are identical in both.

Independent practice

  1. PP: 12, 15, 19, 24, 3012,\ 15,\ 19,\ 24,\ 30; range =18= 18, IQR =9= 9. QQ: 16, 18, 20.5, 23, 2516,\ 18,\ 20.5,\ 23,\ 25; range =9= 9, IQR =5= 5. Center: QQ's median of 20.520.5 is higher than PP's 1919, so a typical QQ value is a little larger. Spread: QQ is about twice as consistent by either measure (99 against 1818 for range, 55 against 99 for IQR). Overlap: the two boxes overlap heavily between 1818 and 2323, but PP holds both the smallest value, 1212, and the largest, 3030.
  2. For example: two bus routes with the same median travel time of 3030 minutes, one with IQR 44 and one with IQR 1818. A typical trip takes the same time on both, but the second route is unpredictable — on any given day it might be much faster or much slower — so a rider who cannot be late should take the first.
  3. You can conclude that one group's typical value is higher than the other's, and that both groups span the same total distance from smallest to largest. You cannot conclude anything about how tightly each group clusters — equal ranges are consistent with very different IQRs — and you cannot conclude anything about the number of values in either group.
  4. a) A circle graph, because the question is about each category's share of one whole. b) A line graph, because the question is about change over time in one quantity, and order matters. c) Two boxplots on one shared axis, because the question compares two groups on both center and spread, and "reliable" is a spread question. d) A dot plot, because the question asks which exact value occurred most often, which requires seeing individual values and repeats.
  5. Because the two axes are not the same, equal-looking widths represent different amounts. A box drawn the same physical size on a 100100-wide axis and on a 300300-wide axis represents three times the spread on the second. Comparison by eye is impossible unless both plots share one axis and one scale.
  6. Any three of: the axis and its scale; the units; the number of values summarized; how the sample was chosen; whether both plots in a comparison share the same axis; whether any outliers were dropped.
  7. The middle half of the values is spread widely while the top quarter and the bottom quarter are packed into narrow intervals just outside the box — so the data clusters at the two ends of the box, with the extremes close by.
  8. Every one of the four sections holds about a quarter of the values, regardless of its length. A wide section means that quarter of the data is spread thinly across a wide interval; a narrow one means that quarter is crowded together. Width shows spread, not count.
  9. For example {10,20,30,40,50,60}\{10, 20, 30, 40, 50, 60\} and a 20-value set with the same five-number summary. Both could produce the same picture, but a conclusion drawn from 6 values is far weaker than the same conclusion from 20 — and neither the count nor the sampling method is visible in the plot, so the reader cannot judge how much to trust it.
  10. Team X: 4, 8, 12.5, 17, 264,\ 8,\ 12.5,\ 17,\ 26; range =22= 22, IQR =9= 9. Team Y: 8, 10.5, 13.5, 16.5, 228,\ 10.5,\ 13.5,\ 16.5,\ 22; range =14= 14, IQR =6= 6. Team Y is more reliable: its IQR of 66 is smaller than X's 99 and its range of 1414 is smaller than X's 2222, so its scores vary less from game to game. Y also has the higher median, 13.513.5 against 12.512.5, and the higher floor, 88 against 44. Team X produced the single best game, 2626 points.
  11. Two boxplots on one shared axis, one per grade, because the question compares two groups on both center and spread. Label the axis "minutes of homework per school night" with a numeric scale, label each plot with its grade, and state the sample size and how students were chosen for each grade.
  12. Box width shows spread, not sample size. A boxplot cannot display how many values it summarizes. A correct statement: "Group 1's box is wider, so the middle half of Group 1's values is spread over a larger interval — Group 1 is less consistent than Group 2."

Exit ticket 16.6

  1. The two groups have the same typical value, 4040, but Group 2 is far less consistent: the middle half of its values spans 2222 units against Group 1's 66.
  2. Two boxplots drawn on one shared axis, because box width is the IQR and whisker length shows the reach of each quarter, so both measures of spread can be compared directly by eye.
  3. Any two of: a stretched or squeezed axis scale; two plots drawn on different scales; a missing or unlabeled axis; missing units; no sample size or sampling method given; an outlier dropped without a note.
  4. Because position on the picture only means a value when both plots use the same scale. On different axes, equal-looking boxes can represent very different spreads, so any visual comparison is meaningless.

Chapter 16 Review

Part A — The data cycle, collecting data, and bias (8.PS.2a, b, c)

  1. Formulate questions — decide what you want to know and word it so data can answer it. Collect or acquire data — gather the numbers yourself or find an existing set. Organize and represent data — order it, compute the summary, draw the boxplot. Analyze data and communicate results — read the plot, draw conclusions, report them.
  2. "What is the mass, in kilograms, of each of 20 randomly chosen eighth graders' backpacks on a normal school day?" Population: all eighth graders at the school. Data recorded: one backpack mass in kilograms per student.
  3. Direct measurement with a scale: number the eighth-grade roster, randomly select 20 students, and weigh each of those students' backpacks on the same scale on the same ordinary school day. Measurement is chosen because a scale gives an accurate number that students could only estimate.
  4. a) Selection bias by timing — on locker-cleanout Fridays backpacks are unusually heavy, so the results are pushed high. b) Voluntary response bias — students who volunteer may be those who think their backpack is remarkably heavy, pushing results high. c) Selection bias — athletes carry sports gear, so the results are pushed high.
  5. A representative sample is one whose values look like the population's values would, so that conclusions about the sample carry over to the population. Bias prevents this because it systematically favors part of the population, shifting the sample's center, spread, or both in a predictable direction that no later arithmetic can undo.

Part B — Organizing, representing, and describing (8.PS.2d, e)

  1. Twenty values, already in order. Median =6+72=6.5= \frac{6+7}{2} = 6.5. Lower half (first ten) 1,2,2,3,4,4,5,5,6,61, 2, 2, 3, 4, 4, 5, 5, 6, 6 gives Q1=4+42=4Q_1 = \frac{4+4}{2} = 4. Upper half (last ten) 7,7,8,8,9,10,11,12,14,187, 7, 8, 8, 9, 10, 11, 12, 14, 18 gives Q3=9+102=9.5Q_3 = \frac{9+10}{2} = 9.5. Summary 1, 4, 6.5, 9.5, 181,\ 4,\ 6.5,\ 9.5,\ 18; range =17= 17; IQR =5.5= 5.5.
  2. Eight values. Median =108+1102=109= \frac{108+110}{2} = 109; Q1=102+1052=103.5Q_1 = \frac{102+105}{2} = 103.5; Q3=115+1202=117.5Q_3 = \frac{115+120}{2} = 117.5. Summary 100, 103.5, 109, 117.5, 125100,\ 103.5,\ 109,\ 117.5,\ 125; range =25= 25; IQR =14= 14.
  3. Rule: order the data, split at the median, exclude the median from both halves when the count is odd, and take the median of each half. Applied: seven values, median is the fourth, 1010; lower half 3,5,83, 5, 8 gives Q1=5Q_1 = 5; upper half 12,15,1812, 15, 18 gives Q3=15Q_3 = 15. Summary 3, 5, 10, 15, 183,\ 5,\ 10,\ 15,\ 18.
  4. Summary 11, 13.5, 17.5, 22, 2811,\ 13.5,\ 17.5,\ 22,\ 28. On the middle axis (scale 00 to 6060, ticks every 55): box from 13.513.5 to 2222, median line at 17.517.5, whiskers to 1111 and 2828.
  5. Summary 6, 11, 16.5, 24, 366,\ 11,\ 16.5,\ 24,\ 36. On the top axis (scale 00 to 4040, ticks every 55): box from 1111 to 2424, median line at 16.516.5, whiskers to 66 and 3636.
  6. Lower extreme 33, Q1=7Q_1 = 7, median 1212, Q3=17Q_3 = 17, upper extreme 2424 books.
  7. Range =243=21= 24 - 3 = 21 books; IQR =177=10= 17 - 7 = 10 books. The range measures the total spread from the quietest day to the busiest. The IQR measures the spread of the middle half of the days, which describes an ordinary day without letting the two most unusual days affect it.

Part C — Outliers, analysis, and comparison (8.PS.2f, g, h)

  1. Observations (any two): the median is 1212 books; the middle half of days falls between 77 and 1717 books; the busiest day saw 2424 and the quietest 33; the upper whisker (1717 to 2424, a span of 77) is longer than the lower whisker (33 to 77, a span of 44). Conclusion: a typical day needs enough stock and staff for about 1212 checkouts, but about a quarter of days exceed 1717, so the librarian should keep reserve capacity for roughly one day in four rather than staffing to the median alone.
  2. The upper whisker spans 77 books (1717 to 2424) and the lower spans 44 (33 to 77). Each holds about a quarter of the days, so the busiest quarter of days varies much more than the quietest quarter — busy days are unpredictable in how busy they get, while slow days are all similarly slow.
  3. The range is the difference of the two extremes. The added value of 4040 became the new maximum, so its entire distance from the old maximum passed straight into the range. The IQR is the difference of Q1Q_1 and Q3Q_3, which are positions well inside the ordered data. Adding one value at the far end moves each of those positions by half a step, so they land on neighboring data values and the IQR is essentially unchanged.
  4. One very large value stretches the plot toward itself: the whisker on that side becomes very long while the box and the median line stay almost exactly where they were. The result is a lopsided, right-skewed picture in which most of the width of the display is occupied by a whisker holding a single value.
  5. Team X: 4, 8, 12.5, 17, 264,\ 8,\ 12.5,\ 17,\ 26; range =22= 22, IQR =9= 9. Team Y: 8, 10.5, 13.5, 16.5, 228,\ 10.5,\ 13.5,\ 16.5,\ 22; range =14= 14, IQR =6= 6. Center: Y's median of 13.513.5 is a point higher than X's 12.512.5, so a typical Y game is slightly better. Spread: X is more spread out by both measures, with a range of 2222 against 1414 and an IQR of 99 against 66, so X's results swing more from game to game.
  6. Team Y is more reliable, because its IQR of 66 and range of 1414 are both smaller than Team X's 99 and 2222 — its scores cluster more tightly, and its lower extreme of 88 means it never had a game as poor as X's 44. Team X produced the single best game, its upper extreme of 2626 points, which is above Y's maximum of 2222.

Part D — Choosing a display and spotting a misleading one (8.PS.2i, j)

  1. a) Two boxplots on one shared axis — the question compares two groups on spread, which box width and whisker length show directly. b) A circle graph — the question is about each option's share of one whole. c) A dot plot — the question requires every individual value to be visible, including repeats. d) A line graph — the question is about change over time, and order matters.
  2. A boxplot places values as positions on a number line and computes a median and quartiles from their order. Categories such as "soccer" and "band" have no numeric value and no meaningful order, so there is nothing to order, no middle value, and no distance to measure.
  3. Any three of: What is the scale on the axis, and what are the units? How many values does the plot summarize? How was the sample chosen, and who could not have been included? Were any extreme values left out? What exactly was measured, and when?
  4. The false impression is about spread: on a narrow axis a data set looks wildly spread out, and on a wide axis the same data set looks tightly clustered, so the reader concludes one group varies far more than the other when they may not. The fix is to redraw both boxplots on a single axis with one scale covering both data sets.
  5. Bias: voluntary response, and selection bias too — the 12 who volunteered are club members, a group that already reads more than average, and among them the volunteers are likely the keenest readers. So the sample cannot possibly represent "the average student." Misleading feature: the boxplot has no comparison group, so a claim about "more than average" is made from a display of only one group; it also hides the sample size and shows no individual values, so 12 volunteers look as authoritative as a proper sample. Better plan: define the population as all students at the school, randomly select up to 20 students from the full roster and up to 20 club members, record the same measurement (minutes read yesterday) the same way for both, and display the two groups as boxplots on one shared axis with the sample sizes and the sampling method stated. Then compare medians and IQRs.

Every data set used in this chapter has at most 20 items, as 8.PS.2d requires, and every summary above matches the figures in ../figures/, which are generated from the same data by virginia-math-sol/figures/make_grade8_ch16.py.