MisleadingCharts
Back to the gallery

The 28-fold spread drawn in the bottom sixteenth of the frame

Showing the misleading chart

Be the first to star this exhibit

A procurement slide we drew puts a storage operator’s published annual failure rates on a linear, zero-based, evenly ticked percentage axis that runs the whole way to 100% — because the quantity is a percentage — and all 31 drive models land in a row of stubs along the floor. Nothing is bent: the ratios between the columns are exact. They are just too small to compare. The field runs from 0.22% to 6.30%, a 28.6-fold range that occupies 61 pixels of a plot a thousand pixels tall, and the six models at or above 3% hold 5.8% of the fleet while producing 21.5% of the year’s 4,317 failures. Bring the top of the axis down to 7.5%, swap the bars for dots, and the flat field becomes a ladder — one whose rungs turn out to be firmware, vibration and age rather than vendor, which is an argument you cannot even start from the first chart.

01The claim

Nothing in this fleet fails, so buy on price. Every figure on the slide is the operator’s own published annual table: 31 drive models, 344,196 drives, 115,638,676 drive days, 4,317 failures, one unit, one publisher, one twelve-month window. Nothing is sampled, smoothed, indexed or rescaled, and no model has been dropped. The axis is linear, evenly ticked every ten points, unbroken, and starts at zero — and it runs the full range the unit can take, 0% to 100%, because the quantity is a percentage. Drawn that way, all 31 models are a row of stubs along the foot of the chart. The whole field lands between 0.22% and 6.30% around a fleet-wide 1.36%, and no model is visibly adrift from any other. What the pack concluded: reliability is not separating the vendors. Recommendation: drop drive reliability from the scoring matrix for the 2026 build and award on price per terabyte. Revisit if any model ever clears 10%.

02The trick

The ratios on this chart are exact. Every column starts at zero, the ticks are even, nothing is broken, clipped, logged or dropped, and the model that fails 28.6 times as often as the best one really is drawn 28.6 times as tall. The trouble is how tall that is. The top of the axis came from the unit rather than from the numbers — a percentage runs to a hundred, so the axis runs to a hundred — and the data occupies the bottom 6.3% of the frame. In this plot, a thousand pixels tall, one percentage point is ten pixels: the tallest column on the chart is 63 of them, the entire distance between the best model and the worst is 61, and 22 of the 31 columns stand shorter than a fortieth of the frame. Rank order survives that — you can still see which end is which — but magnitude does not, and magnitude is the whole question. Nobody reads 28.6 times off a row of stubs, the twenty-two models in the bottom two thirds of the table are mutually indistinguishable, and 93.7% of the box is empty space doing the arguing. What the numbers say instead: the field runs from 0.22% (Seagate ST16000NM002J) to 6.30% (Toshiba MG08ACA16TEY), which is 2.2 against 63 failures per thousand drives a year. The extremes are small samples, so take the largest population in the table against the largest of the high-rate models: WDC WUH722222ALE6L4, 177 failures across 44,577 drives, 0.47% with a 95% interval of 0.41 to 0.55, against HGST HUH721212ALN604, 393 failures across 10,195 drives, 3.95% with an interval of 3.57 to 4.36 — 8.4 times, the intervals nowhere near each other, and per ten thousand drives the difference between 47 replacements a year and 395. The six models at or above 3% hold 19,880 drives, 5.8% of the fleet, and produced 929 of the year’s 4,317 failures, 21.5% of them. Then the part the redraw makes possible: asking why. The operator has already answered. Its report calls three models “red flags” on their fourth-quarter rates and explains each one — the 16TB Toshiba at the top of the ladder had firmware work rolled out with the manufacturer and a 16.95% quarter before this one, which Backblaze describes as “a healthy normalization”; the 8TB HGST is a single sub-Vault of 7.5-year-old drives with suspected vibration; the 10TB Seagate is at end of life. Age runs underneath all of it: the six models at or above 3% have a median average age of 75.5 months against 31.9 for the thirteen under 1%, and pooled, the fifteen models averaging over four years old ran at 1.90% against 0.72% for the eight under two years. So the ladder is not a vendor league table, and reading it as one would be its own mistake. It is firmware, vibration, wear and vintage — every one of which is a question you can only ask once you can see that the rungs are there at all. (The slide is our demonstration in the manner of an internal procurement pack — the desk, the conclusion and the recommendation are invented; every figure behind them is Backblaze’s own, and each model’s rate reproduces from its own failure count and drive days.)

03The fix

Bring the top of the axis down to a little above the largest value — here 7.5%, which also clears the widest interval on the chart, so nothing is clipped — and say in the caption where you put it. That is the whole repair, and it is one line of configuration. The reason it feels illegal is the first rule anybody learns about axes, that bars start at zero, and the way through is to notice that the rule is about the mark rather than about the frame: a bar encodes a length, so it needs the zero and the full run beneath it, while a dot encodes a position, which is all a ranged frame can honestly carry. So change the mark when you change the frame. On the redraw the 31 models become a ladder, and the ladder is what the first chart could not show. Put the uncertainty on it while you are there, because this table mixes 44,577-drive populations with 466-drive ones: the 0.22% at the top of the league is a single failure with an interval running from 0.01% to 1.21%, and a chart that lets you rank on it is a chart that will send you shopping on noise. The 95% intervals here are ours, computed from the published failure counts and drive days; a source that prints them saves you the trouble. Print the thing that explains the spread, too. This table publishes an average-age column, the column the first chart drops, and the models at the wrong end of the ladder are mostly the old ones — so a ladder read as a vendor ranking is a ladder misread, and the operator’s own commentary about firmware, vibration and end of life is the rest of the answer. A ranged frame is not the end of the analysis; it is the thing that makes the analysis possible, and the caption should say which. Then do the multiplication the percentage was hiding. A rate in the low single digits means nothing to anybody until it is in the unit the decision is made in — failures per thousand drives a year, replacements per ten thousand bays, engineer-hours per quarter — and one line of that arithmetic beside the chart does more than any redraw. Keep the full 0-to-100 scale where the hundred is a real destination, such as a quota or a share of a vote, and where you want both, draw both: the whole unit once, and an inset at the size of the data. And do not overcorrect into refitting the axis to whatever is on screen each time, which is the auto-ranged axis and destroys comparability in the other direction. Choose the bounds once, from the question, hold them across every panel and every week, and write them down.