The zero readings at 139, and the 174 from the same visit
Showing the misleading chart
A clinic briefing we drew plots 19,959 systolic blood pressure readings from NHANES, one bar per whole millimetre of mercury, on a linear axis from zero with nothing smoothed or binned — and the distribution comes out a comb. There are no readings at 139, none at 137 or 141, none at any odd millimetre in the file. The desk reads the empty column beneath the 140 threshold as a fact about patients and recommends holding the cut-point there. It is a fact about the mercury column, which is graduated every 2 mmHg; NHANES measured 5,972 of the same participants again at the same visit on an oscillometric device, and recorded 174 readings at exactly 139.
01The claim
Nobody in America has a blood pressure of 139. Every reading on the slide is NHANES’s own — one national survey, one protocol, one field — 19,959 systolic measurements taken by certified physician examiners on 6,717 participants aged 8 and over. Nothing is smoothed, binned, indexed or rescaled: the count axis is linear and starts at zero, and there is one bar for every whole millimetre of mercury from 80 to 200. Read at that resolution the distribution turns out to be a comb. Systolic pressure in this population does not take odd values — not one reading in 19,959 — and the empty column sits directly beneath the threshold the clinic has been arguing about. Read-out for the desk: the distribution has a wall in it. The column at 139 is not low, it is empty, and so is every other odd column across the whole range, so 140 is not a committee’s round number, it is where the data itself resumes. Recommendation: hold the threshold at 140 and close the debate; treat 1 mmHg differences as noise, since pressure does not take those values; and drop the odd integers from the clinic’s index.
02The trick
The count axis runs from zero, nothing is indexed, rebased, logged or smoothed, the bars are one width and one colour, and nothing inside the plotted range has been dropped — the 73 readings that fall outside 80–200 are noted on the panel itself. The drawing is not doing anything. What the chart has drawn, at one bar per whole millimetre, is the mercury column. A mercury sphygmomanometer’s column is graduated every 2 mmHg — that is what the instrument is — and the auscultatory convention is to read it to the nearest mark. NHANES’s own data-processing note for the file records the consequence as an edit rule: systolic and diastolic measurements and the maximum inflation level “can be even numbers only”. An odd reading has no mark to be read from and no field to be stored in. So the comb is a property of the recording system, and every column on it that stands empty is a value the instrument and the form between them cannot express. The scale of it is worth saying plainly: between the lowest reading in the file, 72, and the highest, 238, there are 167 whole millimetres, and 85 of them have no reading at all — 83 because they are odd, and two more, 226 and 230, that simply had nobody at them. It is not a fluke of one cycle either. Put 2013–2014, 2015–2016 and 2017–2018 together and the auscultatory files hold 64,521 systolic readings without a single odd number among them; the diastolic column is the same, 19,959 of 19,959 even in 2017–2018. This is the reverse of the usual version of the trick, and it is worth keeping the two the same way up in your head, because they are one mechanism. In routine clinical practice the well-documented failure is terminal digit preference: a hurried reading gets rounded to a comfortable number and the histogram grows spikes at 120, 130 and 140. NHANES’s examiners are certified against an audio-video standard that requires them to be within ±2 mmHg on 92% of its 24 test measures, so they do not do that, and what shows through instead is the graduation of the instrument itself. Spikes where a person rounds, holes where a scale cannot reach: in both cases the fine structure of the distribution is a picture of how the number was written down. What makes this one worth a slide rather than a footnote is that the desk drew a conclusion from the holes. A threshold is a place where a small difference in the number becomes a large difference in what happens to the patient, and the chart appears to show the population itself declining to stand near this one. It shows nothing of the sort. NHANES ran a blood pressure methodology study in the same 2017–2018 cycle: 5,972 of the same 6,717 participants, at the same visit, order assigned at random, measured a second time by a health technician with an Omron HEM–907XL oscillometric device. That device recorded 174 readings at exactly 139, 9,163 odd readings in all, and 49.91% of its 18,360 readings ending in an odd digit — which is what a scale with odd numbers on it looks like. Pressures at and around 139 are perfectly ordinary, then, and a scale that can express them finds them. It is worth being careful about what that does and does not say, because the tempting sentence — that the people at 139 were there all along and the mercury protocol wrote each of them down as 138 or 140 — is not one the data supports, and it is the same error as the slide’s in a nicer suit. Blood pressure moves minute to minute, and these are two different instruments, read by two different people, at two points in a long visit. The 163 participants who registered a 139 on the device are not 163 people whose mercury reading should have been 139: their 468 mercury readings run from 108 to 180, and only 15.4% of them are 138 or 140. Pair the first reading from each protocol across the 5,601 participants who have both and the mean difference is 1.24 mmHg — but the standard deviation is 10.74, and only 21.4% of pairs agree within 2 mmHg. The offset is the small part and the spread is the large one; a mean difference quoted without its spread is its own small chart crime, and the two protocols are not interchangeable at the level of a person. (On the cleanest cohort — the 5,367 participants with all three readings on both — the mercury protocol reads 1.14 mmHg higher on average, a device difference NCHS examines in its own methodology report. It is a separate finding, and it is not what put the gaps in the first chart.) What the device does establish is the population fact the empty column denied: the values are reachable, and a scale that can reach them lands on them about as often as on their neighbours. And the lattice’s consequence at the threshold does not need any person-level matching at all, because it is arithmetic on one file. An odd reading sits exactly halfway between two marks, so a rounding convention decides it, and the convention is load-bearing: round every odd oscillometric reading down and 2,922 readings, 15.92%, sit at or above 140; round them up and it is 3,096, or 16.86%. Same readings, rounded two ways — 174 of them, and 0.95 points of the share above the line, moved by a convention, with the mean shifting 0.998 mmHg between the two. (Those are shares of readings, not a hypertension prevalence, which would need per-person means, medication status and survey weights. Nor does any of this bear on where the line should sit; 140 against 130 is a live argument among people who treat patients, and a chart about rounding has nothing to contribute to it.) The lattice has one more trap in it, which survives even after you stop reading the comb as biology. Bin a 2 mmHg lattice into 5 mmHg bins and the bins hold three reachable values, then two, then three, alternating forever — so the bars alternate high and low by construction, whatever the data does. Above 120 mmHg the mercury series rises four separate times as the pressure climbs, at 125→130, 135→140, 145→150 and 165→170, and is exactly flat once, at 315 readings in both 155–160 and 160–165. The oscillometric readings from the same visits fall at every single step from 120 to 185. A rising bar in the right-hand tail of a blood pressure distribution is the sort of thing that gets a second look and a hypothesis; here it is the bin width. (NHANES publishes the readings; both drawings are ours, and the desk, the read-out and the recommendation on the first are invented.)
03The fix
Draw at the resolution the measurement actually has. One bar per 2 mmHg is one bar per value the instrument can produce, and the redraw does nothing cleverer than that: the comb becomes a single ordinary unimodal distribution, nothing is dropped, and the axis simply stops asserting a precision that was never recorded. Over it runs the oscillometric line at 1 mmHg — the same survey and the same visits, 5,972 of the same participants — which fills in the odd numbers and shows the shape was right all along. Beneath, two panels answer the two questions the slide got wrong. The first plots the last digit of every reading, ten bars per protocol against a reference line at 10%: the mercury readings put 18.5, 17.8, 22.5, 22.3 and 18.9 per cent on the five even digits and exactly nothing on the five odd ones, while the oscillometric readings sit between 9.34 and 10.44 on all ten. Ten bars is the cheapest diagnostic in statistics and it settles in one glance what a paragraph of caveats cannot. The second draws the upper tail in 5 mmHg bins with the number of reachable values printed under each bin, so the sawtooth and its cause are on the same picture. In general, ask what the measurement was recorded to the nearest of before you draw it, and print the answer beside the unit the way you print the unit: “mmHg, to the nearest 2”, “years, as reported”, “kilograms, self-reported”. It is four words and it is part of what the number means. Choose bin widths that are whole multiples of the recording step — 2 or 10 here, never 5 — so that every bin holds the same number of reachable values and the alternation cannot appear. Treat any threshold that lands on or beside a favoured value as a decision about rounding as much as about medicine, and say which way the rule rounds, because everything within half a step of the line moves on that answer. And where a second, finer measurement of the same thing exists — a scale beside the self-report, an automatic device beside the manual one — draw both, because the disagreement between two instruments is the most honest error bar anyone is going to hand you. The same trap is waiting wherever a person rounds or an instrument steps, and it is nearly always in the data before it is in the chart: census ages ending in 0 and 5, which demographers have scored with Whipple’s index since 1919; self-reported weights at multiples of five; prices ending in 9; species occurrence records sitting on whole-degree coordinates; survey durations at round minutes; star ratings, which have five reachable values and get drawn to two decimal places. The spike and the hole are the same mechanism seen from two sides, and the rule for both is the same: a distribution is a statement about the recording system first and about the population second, so check the form the number was written on before reading anything off its shape.