MisleadingCharts
Back to the gallery

The 23 trials that never reached a journal

Showing the misleading chart

Be the first to star this exhibit

An evidence review we drew pools the published reports of trials of twelve antidepressants: 51 trials, 9,115 patients, 48 of them positive — 94%, for a pooled Hedges’s g of 0.41. Nothing is truncated, smoothed or dropped, both axes start at zero, and every trial on the page is real. But companies had to register these trials with the FDA before using them in a marketing application, so here one whole set of trials is knowable rather than merely searchable: 74 registered short-term, placebo-controlled trials at doses later approved, 12,564 patients, of which the FDA judged 38 positive — 51%. Twenty-three trials and 3,449 patients were never written up, 22 of them ones the FDA did not call positive, and eleven more reached the journals leading with an outcome that was not the registered one.

01The claim

Short-term efficacy for this class is settled. Fifty-one trials of twelve antidepressants have been published and 48 of them are positive — 94%, on an interval running from 84% to 99%. Pooled, those trials give a Hedges’s g of 0.41 (0.36 to 0.45), all twelve agents are ahead of placebo, and dropping any single trial leaves the share at 94% or better. There is no pivotal study to argue about and no axis trick anywhere in sight: the class has been tested 51 times in print, across 9,115 patients, and the answer keeps coming back the same way. Spend the next review cycle on tolerability and dosing rather than on whether the drugs work.

02The trick

Every trial on the slide is real, every effect size is the one its paper reports, both axes start at zero and nothing has been dropped — the filtering happened before the chart existed, in the decision of what got written up. What the page plots is the literature, and the literature is not the set of trials that were run; it is the set that somebody wrote up, and whether a trial was written up depended on how it came out. For these twelve agents that is checkable, which is why this is the dataset the mechanism is taught with. Companies must register with the FDA every trial they intend to use in support of a marketing application, so Turner and colleagues could read all 74 registered phase 2 and 3 trials — 12,564 patients — against the journals. The FDA judged 38 of the 74 positive and 36 not. Of those 36, three were published as not positive, 22 were never published at all, and 11 reached the journals leading with a positive secondary result while the pre-specified primary outcome — non-significant in every one of them — was subordinated in two reports and omitted in nine. Add the one positive trial nobody wrote up and the drawer holds 23 trials and 3,449 patients, 27% of everyone who took part. So the same question gets two answers: 94% of trials positive if you ask the journals, 51% if you ask the register, and the two 95% intervals — 84 to 99 and 39 to 63 — do not touch. Two things are worth being exact about. The 74 are not every trial these agents have ever been through: they are the registered phase 2 and 3 short-term, placebo-controlled trials at doses later approved, which is the set companies had to file, and Turner’s findings are restricted accordingly. And the funnel plot, the test everybody reaches for, would have almost nothing to find here, because the filter ran on the result rather than on the size — median enrolment was 153 patients among the published trials and 146 among the unpublished ones, 5% apart (P = 0.29). That last reading is ours rather than the paper’s; what the paper says about such diagnostics is that none of them is known to detect or rule out reporting bias reliably, which is the whole argument for a register over a shape test. Our IVF exhibit moves a denominator and our recall exhibit lets cohorts age unevenly; both leave a trace in the numbers on the page. This one leaves none, because the sample frame is not a mark on a chart. (This exhibit is our own demonstration in the house style of an internal evidence review, drawn from the two papers’ published figures rather than from any real review.)

03The fix

Count the drawer, then draw it. The honest chart here is the register’s: 74 marks, one per trial, in three columns for what became of them — 40 published in agreement with the FDA, 11 published against it, 23 never published — with the FDA’s own judgement as the fill, so the reader can see that 22 of the 23 missing trials are ones the regulator did not call positive. Once the drawer is on the page the arithmetic changes with it. Pooled from the FDA’s reviews the unpublished trials run at g = 0.15 (0.08–0.22) against 0.37 (0.33–0.41) for the published ones; pooling all 74 gives 0.31 (0.27–0.35) where the journals give 0.41, a 32% inflation, and the same median 32% holds agent by agent, every one of the twelve reading higher in print than in the reviews (sign test P < 0.001, increases of 11% to 69%). A trial the FDA judged positive was 11.7 times as likely to be published in agreement with that judgement as one it did not (95% CI 6.2 to 22.0). Note what does not rescue this: adding studies. A meta-analysis combines what was published, so more of the literature buys precision on a filtered sample and the interval tightens around the wrong number — which is why a register beats a better search, and why the one line worth insisting on is the frame itself. Our psilocybin exhibit ends by saying that what a two-point gap on 59 patients calls for is a bigger trial, and that is still the right answer there; this page is the caveat attached to it, because a bigger trial only helps a reader who finds out it happened. “Fifty-one published trials of 74 registered” travels; a pooled g does not. Where no register exists, and for most questions none does, say so plainly, and mark any study whose reported outcome is not its registered one, since that result usually still exists somewhere and putting it back is arithmetic rather than judgement. Two sentences keep the finding honest in the other direction, and both are the authors’ own. Each of the twelve agents, taken to meta-analysis, was still superior to placebo: what the drawer inflated was the magnitude, not the sign, and an inflated magnitude does its damage later, when the next trial is powered off the published effect and comes out too small to answer anything. And nobody knows who filled the drawer — the paper says outright that it cannot tell whether the bias came from sponsors and authors not submitting manuscripts or from editors and reviewers not accepting them. A drawer does not need a villain; it fills by default. The follow-up is the encouraging part. Turner, with a new set of co-authors, re-ran the exercise in 2022 on four agents approved between 2008 and 2013 — 30 trials, 13,747 patients, 15 positive and 15 negative — and found every positive trial transparently reported and 7 of the 15 negative ones (47%), against 4 of 37 (11%) in the older cohort as that paper classifies it, with the effect-size inflation down from 0.10 to 0.05. Registration and results reporting narrowed the drawer; they have not closed it, and no chart will ever tell you which kind of literature you are holding. Ask the register.