Statistical vs. Clinical Significance: What Boards Misread in Trial Results
A clinical trial can succeed on its own terms and still leave a company with an asset nobody wants. The result clears the pre-specified threshold, the press release goes out, the endpoint was met — and eighteen months later the program is struggling to find a partner, or the label is too narrow to build a business on, or a payer conversation goes badly in a way nobody on the board anticipated.
The gap between statistically significant and clinically meaningful is where a great deal of biopharma value quietly disappears. It disappears with the board's consent, because the board was told the trial succeeded and had no framework for asking the second question. This article is about that second question — what it is, why it is routinely skipped, and how a director without statistical training can ask it well.
The point is not that boards should become statisticians. It is that a small number of conceptual distinctions, held clearly, will let a director hear a positive trial readout and know whether it is the kind of positive that builds a company.
The Distinction, Stated Plainly
Statistical significance answers a narrow question: how likely is it that the difference we observed between treatment and control arose by chance? A p-value below the pre-specified threshold says the effect is probably real. That is all it says. It says nothing whatsoever about whether the effect is large, whether it matters to a patient, or whether anyone would pay for it.
Clinical significance answers a different and much harder question: is the size of this effect enough to change how a physician treats a patient, enough for a patient to notice and value, and enough for a payer to reimburse against the alternatives already available?
These two questions can diverge sharply, and the direction of the divergence matters. A sufficiently large trial can detect a genuine but trivially small effect with great statistical confidence. Sample size is, in effect, a lever that converts small real differences into significant p-values. This is not a flaw in the statistics; it is what the statistics are designed to do. But it means "statistically significant" carries no information about magnitude, and a board that hears the phrase as a synonym for "the drug works well" has misunderstood the readout.
The reverse divergence is also real and matters for a different reason. A trial can show a clinically meaningful effect size that misses statistical significance because the study was underpowered — too few patients, higher-than-expected variability, worse-than-modeled dropout. A board that reads this as simple failure may be discarding a genuine signal. A board that reads it as vindication may be rationalizing. Distinguishing the two is one of the harder judgments in this domain, and it turns almost entirely on whether the effect size seen is consistent with what was assumed in the original power calculation, or whether it is being discovered after the fact.
Where the Misreading Happens
A few specific patterns account for most of the value destroyed in this territory. They recur across companies and therapeutic areas, and a board that recognizes them will catch most of what matters.
The endpoint was met, but the effect size is below what the market requires. Every therapeutic area has a rough threshold — sometimes formalized as a minimum clinically important difference, sometimes just understood by practitioners — below which a change in the measured outcome does not alter clinical practice. A result can be statistically robust and sit below that threshold. The trial succeeded; the drug is not commercially interesting. The board should know, ideally before the trial reads out, what effect size the market actually requires, and how that compares to what the trial was powered to detect. When the trial is powered to detect something smaller than the market needs, the company has designed a study it can win without winning anything.
The comparator was not the real-world alternative. A trial run against placebo, or against an older standard of care, can produce a clean statistically significant result that tells you very little about how the drug performs against what physicians are actually prescribing now — or will be prescribing by the time the drug launches. This is a particularly costly form of the gap, because the trial is legitimately successful and the problem only surfaces in payer and partner conversations later. The board's question is not "did we beat the comparator" but "is the comparator the thing we will be competing against."
The significant result is on a secondary endpoint or a subgroup. When the primary endpoint misses and a secondary endpoint or a subgroup analysis is significant, the enthusiasm in the room is often out of proportion to the evidentiary weight. Pre-specification is the dividing line that matters here. A pre-specified secondary endpoint, with the multiplicity of testing properly accounted for, carries real weight. A subgroup identified after seeing the data does not — it is hypothesis-generating, and treating it as confirmatory is one of the most reliable ways for a company to spend its remaining capital chasing a result that will not replicate. The board's question is blunt and diagnostic: was this analysis specified before we saw the data, and was multiplicity handled? If the answer is no or unclear, the finding is a reason to design the next study, not a reason to fund a pivotal one.
The surrogate endpoint moved, but nobody knows if the outcome will. Many trials measure something that stands in for the thing patients actually care about — a biomarker, an imaging measure, a physiological parameter. Some surrogates are well validated and regulators accept them; others are plausible but unproven. A statistically significant movement in a poorly validated surrogate is a much weaker result than the headline suggests, and the risk lands squarely on the company: the confirmatory outcome study may not follow. The board should always know which category the endpoint sits in, and it should be suspicious when the answer is delivered with more confidence than the evidence supports.
Statistical significance is being used to close a conversation rather than open one. This is the meta-pattern. When "we hit the endpoint" is offered as the complete answer to "how did the trial go," the board is being handed a conclusion instead of a result. It is almost never a deliberate deception. It is a natural compression of a complicated readout into the one fact everyone has been waiting for. But it is the board's job to decompress it.
What the Board Should Ask on Any Readout
None of the following requires statistical training. All of them require the discipline to ask before the celebration has fully set in.
How big was the effect, in units a clinician would recognize? Not the p-value, not the hazard ratio in isolation — the actual magnitude, expressed in the terms the treating physician thinks in. If the answer is hard to give plainly, that difficulty is itself informative.
How does that compare to the threshold that changes practice? And who established that threshold — published consensus, regulatory precedent, our own commercial assumption? The provenance matters, because a company-defined threshold set after the data arrived is not a threshold.
How wide is the confidence interval? A point estimate that looks impressive with an interval running from "barely anything" to "enormous" is a much weaker result than the point estimate alone conveys. The lower bound of the interval is often the more decision-relevant number, because it represents the pessimistic-but-still-consistent-with-the-data case. Boards that learn to look at the bottom of the range rather than the middle of it will make better calls.
Was this analysis pre-specified? Asked of every result being used to justify a decision. It is the single highest-yield question in this entire domain.
What did the safety and tolerability data show alongside the efficacy? Clinical meaningfulness is a net concept. A modest efficacy benefit accompanied by meaningful tolerability problems can be worth less than the efficacy number alone suggests, particularly in indications where alternatives are well tolerated. Efficacy and safety get presented in separate sections of the deck; they have to be evaluated together.
Would this result support the label we need, and the price we assumed? This is the question that connects the statistics to the business, and it is the one most often left un-asked in the readout meeting because it feels like a commercial question for a different agenda item. It is not. It is the question that determines whether the trial result creates value.
The Governance Structure Around It
Two structural habits do most of the work here, and both have to be established before the data arrive.
The first is pre-committing, at the board level, to what a meaningful result looks like. Not just the statistical success criterion — the trial protocol handles that — but the commercial and clinical bar: the effect size that would make this asset worth advancing, worth partnering, worth the next round. Setting that bar in advance costs a single agenda slot and is worth more than any amount of post-hoc analysis, because it gives the board a standard that was fixed when nobody had a stake in where it sat. When the readout arrives and the effect lands below the bar, the conversation is about a known standard rather than a contested interpretation.
The second is ensuring the board has access to a statistical or clinical perspective that is not the team presenting the results. This need not be elaborate — an independent advisor available for the pivotal readouts is usually sufficient. The purpose is not distrust. It is that the people who designed and ran the study are, entirely understandably, the people least able to see it coldly. On the readouts that determine the company's direction, the board should be able to get a second read from someone with no stake in the interpretation.
The Underlying Principle
The p-value is a gatekeeper, not a verdict. It tells the board that an observed effect is probably not noise. Everything the board actually cares about — whether the effect is large enough to change practice, whether the label it supports is a label worth having, whether a partner or a payer will value it — lies on the other side of that gate and requires separate examination.
Boards get into trouble here not through carelessness but through a reasonable-seeming shortcut: the trial succeeded, so the program advances. The discipline that prevents it is unglamorous. Set the meaningfulness bar before the data arrive. Ask for magnitude, not just significance. Ask what was pre-specified. And treat a positive readout as the beginning of the evaluation rather than the end of it.
Lawrence Fine is CEO of AGCP Farmacêuticos and has advised on licensing, regulatory, and partnership strategy across the pharmaceutical and advanced materials sectors.