Citation
Shalhout SZ, Bloom R, Drake L, Miller DM. Evaluation of the fragility of pivotal trials used to support US Food and Drug Administration approval for plaque psoriasis. J Am Acad Dermatol. 2021;84(2):354–360. doi:10.1016/j.jaad.2020.04.057.
Why this study matters
Statistical significance is central to the pivotal trials used to establish efficacy for new therapies, but a significant result does not by itself reveal how readily that conclusion would change if only a small number of participant outcomes differed.
This study applied the fragility index and fragility quotient to pivotal efficacy endpoints supporting FDA approvals for systemic plaque psoriasis therapies, asking whether the evidence base was statistically delicate or unusually robust.
Pivotal efficacy trials supporting FDA approvals in plaque psoriasis were substantially more robust than randomized trials reported across many other clinical fields. The strength of that evidence raises a different question: whether some future development programs could conserve patients and resources without sacrificing statistical persuasiveness.
The analysis at a glance
Non-biosimilar therapies
Applications with reviewable pivotal efficacy data and eligible dichotomous endpoints were included.
Efficacy trials
Each approval program included at least two pivotal trials, with several including three.
Primary or co-primary endpoints
Endpoints were limited to two-arm, dichotomous comparisons suitable for fragility analysis.
Participants in primary-analysis arms
This excluded participants enrolled only in dose-ranging or secondary-analysis arms.
What the fragility measures mean
Fragility index
The minimum number of participant outcomes that would need to change for a statistically significant trial result to become nonsignificant.
Fragility quotient
The fragility index divided by the number of participants in the analyzed trial arms, allowing robustness to be interpreted relative to sample size.
Key findings
Highly robust primary endpoints
A median of 72 outcome changes would have been required to overturn statistical significance.
Robustness after accounting for size
The strength of the results was not explained solely by large trial enrollment.
Median fragility index
Trials supporting non-psoriasis indications for the same therapies had a substantially lower median index.
Median across 3,632 prior trials
A review of 49 reports found much lower fragility indices across the broader clinical-trial literature.
Unlike many fragility analyses that identify statistically delicate evidence, this study found the opposite: pivotal trials in plaque psoriasis were exceptionally resistant to small changes in participant outcomes.
Why psoriasis trials may be so robust
- Modern targeted therapies produce large differences in response compared with placebo.
- Late-phase development benefits from efficacy information accumulated in earlier studies.
- Approval programs commonly included large samples and multiple pivotal trials.
- PASI and Physician Global Assessment endpoints often showed strongly concordant treatment effects.
- Family-wise error control for co-primary endpoints may contribute to larger sample-size requirements.
Implications for trial design
Exceptional statistical robustness is reassuring, but it may also indicate that some trials enrolled more participants than were needed solely to establish efficacy. A post hoc analysis suggested that reducing sample sizes by half—while preserving observed event rates—would still have yielded a median fragility index of 33, with no development program losing statistical significance.
Safety characterization remains an important reason to enroll large populations. Still, the findings support examining whether efficacy programs could be designed more efficiently, especially when treatment effects are large and earlier studies already provide strong evidence of activity.
Robustness should not be reduced merely for its own sake. The relevant question is whether trial design can preserve persuasive evidence and adequate safety assessment while reducing unnecessary exposure, cost, and operational burden.
Potential design opportunities
- Reconsider whether PASI and Physician Global Assessment must always serve as co-primary endpoints in late-phase trials.
- Examine whether a single well-validated primary endpoint could reduce alpha splitting and sample-size inflation.
- Separate efficacy sample-size requirements from the population needed for adequate safety characterization.
- Use active comparators when they address a meaningful scientific, clinical, ethical, or market-access question.
- Align enrollment more closely with the minimum evidence needed to demonstrate clinically meaningful benefit.
Interpretive considerations
- Fragility measures apply only to eligible dichotomous outcomes and do not capture every dimension of trial quality.
- There is no universally accepted threshold defining an adequately robust fragility index.
- Low fragility is not necessarily evidence of a poorly designed trial, because efficient studies are often intentionally powered near the minimum required sample size.
- Several older approval programs lacked sufficient publicly available efficacy or statistical-analysis data.
- Cross-disease comparisons require caution because effect sizes, disease severity, endpoints, and approval standards differ.
The pivotal evidence supporting systemic plaque psoriasis therapies was statistically persuasive. Future regulatory science should focus not only on detecting fragile evidence, but also on identifying when exceptionally robust programs create opportunities for more efficient and patient-conscious trial design.
Authors
Sophia Z. Shalhout, Romi Bloom, Lynn Drake, and David M. Miller.