2026's Biggest Phase 3 Trial Failures: What They Teach About Endpoint Selection
Phase 3 is where promising clinical programmes face their most consequential test.
By this stage, sponsors have already invested years of research, substantial capital, and thousands of patient-participation hours. The expectation is that the clinical development strategy has matured enough to demonstrate whether the therapy delivers a meaningful benefit in the intended population.
Yet Phase 3 failures continue to show that a strong mechanism, encouraging Phase 2 data, or a compelling biomarker does not guarantee success.
In 2026, several high-profile Phase 3 programmes missed their primary endpoints across areas including oncology, cardiovascular disease, pulmonary disease, and rare disease. The reasons differ from programme to programme, and it would be misleading to describe every failed trial as an endpoint-selection failure. Some reflect efficacy, patient population, treatment effect size, background therapy, or other design and execution factors.
But a common lesson emerges: the endpoint must be capable of demonstrating the benefit the therapy is actually expected to deliver, in the population and timeframe in which that benefit can reasonably be detected.
This makes endpoint selection one of the most important components of pharma clinical trial planning. FDA describes endpoints as measurements of what happens to people in a clinical trial and distinguishes direct clinical outcomes from surrogate endpoints that are expected to predict clinical benefit.
For development teams, the 2026 failures provide useful case studies in what can happen when the endpoint, population, comparator, statistical assumptions, or treatment context does not align sufficiently with the underlying clinical question.
What Counts as a Phase 3 Failure?
A Phase 3 failure does not necessarily mean the drug has no biological activity.
A trial can fail its primary endpoint even when:
A numerical treatment difference exists
A secondary endpoint shows a signal
A biomarker changes
A subgroup appears to benefit
Earlier-phase studies were encouraging
The safety profile remains acceptable
The critical issue is whether the prespecified primary endpoint demonstrated the required treatment effect under the trial's statistical framework.
2026 Phase 3 Trial Failures: A Snapshot
Several late-stage failures reported in 2026 illustrate different endpoint and trial-design challenges.
Programme | Disease area | Primary endpoint issue | Key lesson |
Seralutinib | Pulmonary arterial hypertension | Primary endpoint missed | Phase 2-to-Phase 3 assumptions need validation |
Fianlimab + Libtayo | Melanoma | PFS failed to beat comparator | Comparator and treatment landscape matter |
Sigvotatug vedotin | Non-small-cell lung cancer | Overall survival endpoint missed | Biological targeting does not guarantee clinical benefit |
Anselamimab | AL amyloidosis | Overall population missed primary endpoint | Population definition can determine observed treatment effect |
Eplontersen | ATTR cardiomyopathy | Composite CV endpoint missed | Composite endpoints can dilute or complicate treatment effects |
Ziltivekimab | CKD/CV disease | Primary cardiovascular endpoint missed | Mechanistic rationale must translate into patient-relevant outcomes |
Levosimendan | PH-HFpEF | 6-minute walk distance missed | Endpoint sensitivity and population enrichment remain critical |
1. Seralutinib: When a Phase 2 Signal Does Not Translate
Gossamer Bio's Phase 3 PROSERA study of seralutinib in pulmonary arterial hypertension missed its primary endpoint in February 2026. The trial enrolled 390 patients, and the company had previously seen a modest numerical difference in six-minute walk distance in Phase 2. In Phase 3, the treatment-placebo difference increased numerically, but the study still failed its primary endpoint.
The important lesson is not simply that the endpoint was inappropriate.
The six-minute walk distance is an established functional measure in pulmonary hypertension.
The deeper question is whether the effect size expected from Phase 2 was sufficient and reproducible for a pivotal trial, particularly after differences in patient characteristics and background therapy.
This illustrates a critical principle:
A familiar endpoint is not automatically a reliable endpoint for every programme.
Development teams need to understand:
Expected treatment effect
Variability of the measurement
Placebo response
Population characteristics
Background treatment
Duration of follow-up
Clinical relevance of the expected difference
Lesson: Validate the Signal, Not Just the Endpoint
Before Phase 3, teams should ask whether the Phase 2 study demonstrated a sufficiently robust relationship between treatment exposure and the intended clinical outcome.
If the earlier signal is weak, noisy, or highly dependent on a particular subgroup, simply increasing sample size may not solve the underlying problem.
2. Fianlimab: The Comparator Can Change the Meaning of the Endpoint
Regeneron's Phase 3 trial of fianlimab plus Libtayo in first-line unresectable or metastatic melanoma missed its primary progression-free survival endpoint against Merck's Keytruda.
The combination produced a numerically longer median PFS at the higher fianlimab dose—11.5 months versus 6.4 months with Keytruda—but the difference was not statistically significant under the trial's prespecified analysis. Reports indicated that late separation of the curves contributed to the outcome.
This case demonstrates why endpoint selection cannot be separated from comparator selection.
An endpoint may be scientifically appropriate, yet the trial can still fail if the comparator is stronger than expected or the treatment effect is smaller than anticipated.
In competitive therapeutic areas, development teams should therefore evaluate:
Current standard of care
Competitor efficacy
Treatment sequencing
Patient selection
Expected comparator performance
Follow-up duration
Potential crossover effects
Lesson: An Endpoint Exists Inside a Competitive Landscape
A PFS endpoint is not evaluated in a vacuum.
Its ability to distinguish a new treatment depends partly on what the control arm achieves.
This means historical control assumptions should be challenged against the most current evidence available when planning the trial.
3. Sigvotatug Vedotin: Target Biology Must Translate Into Patient Benefit
Pfizer's Phase 3 trial of sigvotatug vedotin in previously treated advanced non-squamous NSCLC failed to meet its primary overall survival endpoint against chemotherapy in June 2026. Pfizer said the trial did not demonstrate a clear correlation between integrin beta-6 expression and patient response.
This is a useful example of the difference between biological plausibility and demonstrated clinical benefit.
A target can appear compelling.
A drug can engage the target.
A biomarker can identify patients with target expression.
Yet the ultimate clinical endpoint can still fail to show benefit.
This raises an important question for trial planning:
Is the biomarker actually predictive, or is it merely associated with the biological target?
That distinction is critical.
A predictive biomarker should help identify patients who are more likely to benefit from treatment. If target expression does not sufficiently distinguish responders from non-responders, using it as the basis for enrichment may not produce the expected clinical effect.
Lesson: Validate Biomarker-to-Outcome Relationships Early
Before using a biomarker as a central component of Phase 3 design, teams should examine evidence connecting:
Biomarker → Target engagement → Biological effect → Clinical outcome
A gap anywhere along this chain can weaken the pivotal trial.
4. Anselamimab: The Overall Population Can Hide a Treatment Signal
AstraZeneca's Phase 3 CARES programme for anselamimab in light-chain amyloidosis provides another important lesson.
The overall study population did not meet its primary endpoint, a hierarchical combination of all-cause mortality and cardiovascular hospitalisations.
However, a prespecified subgroup of patients with kappa light-chain amyloidosis showed a substantial reported treatment effect, including improved survival and fewer cardiovascular hospitalisations.
This does not convert the overall Phase 3 result into a positive trial.
But it raises an important design question:
Was the enrolled population sufficiently aligned with the biological mechanism of the therapy?
If a treatment acts on a particular disease subtype, broad enrolment can dilute the observed effect when biologically distinct populations are combined.
Lesson: Population and Endpoint Must Be Designed Together
Endpoint selection begins with understanding who is expected to benefit.
A highly responsive treatment in a biologically defined subgroup may show a smaller overall effect when that subgroup represents only part of a heterogeneous population.
That means development teams should evaluate whether:
The disease population is biologically homogeneous
Biomarker-defined subgroups are relevant
Enrichment is justified
Stratification is adequate
The primary endpoint is appropriate for all enrolled patients
5. Eplontersen: Composite Endpoints Require Careful Construction
AstraZeneca and Ionis reported in July 2026 that the Phase 3 CARDIO-TTRansform trial of eplontersen in ATTR cardiomyopathy did not meet its primary composite endpoint of cardiovascular mortality and recurrent cardiovascular clinical events through Week 140. The therapy was generally well tolerated, but the addition of eplontersen did not produce a statistically significant benefit on the composite outcome in the overall population.
Composite endpoints can be useful because they allow trials to capture multiple clinically meaningful events and potentially increase event rates.
But composites introduce complexity.
The individual components may:
Occur at different frequencies
Have different clinical importance
Respond differently to treatment
Be influenced by background therapy
If one component dominates the event count, the overall composite result may not represent what investigators initially expected.
Lesson: Understand Every Component of a Composite
Before selecting a composite primary endpoint, teams should examine:
Clinical importance of each component
Expected event frequency
Treatment effect on each component
Relative contribution to the composite
Biological rationale
Statistical weighting or hierarchy
Interpretability of the final result
6. Ziltivekimab: Mechanism Is Not the Same as Outcome
Novo Nordisk's Phase 3 ZEUS study of ziltivekimab in patients with atherosclerotic cardiovascular disease, chronic kidney disease, and inflammation missed its primary endpoint in July 2026. The trial evaluated whether monthly ziltivekimab could reduce cardiovascular outcomes in a population selected around inflammatory risk.
This provides another reminder that a mechanistically attractive hypothesis must ultimately translate into a patient-relevant outcome.
The drug targeted IL-6-related inflammation.
The trial, however, needed to demonstrate that modifying that pathway translated into fewer clinically meaningful cardiovascular events.
That distinction is central to clinical development.
Lesson: Ask What the Endpoint Actually Proves
A biomarker can demonstrate that biology has changed.
A clinical endpoint demonstrates whether patients experienced a meaningful benefit.
The two should not be treated as interchangeable.
A robust development programme should establish how early biological signals are expected to connect to later clinical outcomes.
7. Levosimendan: When a Functional Endpoint Misses
Tenax Therapeutics reported in August 2026 that its Phase 3 LEVEL trial of oral levosimendan in pulmonary hypertension associated with HFpEF failed to reach its primary endpoint of improvement in six-minute walk distance. The key secondary endpoint involving the Kansas City Cardiomyopathy Questionnaire also was not met.
The company reported a prespecified subgroup with greater disease burden showing a stronger treatment effect and highlighted supporting biomarker and pulmonary haemodynamic findings.
Again, the overall primary result remains the critical finding.
But the case illustrates the importance of matching the endpoint to the population most likely to experience measurable benefit.
A functional endpoint may be appropriate, but its sensitivity can depend on baseline disease severity.
If patients have insufficient functional limitation at baseline, there may be less room to demonstrate improvement.
Lesson: Endpoint Sensitivity Depends on the Patient
Endpoint performance is influenced by the population in which it is measured.
Development teams should consider:
Baseline severity
Ceiling and floor effects
Measurement variability
Patient heterogeneity
Natural history
Expected magnitude of change
Clinically meaningful thresholds
What These 2026 Failures Teach About Endpoint Selection
These programmes span different diseases and mechanisms.
Yet several common lessons emerge.
1. Choose the Endpoint From the Clinical Question
Start with:
What benefit should the patient experience if the drug works?
Then determine how that benefit can be measured reliably.
FDA's patient-focused drug development guidance emphasises selecting clinical outcome assessments that are fit for purpose and capable of measuring outcomes important to patients.
The process should not start with:
Which endpoint is commonly used in this disease?
Instead, ask:
Which endpoint best captures the expected treatment effect for this mechanism, population, and development context?
2. Do Not Confuse Surrogates With Clinical Benefit
Surrogate endpoints can make trials faster and more efficient when they reliably predict clinical benefit.
But their predictive value is context dependent.
FDA explicitly notes that a surrogate endpoint appropriate in one development programme may not automatically be appropriate in another because acceptability depends on factors including disease, patient population, mechanism of action, and available treatments.
This is one of the most important lessons for modern clinical development.
A validated endpoint in one context is not automatically transferable to another.
3. Validate Endpoint Sensitivity Before Phase 3
Phase 2 should not only answer whether the drug appears active.
It should help answer:
Can the planned Phase 3 endpoint detect the expected treatment effect?
This requires examining:
Effect size
Variability
Dose response
Timing
Patient characteristics
Measurement reliability
Relationship to mechanism
4. Align the Endpoint With the Mechanism
Different mechanisms may produce different types of clinical benefit.
A treatment designed to reduce inflammation may first affect biomarkers.
A neurodegenerative therapy may affect disease progression over a longer period.
An oncology treatment may delay progression before an overall survival effect becomes visible.
A symptomatic treatment may produce relatively rapid patient-reported improvements.
The endpoint must reflect the expected biology and treatment timeline.
5. Design Around the Patient Population
Endpoint selection and population selection are inseparable.
A broad population can increase recruitment and generalisability.
But it can also dilute treatment effects when the therapy benefits only a biologically defined subgroup.
Development teams should determine whether enrichment, stratification, or biomarker selection is needed before finalising the pivotal design.
6. Understand the Comparator
A treatment can fail because the control arm performs better than expected.
This is especially important in competitive therapeutic areas.
Historical control assumptions should therefore be challenged using recent evidence.
The comparator should reflect the treatment landscape that patients will actually face when the therapy reaches market.
7. Predefine How Multiple Endpoints Will Be Handled
Modern trials frequently measure multiple outcomes.
The more endpoints analysed, the more important it becomes to control multiplicity and establish a prespecified hierarchy.
FDA's multiple-endpoint guidance recommends strategies for grouping and ordering endpoints and applying statistical methods that control the chance of erroneous conclusions.
This matters when teams later point to a positive secondary endpoint after the primary endpoint has failed.
A statistically significant secondary result is not automatically sufficient to establish efficacy.
AI Can Improve Endpoint Selection
This is where AI can become useful during clinical development.
A traditional team may need to manually compare:
Previous trials
Endpoint definitions
Historical event rates
Trial populations
Comparator performance
Regulatory precedents
Published evidence
Competitor protocols
Biomarker data
An AI-enabled clinical intelligence workflow can help organise and connect this evidence.
For example, teams can investigate how an endpoint has performed across similar trials and identify patterns in:
Population selection
Event rates
Effect sizes
Trial duration
Comparator arms
Primary endpoint definitions
Secondary endpoint behaviour
This does not mean AI should choose the endpoint.
Rather, AI can help development teams make the evidence review more comprehensive.
Pienomial's clinical trial intelligence capabilities are relevant to this type of workflow, helping teams connect clinical trial information and competitive evidence when evaluating development strategies.
Using AI for Protocol Benchmarking
Endpoint selection should also be considered as part of broader protocol benchmarking.
A development team can compare a proposed protocol against relevant historical and contemporary studies.
Questions may include:
Is the proposed endpoint commonly used?
How did similar trials perform?
What event rate assumptions were made?
What patient characteristics affected outcomes?
What comparator was used?
How long were patients followed?
Were endpoints modified during development?
Did similar programmes succeed or fail?
Building Better Trial Design Lessons Into Development
The most useful output of a failed Phase 3 trial is not simply a post-mortem.
It is a reusable lesson.
Organisations can build institutional knowledge around:
Endpoint Performance
Which endpoints have historically produced reliable treatment differentiation?
Population Selection
Which patient characteristics influence treatment response?
Comparator Behaviour
How has standard of care changed?
Trial Duration
How long does the expected treatment effect take to emerge?
Effect Size
What magnitude of difference is realistic?
Biomarkers
Which biomarkers appear predictive rather than merely prognostic?
Regulatory Context
Which endpoint types have been accepted for similar programmes?
An enterprise intelligence approach can help preserve this knowledge rather than leaving it inside individual project teams.
A Practical Framework for Phase 3 Endpoint Selection
Before finalising a pivotal protocol, teams can ask seven questions.
1. Is the endpoint clinically meaningful?
Does it measure an outcome that matters to patients, clinicians, or regulators?
2. Is it sensitive to the mechanism?
Can the endpoint detect the type of benefit the treatment is expected to produce?
3. Is it appropriate for the population?
Will baseline characteristics allow meaningful change to be detected?
4. Is the timing appropriate?
Will the endpoint be assessed after enough time for the expected treatment effect to emerge?
5. Is the comparator realistic?
Are control-arm assumptions supported by current evidence?
6. Is the statistical strategy robust?
Have multiplicity, missing data, estimands, and sensitivity analyses been considered appropriately?
ICH E9(R1) emphasises the importance of defining estimands and sensitivity analyses because unclear estimands can create problems in trial design, analysis, interpretation, and decision-making.
7. Has the endpoint been challenged?
Have teams actively searched for evidence that the endpoint may fail to detect the expected treatment effect?
This final question is often overlooked.
A strong protocol review should attempt to disprove its own assumptions.
What Pharma Clinical Trial Planning Looks Like in 2026
Clinical trial planning is becoming increasingly evidence-driven.
Sponsors have access to more historical trial data, real-world evidence, scientific literature, regulatory information, and competitor intelligence than ever before.
The challenge is turning that information into useful decisions.
AI can help development teams move from:
Search → Read → Summarise
toward:
Compare → Benchmark → Challenge → Decide
This is especially useful for endpoint selection because the strongest evidence may not exist in a single document.
It may be distributed across dozens or hundreds of comparable studies.
An AI-enabled platform can help teams connect those studies and identify patterns that deserve expert review.
The Role of Pienomial in Clinical Development Intelligence
Pienomial provides AI-powered intelligence capabilities for organisations working with complex information.
Its clinical trial intelligence solution can support teams investigating clinical development activity, trial designs, endpoints, competitors, and emerging evidence.
The broader Pienomial platform is designed to help organisations connect information and create intelligence workflows rather than treating every research question as an isolated search.
For clinical development teams, that distinction can matter.
The goal is not simply to find another Phase 3 protocol.
It is to understand why similar protocols succeeded or failed and whether those lessons apply to the programme being planned today.
Conclusion
The biggest Phase 3 failures of 2026 reinforce a fundamental lesson in drug development: a trial can be scientifically ambitious and still fail if the endpoint, population, comparator, timing, or statistical framework is not sufficiently aligned with the clinical question.
Seralutinib illustrates the challenge of translating an earlier signal into a pivotal result.
Fianlimab demonstrates how comparator performance can reshape the meaning of an endpoint.
Sigvotatug vedotin highlights the difference between target biology and demonstrated clinical benefit.
Anselamimab shows how population heterogeneity can affect the observed treatment effect.
Eplontersen demonstrates the importance of understanding composite endpoints.
Ziltivekimab reinforces that mechanistic rationale must ultimately translate into patient-relevant outcomes.
Levosimendan highlights how endpoint sensitivity can depend on disease severity and population characteristics.
The lesson is not that these endpoints were universally wrong.
It is that endpoint selection is contextual.
A strong endpoint for one drug, disease, mechanism, population, or treatment setting may not be the right endpoint for another.
That is why modern pharma clinical trial planning should combine scientific expertise with systematic evidence review, protocol benchmarking, competitive intelligence, and rigorous statistical planning.
AI can strengthen this process by helping teams compare historical trials, identify relevant endpoint patterns, examine competitor protocols, and challenge assumptions before a pivotal study begins.
The most valuable AI use case is therefore not predicting whether a Phase 3 trial will succeed.
It is helping teams ask better questions before they commit patients, time, and capital to the trial design.
The best endpoint is not necessarily the most familiar one.
It is the one that gives the development team the strongest, clearest, and most clinically meaningful opportunity to demonstrate whether the treatment works.
Frequently Asked Questions
1. What are the biggest Phase 3 trial failures of 2026?
Several high-profile Phase 3 programmes missed their primary endpoints in 2026, including seralutinib in pulmonary arterial hypertension, fianlimab plus Libtayo in melanoma, sigvotatug vedotin in NSCLC, eplontersen in ATTR cardiomyopathy, ziltivekimab in cardiovascular/CKD disease, and levosimendan in PH-HFpEF. These failures had different causes and should not all be attributed to endpoint selection.
2. Why does endpoint selection matter so much in Phase 3?
The primary endpoint determines how the trial will formally assess treatment benefit. If it does not adequately capture the expected clinical effect, a biologically active therapy may fail to demonstrate efficacy.
3. What is a clinical trial endpoint?
An endpoint is a prespecified measurement used to assess an outcome in a clinical trial. It may directly measure clinical benefit or use a surrogate expected to predict clinical benefit.
4. Can AI help with endpoint selection?
AI can help teams review historical trials, compare endpoint definitions, analyse competitor protocols, identify relevant evidence, and benchmark design assumptions. Final endpoint selection should remain a scientific, statistical, clinical, and regulatory decision.
5. What is the biggest endpoint-selection lesson from 2026?
There is no universally optimal endpoint. Endpoint suitability depends on the disease, mechanism, patient population, treatment setting, comparator, expected effect, timing, and regulatory context.
6. How can teams reduce the risk of Phase 3 failure?
Teams can validate assumptions using Phase 2 data, benchmark against comparable trials, assess endpoint sensitivity, challenge population and comparator assumptions, predefine statistical strategies, and seek appropriate regulatory input before finalising the pivotal protocol.
References
[1] U.S. Food and Drug Administration (FDA). Biomarkers and Surrogate Endpoints. FDA explains the distinction between clinical outcomes and surrogate endpoints and why surrogate endpoints need evidence supporting their ability to predict clinical benefit.
FDA — Biomarkers and Surrogate Endpoints
[2] U.S. Food and Drug Administration (FDA). Multiple Endpoints in Clinical Trials. The guidance discusses grouping, ordering, and statistically evaluating multiple endpoints to reduce the risk of erroneous conclusions.
FDA — Multiple Endpoints in Clinical Trials
[3] U.S. Food and Drug Administration (FDA). Table of Surrogate Endpoints That Were the Basis of Drug Approval or Licensure. FDA notes that the appropriateness of surrogate endpoints is context dependent and should be considered with the relevant review division.
FDA — Surrogate Endpoint Table
[4] U.S. Food and Drug Administration (FDA). Patient-Focused Drug Development: Selecting, Developing, or Modifying Fit-for-Purpose Clinical Outcome Assessments. The guidance addresses the selection and development of clinical outcome assessments for measuring outcomes important to patients.
FDA — Fit-for-Purpose Clinical Outcome Assessments
[5] U.S. Food and Drug Administration (FDA). E6(R3) Good Clinical Practice (GCP). The 2025 final guidance incorporates risk-based approaches and modern trial design and technology while emphasising quality by design and reliable trial results.
FDA — E6(R3) Good Clinical Practice
2026 Trial Sources
The 2026 trial examples above were checked against sponsor disclosures and contemporary reporting, including AstraZeneca's disclosures for CARDIO-TTRansform and CARES, and reporting on the seralutinib, fianlimab, sigvotatug vedotin, ziltivekimab, and levosimendan programmes.






