Imaging Endpoints Are a Measurement System | The Bracken Group

Imaging Endpoints Are a Measurement System: How to Design Them Like One

In radiopharma, where the isotope sets the clock, that judgment is the difference between a program that can prove its case and one that cannot.

Executive Summary

  • An imaging endpoint is a measurement system, not a scan. Its result depends on the interaction of acquisition, disease biology, assessment criteria, reader behavior, visit timing, data handling, and statistical assumptions.
  • When those elements are designed separately, or finalized after the protocol’s central claims are set, the endpoint can be scientifically reasonable and still be operationally unstable, statistically inefficient, or difficult to defend.
  • The fix is to treat the endpoint as a measurement system and design it upstream of protocol lock.

Session 1 of Bracken’s Radiopharma Roundtable named the underlying problem:6 radiopharmaceutical trials have grown complex enough to put themselves at risk, and imaging, properly designed, is the discipline that manages that complexity.

2 → 40Endpoints per trial. Radiopharma trials that began with two now carry 15, 20, even 40.
5–10 yrsPhase 3 timelines, against 2–3 years for comparable non-radiopharma work.
3–4Make-or-break questions imaging can answer early, before the cost of an extended Phase 1.

As characterized by the Session 1 panel: practitioner observations of the trend, not a published dataset. The broader pattern is independently documented: trial design complexity has risen sharply across the industry,7 and the number of Phase 3 radiopharmaceutical-therapy trials is climbing quickly.8

None of this principle is new. That an imaging endpoint is a system rather than a scan is the premise of the FDA’s imaging endpoint process standards1 and of two decades of quantitative-imaging work.5 Programs do not come apart on the principle; they come apart in the few places where it gets set aside to save time. This paper treats the imaging endpoint as the measurement system it is, so it can carry the clinical claim placed on it, and it covers those places in order: the claim, reliability, the estimand, acquisition, the charter and reader model, pre-launch testing, and the alignment of the full evidence package.

That first session also parked two questions for the follow-up, and both are measurement-system questions in disguise. One is whether progression-free survival can serve as a primary endpoint in radioligand therapy. The other is how to read RECIST when uptake within a single lesion is non-homogeneous. Neither is answered by picking a modality; each depends on how the endpoint is defined, how acquisition and quantification are controlled, and how readers are instructed to handle the ambiguous case. They belong upstream of protocol lock, and they are on the agenda for Session 2.

The useful question is not “can imaging support this trial?” It is “what evidence must the imaging system generate, under what conditions, and with what level of uncertainty?”

Start with the claim, not the modality

The first conversation about imaging should not be about PET versus SPECT versus MRI. It should be about the claim the endpoint is expected to support. What biological or clinical change is it meant to represent, and is that change biologically direct or a surrogate for the outcome that matters? Keep that separate from the job the endpoint does in the trial, whether primary, secondary, or exploratory; conflating the two is how a measure ends up asked to carry more than it can. At what level does the change occur, patient, lesion, organ, or whole body? And the question worth sitting with before any of it: is the expected treatment effect likely to exceed the measurement variability? Endpoints that look similar can demand very different measurement strategies. A biomarker sensitive enough for exploratory work may be nowhere near standardized enough to anchor a pivotal trial.

Clinical importance is not measurement reliability

These two get conflated constantly, and the gap between them is where trials get into trouble. An endpoint can matter enormously to patients and still be too noisy to measure reliably across a real site network. Before the protocol locks, the endpoint should be tested against the sources of variance, not only the size of the effect you hope to see. Is it reproducible across scanners, sites, readers, and timepoints? Is it robust to the disease heterogeneity you actually expect? Is it interpretable when lesions disappear, merge, split, or become non-measurable?

This is where the common instinct misleads: a nominally objective quantitative endpoint is not automatically the safer choice. It can be quietly vulnerable to segmentation, reconstruction, motion, calibration drift, partial-volume effects, or a single software-version change. A categorical reader endpoint, by contrast, can be more reliable than a quantitative one when the decision rules and case presentation are tightly controlled. Objective is not the same as reliable, and which one you have is decided by control, not by whether the output happens to be a number.

Tie the endpoint to the estimand

This decision too often gets left to the statisticians alone, and it should not be. Imaging strategy and the treatment effect being estimated are one conversation, not two.3 Intercurrent events decide which images matter: whether scans after treatment discontinuation remain relevant, how progression, treatment switching, surgery, or death are handled, whether the endpoint depends on confirmed response. These are not statistical afterthoughts to settle once the images are in hand. They determine which images must be collected, when, and how readers should interpret them. Working the estimand and the imaging plan together is what keeps the analysis from inheriting a measurement it cannot actually use.

Acquisition variability is endpoint variability

Acquisition is not an upstream quality-control matter separate from the endpoint. It is part of the endpoint’s error structure. The task is to decide which parameters must be standardized, which can be harmonized analytically, and which deviations make an assessment non-evaluable: scanner qualification and calibration, reconstruction and post-processing, contrast and uptake timing, positioning, slice thickness, motion management, and longitudinal consistency as equipment changes over a multi-year study. The standard is not “control everything.” It is sufficient control of the variables that materially move your endpoint.

Radiopharmaceutical programs feel this more acutely than any other, because the isotope sets the clock. A short half-life compresses acquisition timing, dose preparation, uptake intervals, and site readiness into windows that cannot be rescheduled, and in radioligand therapy the scan is frequently the gating step: patient selection for PSMA-targeted therapy is made by PSMA-PET, so its acquisition and quantification standards are the clinical decision about who gets treated at all.4 This is also why a radiopharma endpoint should not be lifted wholesale from an oncology protocol, a point Session 1 made bluntly. The radiation-safety limits, the decay-driven supply chain, and the patient flow are different enough that a borrowed design tends to misfit in ways that surface late. Give the field more speed and flexibility, which is the direction of travel, and the gap between getting the imaging right and getting it wrong only widens.

Radiopharma Roundtable · Session 2

Short Half-Life, Long To-Do List

The practical version of these decisions, under a short half-life and a long to-do list, is exactly what our expert panel takes on live on September 16.

Session 2 · Wednesday, September 16, 2026 · 11:00 AM ET

The charter and reader model decide the close calls

Many imaging charters are operationally complete and scientifically underdeveloped. They describe how cases move through a system without resolving how the hard cases should be read. A charter earns its keep by making the endpoint’s logic executable where reasonable experts would otherwise disagree: baseline and lesion selection, longitudinal tracking, handling of new or equivocal findings, previously treated or irradiated disease, non-evaluable anatomy, assessments that fall outside the protocol window, and what triggers adjudication. Concentrate it on the scenarios most likely to create variability, not on restating criteria that already live elsewhere.

The reader model deserves the same scrutiny, and it is too often chosen by convention: one reader for exploratory work, two plus adjudication for anything higher stakes. That is not always the most defensible design. The read model should follow the endpoint’s failure modes, expected inter-reader variability, disease complexity, the risk of reader drift over a long study, and whether adjudication resolves a real disagreement or merely supplies a third opinion. In radioligand therapy this is concrete, not hypothetical: how a reader is instructed to handle non-homogeneous uptake within a single lesion, or an early apparent progression that later resolves, can move the endpoint as much as the drug does. Oncology meets the same problem under a different name, in pseudoprogression and criteria such as iRECIST.2

Test the endpoint before the trial does

Reader training should not be the first end-to-end test of the imaging strategy. Before launch, run pilot reads on representative and deliberately ambiguous cases, simulate acquisition deviations, test the workflow across complete and incomplete timepoints, and look hard at reader agreement and adjudication frequency. The point is to learn whether the endpoint can be applied consistently under realistic conditions, while there is still time to change it. When a pilot produces frequent disagreement, the instinct is to retrain the readers. That is often the wrong move. The endpoint definition, the case presentation, the criteria, or the workflow may be what needs to change, and no amount of training fixes an endpoint that is underspecified.

Choose the imaging partner against the strategy, not by default

A core lab or imaging vendor should execute and strengthen the measurement strategy, not become its source by default. That reframes vendor selection. The question is not only which partner has the longest feature list or the prior relationship; it is which partner can be measured against the endpoint’s anticipated risks. Look for real experience with the specific disease and endpoint, scientific and medical leadership with the standing to challenge an underdeveloped endpoint, disciplined reader selection and performance management, site qualification and acquisition oversight, technology validation and change control, and transparency of operational and reader metrics. The decisive capability is the one sponsors ask about last: the ability to preserve endpoint consistency across amendments, software changes, and study expansion, which is where long radiopharma programs are most exposed.

Keep the evidence package aligned

Imaging strategy is spread across the protocol, the charter, the acquisition manual, the statistical analysis plan, the reader-training materials, and the eventual clinical study report. When an endpoint is defined one way in the protocol and another in the charter, or reader rules do not match the analysis, or adjudication outputs are not clearly represented in the dataset, the result is avoidable uncertainty about what was actually measured. That uncertainty tends to surface at the least convenient moment, when a reviewer or auditor asks how a result was produced and whether it would hold up a second time. Consistency across those documents is not administrative tidiness. It is part of whether the evidence is defensible.

Imaging strategy is measurement strategy

Imaging endpoints do not fail only because scans go missing or a site drifts from a manual. They fail when the measurement system cannot reliably support the clinical claim placed on it, and that risk is set early: when the endpoint is defined, when the estimand is fixed, when acquisition tolerances are chosen, when the reader model is structured, and when someone decides how much variability is acceptable. That is why imaging expertise belongs at the table before protocol finalization, not to operationalize a finished design, but to help decide whether the intended evidence can be generated at all. In radiopharma the margin is thinner and the clock is external, so that judgment is not a refinement. It is the difference between a program that can prove its case and one that cannot.

This is the work Bracken’s imaging and clinical-trial strategy team does with sponsors: developing fit-for-purpose imaging endpoints, assessing measurement risk, aligning protocols and charters, designing reader workflows, evaluating imaging partners, and strengthening the scientific and regulatory defensibility of imaging data. The next live conversation about it is Session 2.

Common Questions

When should imaging experts be involved in trial design?

Before protocol finalization, not after. Once the endpoint, the estimand, and the acquisition tolerances are fixed, an imaging expert can only operationalize decisions that are already made. The decisions that determine whether the endpoint can support its clinical claim are made upstream, which is where the expert input has to happen.

Why do imaging endpoints fail in clinical trials?

They rarely fail only because scans are missing or a site deviates from a manual. They fail when the measurement system cannot reliably support the clinical claim placed on it: the treatment effect is smaller than the measurement variability, the reader model does not match the endpoint’s failure modes, or acquisition is not controlled where it materially affects the result. Those are design problems, and they are cheapest to fix before the trial starts.

Why is imaging so critical in radiopharmaceutical and radioligand therapy trials?

In radioligand therapy the scan often decides who is treated, not just how disease is described. PSMA-PET, for example, selects patients for PSMA-targeted therapy, so the image is a gating step. The short half-life of many radiopharmaceuticals also compresses acquisition timing, dosing, and site readiness into narrow windows, which raises both the value of getting the imaging right and the cost of getting it wrong.

How does imaging reduce risk in early radiopharma development?

A well-designed imaging endpoint can answer the make-or-break biological questions early, before the cost of an extended Phase 1, and let an ineffective agent fail fast. That only works if the endpoint is designed as a measurement system from the start, with acquisition, reader model, and estimand aligned to the decision the trial has to make.

Radiopharma Roundtable · Session 2

Short Half-Life, Long To-Do List: Solving Early Radiopharma Development Challenges

Wednesday, September 16, 2026 · 11:00 AM ET

Sources

  1. 1.U.S. Food and Drug Administration. Clinical Trial Imaging Endpoint Process Standards: Guidance for Industry. 2018.
  2. 2.Eisenhauer EA, Therasse P, Bogaerts J, et al. New response evaluation criteria in solid tumours: revised RECIST guideline (version 1.1). Eur J Cancer. 2009;45(2):228–247.
  3. 3.ICH. E9(R1): Addendum on Estimands and Sensitivity Analysis in Clinical Trials. 2019.
  4. 4.U.S. Food and Drug Administration. Approval of Pluvicto (lutetium Lu 177 vipivotide tetraxetan), a PSMA-targeted radioligand therapy with patient selection by PSMA-PET. 2022.
  5. 5.Miller CG, Krasnow J, Schwartz LH, eds. Medical Imaging in Clinical Trials. London: Springer; 2014.
  6. 6.The Bracken Group. Radiopharma Roundtable, Session 1 recap. Recorded June 11, 2026.
  7. 7.Markey N, Howitt B, et al. Clinical trials are becoming more complex: a machine learning analysis of data from over 16,000 trials. Sci Rep. 2024;14:3514.
  8. 8.An overview of current phase 3 radiopharmaceutical therapy clinical trials. Front Med. 2025;12:1549676.
The Radiopharma Roundtable Series

Candid, expert conversations on the hardest problems in radiopharma development.

Each session convenes expert peers working through a real challenge in the field, one problem at a time. Get digital recaps, watch past sessions on-demand, and see what’s coming next.

Explore the series