The argument in brief
- SUV is the default PET metric, but it stays highly variable (within-subject CV ~10–12%, repeatability ±28–39%) even after harmonization.
- Harmonization frameworks (EARL, QIBA, UPICT) are complex and costly, and add little beyond expert nuclear medicine reads.
- In Phase 0 trials, SUV should not be a primary endpoint. Qualitative uptake, volumetric measures, and kinetic modeling fit the exploratory aims better.
Abstract
- Rationale
- Standardized uptake values (SUVs) are widely used in PET quantitation but remain highly variable due to technical, biological, and interpretive factors. In multicenter trials, harmonization frameworks (EARL, QIBA, UPICT) aim to reduce variability, yet require complex calibration and quality control. The utility of SUVs as primary endpoints in Phase 0 clinical trials is uncertain.
- Methods
- We reviewed publications from 2008–2023 using PubMed and Scopus with the terms PET, SUV harmonization, EARL, QIBA, and Phase 0 clinical trial. Foundational guidelines (QIBA FDG-PET/CT Profile, UPICT protocol, EANM/EARL standards) and recent multicenter trials were critically appraised.
- Results
- SUV repeatability in multicenter trials is limited, with within-subject CV of ~10–12% and repeatability coefficients of ±28–39%. Harmonization requires cross-calibration, phantom benchmarking, uptake-time enforcement, and strict reconstruction protocols. Despite this burden, SUVs add little beyond trained nuclear medicine reads. In lymphoma, Deauville scoring remains standard, and visual assessment often outperforms strict quantitative thresholds.
- Conclusions
- SUVs should not serve as primary endpoints in Phase 0 clinical trials. Qualitative uptake, volumetric measures, and kinetic PET modeling better align with the exploratory, mechanistic aims of Phase 0 studies.
Keywords: Positron Emission Tomography; Phase 0 Clinical Trials; Standardized Uptake Value; Harmonization; Nuclear Medicine
Introduction
Phase 0 clinical trials, or exploratory IND studies, are designed to provide early insights into drug pharmacodynamics and target engagement. PET imaging is particularly suited for these aims, providing noninvasive readouts of tracer biodistribution and target binding. Standardized uptake values (SUVs) have become the default PET metric, but they are confounded by multiple sources of variability; even a patient's measured weight can vary between scanners and on different days. While harmonization frameworks attempt to standardize SUV measurements across centers and within institutions, it is hard to justify this complexity in a Phase 0 setting. This whitepaper reviews SUV harmonization, its limitations, and the case for non-SUV endpoints in Phase 0 trials.
Methods
We conducted a narrative review of peer-reviewed publications between 2008 and 2023 using PubMed and Scopus. Search terms included "PET," "SUV harmonization," "EARL," "QIBA," and "Phase 0 clinical trial." We also reviewed foundational technical documents such as the QIBA FDG-PET/CT Profile, UPICT guidelines, and EANM/EARL accreditation protocols. Articles were included if they reported on SUV variability, harmonization methods, or clinical trial applications. Key findings were synthesized into themes regarding sources of variability, harmonization workflows, and comparisons with expert nuclear medicine interpretation.
Results
SUV variability arises from technical, biological, and interpretive factors. Harmonization strategies such as EARL and QIBA require:
- Dose calibrator to scanner cross-calibration
- Phantom recovery-coefficient benchmarking
- Uptake-time enforcement (±10 minutes)
- Reconstruction and smoothing standardization
- Ongoing quality control after scanner upgrades
Biological factors compound the problem: body weight itself shifts with season and, in patients with cancer, with disease stage, sarcopenia, and chemotherapy, all of which change the weight and muscle mass that SUV normalization depends on.
Despite these efforts, within-subject coefficient of variation remains ~10–12% and repeatability limits extend to ±28–39%. Visual interpretation methods, such as Deauville scoring in lymphoma, continue to outperform or complement quantitative thresholds.
Despite the burden of harmonization, SUVs add little beyond a trained nuclear medicine read.
Figures
- Dose calibrator to scanner cross-calibration
- Phantom QC (recovery coefficients)
- Patient prep & uptake-time enforcement
- Reconstruction alignment / smoothing
- Ongoing QC & software version control
Discussion
This review demonstrates that while harmonization improves SUV reproducibility in multicenter trials, the complexity and cost are substantial. For Phase 0 trials, where small cohorts and exploratory objectives dominate, SUVs are poorly suited as primary endpoints. Non-SUV endpoints, such as qualitative uptake, volumetric measures, and kinetic modeling, provide more robust mechanistic insights with fewer operational constraints. Expert nuclear medicine readers often outperform rigid SUV thresholds, reinforcing the limited incremental value of SUVs in this context.
Conclusion
SUV harmonization reduces variability but remains fragile and resource-intensive. In Phase 0 clinical trials, SUVs should not be used as primary endpoints. Non-SUV PET endpoints better align with the mechanistic goals of early-phase drug development, supporting exploratory science without the burden of harmonization frameworks.
Frequently asked questions
Should SUV be a primary endpoint in Phase 0 PET trials?
No. Even after harmonization, within-subject coefficient of variation remains roughly 10–12% with repeatability limits of ±28–39%. Given the small cohorts and exploratory, mechanistic aims of Phase 0 studies, that variability and the operational burden are hard to justify. Qualitative uptake, volumetric measures, and kinetic modeling align better with the science.
Does EARL or QIBA harmonization make SUV reliable enough for early-phase decisions?
Harmonization improves reproducibility but does not remove the core limitations. It requires cross-calibration, phantom benchmarking, uptake-time enforcement, strict reconstruction protocols, and ongoing quality control, and meaningful residual variability still persists. In practice, SUV adds little beyond a trained nuclear medicine read, and visual methods such as Deauville scoring often outperform strict quantitative thresholds.
What PET endpoints fit Phase 0 objectives better than SUV?
Qualitative uptake assessment, volumetric measures, and kinetic PET modeling. These match the target-engagement and pharmacodynamic questions Phase 0 studies are designed to answer, with fewer harmonization constraints and less sensitivity to the technical and biological noise that degrades SUV.
References
- Graham MM, Wahl RL, Hoffman JM, et al. Summary of the UPICT Protocol for 18F-FDG PET/CT Imaging in Oncology Clinical Trials. J Nucl Med. 2015;56(6):955-961.
- Kinahan PE, Fletcher JW. PET/CT Standardized Uptake Values in Clinical Practice and Assessing Response to Therapy. Semin Ultrasound CT MR. 2010;31(6):496-505.
- QIBA FDG-PET/CT Profile. Quantitative Imaging Biomarkers Alliance (QIBA), RSNA; 2013.
- Kaalep A, Sera T, Rijnsdorp S, et al. Quantitative implications of the updated EARL (2019) accreditation. EJNMMI Phys. 2019;6:28.
- Hicks RJ, Aide N. Harmonization strategies in PET quantification: from daily practice to multicenter trials. J Nucl Med. 2016;57(Suppl 2):15S-22S.
- Boellaard R, Delgado-Bolton R, Oyen WJ, et al. FDG PET/CT: EANM procedure guidelines for tumor imaging: version 2.0. Eur J Nucl Med Mol Imaging. 2015;42(2):328-354.
- Wahl RL, Jacene H, Kasamon Y, Lodge MA. From RECIST to PERCIST: evolving considerations for PET response criteria in solid tumors. J Nucl Med. 2009;50(Suppl 1):122S-150S.


