Patient-reported tolerability: insights from a secondary analysis of the LIBRETTO-531 study
The management of advanced RET-mutant medullary thyroid cancer (MTC) has been transformed by the advent of selective RET inhibition. The LIBRETTO-531 trial established the superiority of selpercatinib over cabozantinib/vandetanib in progression-free survival (PFS) [hazard ratio (HR) 0.28], treatment failure-free survival (HR 0.25), and overall response rate (69.4% vs. 38.8%), with a more favorable safety profile (1). Elisei et al. extended these findings by reporting a novel patient-reported tolerability (PRT) endpoint which describes the proportion of time on treatment (PTT) with “high side-effect burden”. This metric suggested that patients receiving selpercatinib spent significantly less time burdened by treatment toxicity (8% vs. 24%, P=0.0001) (2). This secondary analysis represents a meaningful example of how the patient experience is quantified in oncology clinical trials, and it also raises important methodological questions.
Simplicity as both strength and limitation
The PRT metric is built upon a single item from the Functional Assessment of Cancer Therapy-General (FACT-G) Physical Well-Being scale: GP5, “I am bothered by side effects of treatment”, rated on a 5-point Likert scale from 0 (“not at all”) to 4 (“very much”) (3). The appeal of GP5 is because of simplicity: a single question imposes minimal respondent burden, facilitates high compliance, and can be administered weekly without the fatigue that often accompanies longer instruments. The FACT-G itself was designed with a 7-day recall period, and the weekly administration of GP5 in LIBRETTO-531 aligns well with this window, decreasing some concerns about recall bias that can challenge instruments administered at longer intervals (4).
The validity evidence supporting GP5 as a tolerability indicator has grown substantially. Pearman et al. demonstrated significant associations between GP5 scores and clinician-rated adverse event grade across multiple cancer types (3). Peipert et al. showed that “high” GP5 scores [3–4] were associated with 2.2- to 4.7-fold greater odds of early treatment discontinuation due to adverse events in the ENDURANCE myeloma trial (5). Griffiths et al. provided real-world validation across 6,755 patients with 10 cancer types in 6 countries (6). Importantly, Payakachat et al. conducted qualitative content validity interviews specifically in MTC patients enrolled in LIBRETTO-531, confirming that the concepts of side-effect bother, burden, and tolerability were “highly relevant and related” and that the dichotomization threshold of scores 3–4 as “high side-effect burden” was endorsed by 60% of participants (7).
Nonetheless, can a single item instrument adequately capture the multidimensional experience of treatment toxicity? The GP5 collapses the heterogeneous landscape of adverse effects into a single ordinal response, sacrificing granularity. The complementary PRO-CTCAE symptomatic adverse event data reported by Elisei et al. partially address this concern, revealing that the tolerability advantage of selpercatinib was driven by substantially less time with severe diarrhea (5% vs. 38%), fatigue (6% vs. 21%), taste change (3% vs. 15%), decreased appetite (2% vs. 15%), and hand-foot syndrome (2% vs. 9%) (2). These symptom-level data are arguably as clinically informative as the composite GP5 metric, and future analyses would benefit from more granular exploration of which specific toxicities most strongly drive the GP5 signal. The preliminary I-SPY2 breast cancer experience (see referenced abstract) offers a useful model as higher-grade abdominal pain, decreased appetite, dizziness, and blurry vision were significantly associated with higher GP5 scores, while nail loss, acne, hair loss, and rash were not. This may demonstrate that not all toxicities contribute equally to patient-reported bother (8).
Score distribution and the question of floor effects
A critical methodological consideration is the distribution of GP5 responses. The FACT-G Physical Well-Being scale, from which GP5 is drawn, characteristically exhibits left-skewed distributions in clinical trial populations, with the majority of responses clustering at the lower end of the scale (scores 0–2) (3,4). This distributional characteristic has implications for the PRT metric. If the majority of patients in both arms report scores of 0–2 at most timepoints, the PTT with “high side-effect burden” will be low in both groups, and the absolute difference—while statistically significant—may reflect the experience of a relatively small proportion of patients at any given assessment. The 8% versus 24% PTT difference reported by Elisei et al. is clinically meaningful, but readers should consider what this means for the patients: even in the control arm, patients spent roughly three-quarters of their treatment time reporting low side-effect burden (2). Whether this reflects genuine tolerability, adaptation, response shift, or a flaw in a single item instrument remains an open question.
Compliance and missingness
The authors report post-baseline compliance rates “generally greater than 80%” for PRO questionnaires in both treatment groups (2). This exceeds the benchmarks reported in a recent scoping review by Krepper et al., who found mean baseline PRO completion rates of approximately 92% declining to 82% at first post-baseline assessment across 222 oncology randomized controlled trials (RCTs), with open-label trials exhibiting lower completion than double-blind trials (9). The sustained compliance in LIBRETTO-531 strengthens confidence in the representativeness of the data and reduces concerns about data missingness.
The pattern of missingness deserves a call out. Arizmendi et al. demonstrated that while the responsiveness of GP5 to side effects and CTCAE grade is generally robust to missing assessments, the number of missing assessments can impact the trajectory of GP5 over time (10). Patients who discontinue treatment due to intolerable toxicity are, by definition, no longer contributing GP5 data. Given that treatment discontinuation due to adverse events was dramatically higher in the control arm (26.8% vs. 4.7%), the surviving PRO data in the control arm may paradoxically underestimate the true tolerability burden by excluding the most severely affected patients (1). The PTT metric does addresses this by measuring the PTT rather than absolute scores, but the potential for survivor bias in the denominator warrants acknowledgment.
Subgroup considerations
The control arm comprised 56 patients receiving cabozantinib and 25 receiving vandetanib, a distribution influenced by the fluctuating availability of vandetanib during the trial (1). These two multikinase inhibitors have overlapping but distinct toxicity profiles and pooling them into a single control group introduces heterogeneity that complicates interpretation. Whether the PRT difference is driven primarily by the cabozantinib subgroup (which constituted 69% of the control arm) or is consistent across both comparators is an important question for subgroup analysis. The parent trial demonstrated that PFS favored selpercatinib over each control agent individually, but whether the same holds for PRT is less clear from the available data (1).
Were major adverse events similar between groups?
The safety profiles of the two treatment strategies were fundamentally different in character. In the parent trial, the most common grade ≥3 adverse events with selpercatinib were alanine aminotransferase (ALT) increase and QT prolongation, which are both monitorable and manageable with dose modification (1). In contrast, the control arm experienced high rates of symptomatic grade ≥3 toxicities including palmar-plantar erythrodysesthesia and mucosal inflammation. The overall incidence of adverse events, including grade ≥3 events, was higher with cabozantinib/vandetanib (1). This asymmetry in the nature of adverse events is the type of distinction where patient-reported measures add the most value. The PRO-CTCAE validation study by Dueck et al. demonstrated that patient and clinician perspectives on adverse events frequently diverge, particularly for subjective symptoms (11). The PRT metric, by centering the patient’s global assessment of bother, provides a complementary lens that integrates the cumulative impact of diverse toxicities into a single, patient-centered measure.
The broader significance: PRT as a regulatory and clinical tool
Perhaps the most consequential aspect of this analysis is its positioning within the regulatory framework. The Food and Drug Administration (FDA) has increasingly emphasized the importance of PRT data, as reflected in their Core Patient-Reported Outcomes in Cancer Clinical Trials guidance, which identified overall adverse event impact and symptomatic adverse events as key PRO concepts (12). Fiero et al. reviewing FDA analyses of PRO data in lung cancer approvals, highlighted the need for prespecified sensitivity analyses and standardized approaches to missing data (13).
The adoption of PRT as a complement to traditional endpoints is particularly compelling in diseases like MTC, where treatments are administered over prolonged periods and cumulative toxicity may be particularly influential in determining real-world treatment persistence. The treatment failure–free survival endpoint in LIBRETTO-531 already captured some of this signal by incorporating treatment discontinuation due to adverse events. The PRT metric adds temporal resolution by estimating how much of their time on treatment was spent in a state of high side-effect burden.
Conclusions
Elisei et al. have demonstrated that selpercatinib is not only more effective but also more tolerable than cabozantinib/vandetanib from the patient’s perspective, using a simple, feasible PRO endpoint. The strong questionnaire compliance, concordance with objective safety data, and growing validation evidence for GP5 collectively support the robustness of these findings. The systematic quantification of how patients feel during treatment is a welcome and overdue evolution. Importantly, future work should consider whether there are specific drug profiles that may be well suited to evaluation using PRT such as drugs with well established tolerability profiles versus innovative compounds with insufficient data on tolerability. As natural language processing of patient feedback (medical record, social media, etc.) becomes easily accessible, metrics like PRT which can collapse multidimensional data and enable comparisons in tolerability will become increasingly useful.
Acknowledgments
OpenEvidence was used to generate a list of potential references using the following prompt: give some articles that would be appropriate to review and reference when writing an editorial on the article “Patient-Reported Tolerability of Selpercatinib Compared to Cabozantinib/Vandetanib: A Secondary Analysis of the LIBRETTO-531 Randomized-Controlled Trial in RET-Mutant Medullary Thyroid Cancer”; query performed on 3/13/26. Grammarly was used for syntax and text editing throughout various versions of the writing process. The opinions, comments, and editorial insights expressed in this submission are solely those of the authors and were human-only generated.
Footnote
Provenance and Peer Review: This article was commissioned by the editorial office, Annals of Thyroid. The article has undergone external peer review.
Peer Review File: Available at https://aot.amegroups.com/article/view/10.21037/aot-2026-0015/prf
Funding: None.
Conflicts of Interest: Both authors have completed the ICMJE uniform disclosure form (available at https://aot.amegroups.com/article/view/10.21037/aot-2026-0015/coif). K.M. reports a grant from Stanford University, publication royalties from Elsevier and Taylor & Francis, and payment for expert testimony. She serves as an unpaid AAO-HNS POEC Chair. None of the disclosures is related to this work or this topic. The other author has no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Hadoux J, Elisei R, Brose MS, et al. Phase 3 Trial of Selpercatinib in Advanced RET-Mutant Medullary Thyroid Cancer. N Engl J Med 2023;389:1851-61. [Crossref] [PubMed]
- Elisei R, Wirth LJ, Capdevila J, et al. Patient-Reported Tolerability of Selpercatinib Compared to Cabozantinib/Vandetanib: A Secondary Analysis of the LIBRETTO-531 Randomized-Controlled Trial in RET-Mutant Medullary Thyroid Cancer. Thyroid 2025;35:1162-72. [Crossref] [PubMed]
- Pearman TP, Beaumont JL, Mroczek D, et al. Validity and usefulness of a single-item measure of patient-reported bother from side effects of cancer therapy. Cancer 2018;124:991-7. [Crossref] [PubMed]
- Cella DF, Tulsky DS, Gray G, et al. The Functional Assessment of Cancer Therapy scale: development and validation of the general measure. J Clin Oncol 1993;11:570-9. [Crossref] [PubMed]
- Peipert JD, Zhao F, Lee JW, et al. Patient-Reported Adverse Events and Early Treatment Discontinuation Among Patients With Multiple Myeloma. JAMA Netw Open 2024;7:e243854. [Crossref] [PubMed]
- Griffiths P, Peipert JD, Leith A, et al. Validity of a single-item indicator of treatment side effect bother in a diverse sample of cancer patients. Support Care Cancer 2022;30:3613-23. [Crossref] [PubMed]
- Payakachat N, Gilligan AM, Altman D, et al. Assessing side-effect bother, burden, and tolerability: A qualitative study exploring the content validity of the Functional Assessment of Cancer Therapy - Item GP5. J Geriatr Oncol 2025;16:102304. [Crossref] [PubMed]
- Basu AB, Umashankar S, Wolf DM, et al. Evaluation of Drug Tolerability as a Function of Toxicity and Quality of Life in Patients Enrolled on I-SPY2. J Clin Oncol 2024;42:11113.
- Krepper D, Hubel NJ, Vorbach SM, et al. What’s missing in patient-reported outcome reporting? A scoping review and aggregated trial-level analysis of completion rates in oncology randomized controlled trials. Crit Rev Oncol Hematol 2026;218:105100.
- Arizmendi C, Zhu Y, Khan M, et al. The FACT-GP5 as a global tolerability measure: responsiveness and robustness to missing assessments. Qual Life Res 2024;33:2869-80. [Crossref] [PubMed]
- Dueck AC, Mendoza TR, Mitchell SA, et al. Validity and Reliability of the US National Cancer Institute’s Patient-Reported Outcomes Version of the Common Terminology Criteria for Adverse Events (PRO-CTCAE). JAMA Oncol 2015;1:1051-9. [Crossref] [PubMed]
- Commissioner O of the. Core Patient-Reported Outcomes in Cancer Clinical Trials. FDA; 2024. Available online: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/core-patient-reported-outcomes-cancer-clinical-trials [Last accessed: 3/25/2026].
- Fiero MH, Roydhouse JK, Vallejo J, et al. US Food and Drug Administration review of statistical analysis of patient-reported outcomes in lung cancer clinical trials approved between January, 2008, and December, 2017. Lancet Oncol 2019;20:e582-9. [Crossref] [PubMed]
Cite this article as: Golchin A, Meister K. Patient-reported tolerability: insights from a secondary analysis of the LIBRETTO-531 study. Ann Thyroid 2026;11:13.

