The human variable: interpreting real-world integration and operator dependence in AI-assisted diagnosis
Introduction
Thyroid ultrasound (US) is a cornerstone of thyroid nodule evaluation, yet diagnostic performance remains highly variable due to differences in operator experience, scanning technique, and subjective interpretation (1,2). Although standardized reporting systems such as the European Thyroid Imaging Reporting and Data System (EU-TIRADS) were developed to improve risk stratification and guide fine needle aspiration (FNA) decisions, their effectiveness depends critically on consistent image acquisition and expert interpretation (3,4). To address these challenges, artificial intelligence (AI)-based computer-aided diagnostic tools, including S-Detect for Thyroid, have been introduced with the promise of providing automated feature extraction and risk assessment to support clinical decision making (5,6). However, most evaluations of these systems are retrospective and rely on curated static images, thereby failing to capture the variability and complexity of real-world US acquisition (7).
The study “Human-AI collaboration for ultrasound diagnosis of thyroid nodules: a clinical trial” offers a timely and clinically relevant contribution by prospectively examining how an established AI system performs when used by operators with varying levels of US expertise (8). By assessing the interaction between human users and AI in a controlled yet realistic setting, the study moves beyond isolated model performance and instead interrogates the full diagnostic pipeline from image acquisition to final risk stratification. The central question posed by this work is not merely whether AI can classify thyroid nodules accurately, but whether it can meaningfully enhance diagnostic decision-making in routine clinical practice or instead remains constrained by operator-dependent variability.
Study overview and results
This clinical trial evaluated how an AI-based diagnostic tool functions within real US workflows by examining its use by 20 participants representing a broad range of experience levels, including eight medical students, three novice US users, and nine experienced physicians. Each participant independently scanned the same five female patients who volunteered for the study, whose thyroid nodule types were later characterized through cytology and histopathology. Using similar US machines with S-Detect installed, participants first performed a focused thyroid US examination and assigned each nodule an EU-TIRADS category based on their own interpretation (3,9). They then applied S-Detect to the same lesion and were asked to reconsider and potentially revise their EU-TIRADS assessment (5,10,11). This sequential design allowed the investigators to capture both the initial human judgment and the subsequent interpretation informed by AI output. A unique aspect of the study is that every participant evaluated the same set of nodules using the same equipment, thereby isolating operator-related variability in image acquisition and decision making. By holding patient and hardware factors constant, the trial provides a controlled environment to examine how differences in scanning technique influence human assessment and downstream AI-assisted interpretation.
Rather than focusing on absolute diagnostic accuracy, the primary finding of this study is the lack of aggregate improvement in participant performance after exposure to AI assistance. Across all 20 participants, mean biopsy recommendation accuracy remained virtually unchanged, shifting from 69.8% (range, 40–100%) at baseline to 69.3% (range, 33–100%) following review of the S-Detect output (P=0.75). While no significant differences were observed between medical students, novice US users, and experienced physicians (P=0.31), this overall stability masks the minimal underlying variability in decision-making (11). Participants altered their biopsy recommendations in only 11 specific instances after AI consultation, and these shifts were not systematically beneficial: while the AI helped participants correct 4 false negatives into true positives, it simultaneously induced 7 detrimental changes (converting 4 true positives to false negatives and 3 true negatives to false positives). Consequently, errors were redistributed rather than reduced. This effect varied by operator group, with the Student group experiencing a slight decline in accuracy (74.4% to 67.1%) due to induced errors, whereas the US novice group showed some improvement (60% to 71.7%), though not statistically significant (P=0.37). Crucially, the dependence on operator acquisition was most evident in the experienced physician group, where the diagnostic accuracy of the AI system itself varied dramatically from 40% to 100% (mean 72.2%), suggesting that even expert-level scanning differences propagate directly into algorithmic predictions. These divergent outcomes, despite identical patient cases and US hardware, underscore the strong dependence of AI-assisted performance on operator interpretation and interaction.
Strengths and limitations
This study has several notable strengths that enhance its relevance for understanding how AI systems function within real clinical workflows. First, the prospective design provides a realistic assessment of human and AI performance, avoiding the limitations of retrospective image selection that often overestimate diagnostic accuracy (12-14). By recruiting participants with diverse levels of US experience, the investigators captured a broad spectrum of acquisition techniques and interpretative behaviors, which more accurately reflect the variability encountered in routine practice. Another key strength is the explicit measurement of human-to-AI interaction. Rather than treating S-Detect as an isolated classifier, the study evaluated how clinicians use the AI output to reconsider their own EU-TIRADS judgments, offering insight into both cognitive and technical aspects of collaboration. Together, these strengths position the study as an important step toward realistic evaluation of AI-enabled US diagnostics.
Despite its strengths, the study has several important limitations that should be considered when interpreting the findings. Foremost among these is the very small patient cohort, consisting of five patients with five thyroid nodules, only one of which was malignant. Such a limited sample substantially reduces statistical power and restricts the generalizability of the results to routine clinical populations. With so few cases, performance metrics become highly sensitive to individual outcomes, making it difficult to draw reliable conclusions about diagnostic accuracy or the true impact of AI assistance. While participants received standardized training in the use of S-Detect, US acquisition itself was not formally standardized beyond routine clinical practice, making it difficult to disentangle acquisition variability from interpretative differences, particularly in the context of the small dataset. Additionally, not all benign nodules were confirmed with postoperative histopathology, relying instead on cytology in some cases, which introduces uncertainty in the reference standard and should be clearly specified in the methods. Finally, the immediate re-evaluation of EU-TIRADS scores after AI output may introduce anchoring effects and does not fully reflect real-world clinical workflows. Collectively, these limitations underscore the need for larger, prospectively designed studies with clearer ground truth definitions to rigorously assess AI-assisted thyroid US diagnosis.
Why AI failed to improve decisions
The largest factor leading to the lack of measurable improvement with AI support in this study is likely the limited size and composition of the patient cohort. With only five nodules, including a single malignant case, the study is inherently underpowered, making performance estimates noisy and highly sensitive to case selection. This constraint likely dominated any potential benefit that AI assistance could provide. Beyond cohort size, the strong dependence of S-Detect on image acquisition quality remains relevant, not because it selectively disadvantages novice operators, but because variability in acquisition affects AI output across all experience levels. Even experienced users produced AI accuracies spanning a wide range, suggesting that acquisition-related variability propagates directly into algorithmic predictions. Additionally, participants may have been reluctant to revise their initial judgments in response to discordant AI outputs, particularly in the absence of transparent explanations. Taken together, these factors, with cohort size and case selection as the primary limitations, likely account for the observed lack of improvement in diagnostic decision making.
A central insight from this study is that diagnostic performance in AI-assisted US is not solely a property of the algorithm but is co-determined by the human operator. Despite all participants scanning the same five patients using identical systems, we observed a wide range in diagnostic accuracy following AI exposure, a variability that persisted even among experienced physicians. Since the patients and the AI software were held constant, this divergence clearly isolates the operator as the primary source of variation. Potentially, performance differences stem either from the quality of the image acquisition that dictates the input quality for AI, or from how individuals choose to weigh and integrate the AI’s recommendation. While acquisition-related factors remain relevant in US, including sensitivity to probe angle, depth, and frame selection, their contribution in this study is difficult to disentangle from the dominant effect of limited sample size. Nevertheless, the results highlight the importance of evaluating human-dependent variability alongside model behavior in prospective studies, rather than treating AI accuracy as a fixed metric.
Implications for clinical deployment and future AI research
Given the limited sample size, it is difficult to formulate definitive recommendations based solely on these results. However, the observed trends highlight critical considerations for hospitals, regulatory bodies, and training programs evaluating AI-enabled US tools. First, the substantial operator dependence observed suggests that AI systems will not meaningfully reduce diagnostic variability unless acquisition techniques are standardized across users. Institutions may therefore need to adopt structured protocols that incorporate “AI aware scanning”, where clinicians are trained not only in traditional US technique but also in how acquisition choices influence downstream algorithmic performance. To support this shift, future diagnostic tools may require real-time guidance capabilities, such as automated view classification, probe angle feedback, or quality scoring, so that operators can correct suboptimal scans during acquisition rather than relying on post hoc analysis. The trial also underscores the need for CAD frameworks to evolve beyond static classifiers and instead function as adaptive, interactive assistants that respond to user actions and image quality in real time. From a regulatory perspective, frameworks for evaluating and reimbursing AI technologies must acknowledge that performance is jointly determined by the model and the operator. Policies that treat AI outputs as operator-independent may misrepresent true clinical effectiveness. These insights highlight the need for holistic deployment strategies that integrate training, workflow redesign, and updated oversight mechanisms.
The results also highlight several important research directions for advancing AI in thyroid US and in operator-dependent imaging more broadly. First, the limitations of single-frame or single-nodule classification systems suggest a need for multimodal or cine-based models that capture the full dynamic context of thyroid scanning rather than relying on isolated static images (15-17). Training algorithms on datasets that reflect diverse operator styles, including suboptimal acquisitions, will be essential for improving robustness and reducing sensitivity to technique-related variability. Emerging approaches in self-supervised learning and multimodal modeling, such as combining US with clinical text reports, may help reduce dependence on specific views and manual frame selection (18,19). Foundation models trained on large scale heterogeneous US data offer another promising direction for mitigating acquisition noise (20-23). The trial also underscores the importance of prospective multicenter studies that collect standardized metadata and acquisition logs, enabling researchers to directly examine how operator behavior influences AI performance. Integration of automated quality control systems that assess image adequacy before diagnosis could further enhance reliability. Finally, incorporating uncertainty quantification may allow AI tools to flag low confidence predictions that stem from poor acquisition, guiding users toward rescanning and more informed decision making.
AI-assisted US machines are frequently promoted as technologies that can support novice users and help bridge gaps in experience. However, this study indicates that the gap remains difficult to close. Experts derived no benefit as they rarely altered their initial assessments, while novices paradoxically limited their own support by providing lower-quality image inputs that degraded the AI’s diagnostic accuracy. This dynamic does challenge the common assumption that decision support systems operate independently of the human users who operate them. Instead, meaningful collaboration requires bidirectional interaction in which the AI does not merely process images but also provides feedback to guide acquisition. Real-time prompts related to probe alignment, view adequacy, or representative frame selection could help novices produce inputs that the AI can interpret reliably. At the same time, AI outputs must be accompanied by interpretable explanations that integrate smoothly into clinical decision making. Without these elements, the collaboration remains incomplete, with the AI functioning more as a passive reviewer than an active partner in clinical reasoning.
Conclusions
This prospective study offers an informative but ultimately inconclusive examination of AI-assisted thyroid US diagnosis due to its limited sample size. Nevertheless, the experimental design and observed trends highlight important considerations for the field. In particular, the results suggest that variability introduced by human acquisition and interpretation may substantially influence the effectiveness of AI systems in US. These findings reinforce the distinction that AI is a dependent tool rather than a ‘plug-and-play’ solution. Because its utility is inextricably linked to operator performance, successful clinical implementation requires not only accurate algorithms but also a focus on system usability and dedicated operator training. Future work should focus on developing AI systems that support bidirectional interaction, including guidance during image acquisition and outputs that are interpretable and responsive to user behavior. Larger, rigorously designed prospective studies will be essential to determine whether such approaches can meaningfully improve thyroid cancer risk stratification in routine clinical practice.
Acknowledgments
None.
Footnote
Provenance and Peer Review: This article was commissioned by the editorial office, Annals of Thyroid. The article did not undergo external peer review.
Funding: None.
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://aot.amegroups.com/article/view/10.21037/aot-2026-1-0002/coif). The authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Grani G, Lamartina L, Cantisani V, et al. Interobserver agreement of various thyroid imaging reporting and data systems. Endocr Connect 2018;7:1-7. [Crossref] [PubMed]
- Solymosi T, Hegedűs L, Bonnema SJ, et al. Considerable interobserver variation calls for unambiguous definitions of thyroid nodule ultrasound characteristics. Eur Thyroid J 2023;12:e220134. [Crossref] [PubMed]
- Russ G, Bonnema SJ, Erdogan MF, et al. European Thyroid Association Guidelines for Ultrasound Malignancy Risk Stratification of Thyroid Nodules in Adults: The EU-TIRADS. Eur Thyroid J 2017;6:225-37. [Crossref] [PubMed]
- Russ G, Trimboli P, Buffet C. The New Era of TIRADSs to Stratify the Risk of Malignancy of Thyroid Nodules: Strengths, Weaknesses and Pitfalls. Cancers (Basel) 2021;13:4316. [Crossref] [PubMed]
- Barczyński M, Stopa-Barczyńska M, Wojtczak B, et al. Clinical validation of S-Detect(TM) mode in semi-automated ultrasound classification of thyroid lesions in surgical office. Gland Surg 2020;9:S77-85. [Crossref] [PubMed]
- Zhang D, Jiang F, Yin R, et al. A Review of the Role of the S-Detect Computer-Aided Diagnostic Ultrasound System in the Evaluation of Benign and Malignant Breast and Thyroid Masses. Med Sci Monit 2021;27:e931957. [Crossref] [PubMed]
- Yang K, Chen J, Wu H, et al. S-Thyroid Computer-Aided Diagnosis Ultrasound System of Thyroid Nodules: Correlation Between Transverse and Longitudinal Planes. Front Physiol 2022;13:909277. [Crossref] [PubMed]
- Edström AB, Makouei F, Wennervaldt K, et al. Human-AI collaboration for ultrasound diagnosis of thyroid nodules: a clinical trial. Eur Arch Otorhinolaryngol 2025;282:3221-31. [Crossref] [PubMed]
- Chung SR, Ahn HS, Choi YJ, et al. Diagnostic Performance of the Modified Korean Thyroid Imaging Reporting and Data System for Thyroid Malignancy: A Multicenter Validation Study. Korean J Radiol 2021;22:1579-86. [Crossref] [PubMed]
- Warm JJ, Melchiors J, Kristensen TT, et al. Head and neck ultrasound training improves the diagnostic performance of otolaryngology residents. Laryngoscope Investig Otolaryngol 2024;9:e1201. [Crossref] [PubMed]
- Yoo YJ, Ha EJ, Cho YJ, et al. Computer-Aided Diagnosis of Thyroid Nodules via Ultrasonography: Initial Clinical Experience. Korean J Radiol 2018;19:665-72. [Crossref] [PubMed]
- Wildman-Tobriner B, Buda M, Hoang JK, et al. Using Artificial Intelligence to Revise ACR TI-RADS Risk Stratification of Thyroid Nodules: Diagnostic Accuracy and Utility. Radiology 2019;292:112-9. [Crossref] [PubMed]
- Gichoya JW, Thomas K, Celi LA, et al. AI pitfalls and what not to do: mitigating bias in AI. Br J Radiol 2023;96:20230023. [Crossref] [PubMed]
- Lee SE, Hong H, Kim EK. Diagnostic performance with and without artificial intelligence assistance in real-world screening mammography. Eur J Radiol Open 2024;12:100545. [Crossref] [PubMed]
- Athreya S, Melehy A, Suthahar SSA, et al. Combining Ultrasound Imaging and Molecular Testing in a Multimodal Deep Learning Model for Risk Stratification of Indeterminate Thyroid Nodules. Thyroid 2025;35:590-4. [Crossref] [PubMed]
- Zhuang L, Ivezic V, Feng J, et al. Patient-level thyroid cancer classification using attention multiple instance learning on fused multi-scale ultrasound image features. AMIA Annu Symp Proc 2023;2023:1344-53.
- Radhachandran A, Vittalam A, Ivezic V, et al. ThyGraph: A Graph-Based Approach for Thyroid Nodule Diagnosis from Ultrasound Studies. In: Medical Image Computing and Computer Assisted Intervention – MICCAI 2024. 2024;753-63.
- Xiao X, Zhou Y, Zhu Y, et al. A multi-modal prompt-tuning method of ultrasound diagnosis for thyroid nodule. Front Med (Lausanne) 2025;12:1686374. [Crossref] [PubMed]
- Dadoun H, Delingette H, Rousseau AL, et al. Joint Representation Learning from French Radiological Reports and Ultrasound Images. In: 2023 IEEE 20th International Symposium on Biomedical Imaging (ISBI). IEEE; 2023:1-5.
- Jiang Y, Feng CM, Ren J, et al. From pretraining to privacy: federated ultrasound foundation model with self-supervised learning. NPJ Digit Med 2025;8:714. [Crossref] [PubMed]
- Kang Q, Lao Q, Gao J, et al. URFM: A general Ultrasound Representation Foundation Model for advancing ultrasound image diagnosis. iScience 2025;28:112917. [Crossref] [PubMed]
- Jiao J, Zhou J, Li X, et al. USFM: A universal ultrasound foundation model generalized to tasks and organs towards label efficient image analysis. Med Image Anal 2024;96:103202. [Crossref] [PubMed]
- Meyer A, Murali A, Zarin F, et al. Ultrasam: a foundation model for ultrasound using large open-access segmentation datasets. Int J Comput Assist Radiol Surg 2026;21:93-102. [Crossref] [PubMed]
Cite this article as: Athreya S, Radhachandran A, Ivezić V, Arnold CW, Speier W. The human variable: interpreting real-world integration and operator dependence in AI-assisted diagnosis. Ann Thyroid 2026;11:3.

