National Radiology Turnaround Time: A Perspective

SCHOLARLY COMMENTARY · DIGITAL EDITION

Interpreting National Radiology
Turnaround Time Trends

Kelly Emrick, DHSc, PhD, MBA, BSRT(ARRT)R

Contents & search

Report page 1: Title and abstract
Report page 2: Study contribution and measurement
Report page 3: Measurement validity and capacity
Report page 4: Modality differences
Report page 5: Evidence and technology
Report page 6: Diagnostic completion and operations
Report page 7: Accountability and implementation
Report page 8: Research priorities and conclusion
Report page 9: References

Continue the conversation

Interpreting National Radiology Turnaround Time Trends

Kelly Emrick, DHSc, PhD, MBA, BSRT(ARRT)R

kellyemrick.com

Use Reading view for selectable text and reference links.

SCHOLARLY COMMENTARY

Interpreting National Radiology Turnaround Time Trends

Measurement validity, workforce capacity, and responsibility for timely diagnosis

Kelly Emrick, DHSc, PhD, MBA, BSRT(ARRT)R

2026

Abstract

Increasing delays in imaging interpretation warrant an operational response, but the strength of that response depends on distinguishing an observed trend from its proposed cause. This commentary examines Christensen and colleagues’ analysis of Medicare fee-for-service imaging during 2014 to 2023. It argues that claims-derived calendar-date intervals provide a useful surveillance signal while requiring validation against clinical timestamps before they can support precise statements about waiting time. Workforce saturation is a plausible explanation for deteriorating responsiveness; establishing it requires information about staffed capacity, examination complexity, demand patterns, and workflow. Complementary research on radiologist workload, fatigue, attrition, artificial intelligence, and follow-up of abnormal results supports a broader assessment of diagnostic reliability. The proposed leadership response combines measurement validation with targeted workflow changes, explicit responsibility for actionable findings, and monitoring of accuracy and workforce well-being. National trends should prompt investigation and resource allocation without becoming an unsupported local productivity target.

Figure 1. The clinical responsibility behind a turnaround measure

Figure from the original report

The contribution of the study

Christensen, Drake, et al. (2026) analyzed 2,578,953 selected office and hospital outpatient imaging studies from a 5% Medicare fee-for-service sample. They linked technical and professional claims to measure calendar-date differences between acquisition and interpretation. Mean intervals increased from 0.091 days in 2014 to 0.193 days in 2023, an 113% increase. Most growth occurred in the final two years. CT and MR showed particularly large relative increases, and disadvantaged communities generally had longer intervals. The authors interpret these patterns as suggesting that the radiology workforce has reached maximum capacity.

The study makes a valuable contribution by providing a national indicator against which local experience can be examined. A sustained change in such an indicator deserves attention even when its causal explanation remains unsettled. Leaders do not need definitive proof of the mechanism before investigating an aging worklist, verifying coverage, or protecting patients with urgent findings. They do need stronger evidence before attributing every delay to insufficient radiologist effort or prescribing a single remedy across settings.

Figure 2. Change in the mean claims-derived calendar-date interval

Figure from the original report

Source: Christensen, Drake, et al. (2026). Endpoints only; no intervening annual values are inferred. The authors’ 113% estimate is retained because the displayed means are rounded. These values are not exact elapsed-time measurements.

A national average should initiate a local inquiry. It cannot establish the appropriate report deadline for a suspected cord compression, a routine surveillance examination, and an incidental finding requiring interval follow-up. Those decisions depend on clinical urgency and the consequences of delay. A useful response therefore connects statistical surveillance to the clinical decisions that timely interpretation is intended to support.

Measurement validity comes first

The distinction between a date and a timestamp changes what a turnaround statistic can mean. Consider the illustrative cases in Table 1. A calendar-date indicator can assign the same value to clinically different waits and different values to nearly identical waits. Multiplying its average by 24 would change the unit label without recovering the unobserved hours. It could encourage a misleading claim about the exact time patients waited.

Table 1. Why calendar-date differences and elapsed time answer different questions

Illustrative acquisition and interpretationDate differenceActual elapsed time
Monday 09:00 to Monday 09:050 days5 minutes
Monday 00:01 to Monday 23:590 days23 hours 58 minutes
Monday 23:59 to Tuesday 00:011 day2 minutes

Original hypothetical examples demonstrating measurement properties. They are not observations from the Medicare study.

Date-based surveillance can still be useful if its limitations are explicit and its relationship to clinical timing remains sufficiently stable. A validation study should compare the claims indicator with examination completion, preliminary interpretation, final signature, and result-release timestamps in radiology information systems and picture archiving and communication systems. It should assess whether agreement differs by site, billing arrangement, modality, and acquisition time. A shift toward later scanning, for example, could change the probability of crossing midnight even if elapsed reporting time changed less.

Claim linkage also requires careful scrutiny. Relevant questions concern the handling of repeated procedure codes, matching windows, corrected claims, combined billing, and examinations without a matched interpretation. These are appraisal questions, not established defects in the investigators’ methods. Their importance lies in determining which services contribute to the denominator and whether that selection changes over time. A precise estimate from a large cohort cannot, by itself, resolve systematic differences between the measured indicator and the clinical construct of interest.

Future reporting should distinguish the share of examinations crossing a date boundary from the duration of the longest waits. Actual timestamp data should support medians, upper percentiles, and urgency-specific deadline breaches. A median can remain stable while a vulnerable minority experiences prolonged delay. Conversely, a changing mean can reflect relatively few outliers. Both patterns require a different response from a uniform increase across the entire distribution.

Workforce capacity is a hypothesis to test

A rise in delayed interpretations is compatible with a demand-capacity imbalance. It is also compatible with uneven allocation of existing capacity, inadequate overnight coverage, competing procedural responsibilities, the unavailability of comparison studies, or changes in case complexity. These mechanisms can coexist. An observational time trend cannot determine their separate contributions without additional measurement and a defensible comparison strategy.

Projected demand adds context without proving saturation. Christensen et al. (2025) modeled future imaging use under alternative assumptions about population and utilization. Such forecasts should be used as planning scenarios with explicit assumptions, rather than treated as predetermined demand. A practice should update its own forecast when referral patterns, local demographics, or clinical pathways change.

Capacity should therefore be measured in available clinical work time and relevant expertise, with allowances for work beyond report production. The same headcount can supply different coverage when leave, part-time practice, supervision, procedures, or subspecialty requirements change. The operational question is whether the right expertise is available when a particular class of examinations enters the worklist. A department may have sufficient weekly hours and still experience recurrent shortages at predictable times.

Figure 3. Relative increases differ substantially by modality

Figure from the original report

Source: Christensen, Drake, et al. (2026). Relative changes from 2014 to 2023. Modality-specific baseline intervals are required to compare absolute waiting times; these percentages do not supply them.

A stronger research design would link examination-level timing to staffed hours, case complexity, workload arrivals, and backlog within the same practice over time. Comparisons before and after a staffing or workflow intervention should account for concurrent changes and use comparable sites where feasible. Interrupted time-series analysis or a carefully justified difference-in-differences design could strengthen inference. Neither method substitutes for checking its assumptions, especially changing case mix and preintervention trends.

Relevant evidence extends beyond report volume

Workload research illustrates why a report count is an incomplete measure of capacity. MacDonald et al. (2013) observed substantial professional activity outside reporting, including procedures, supervision, consultation, and teaching. Those single-department proportions should not be transferred directly to a United States staffing model. The applicable principle is to measure the work that a service actually requires before determining its staffing needs.

Evidence on fatigue gives a further reason to avoid indiscriminate acceleration. Krupinski et al. (2010) found lower fracture-detection performance after a clinical workday in an experimental reader study. That experiment does not estimate present national error rates, but it supports treating diagnostic accuracy and reader fatigue as outcomes to monitor when productivity changes. An intervention that shortens the queue by making interpretation less reliable has not demonstrated better diagnostic service.

Table 2. Complementary empirical evidence and its appropriate use

Study and designContribution to this responseLimit on interpretation
MacDonald et al. 2013 Departmental workload studyMeasure procedures, supervision, and consultation alongside reporting.One New Zealand department; workload shares are context-specific.
Krupinski et al. 2010 Experimental reader studyInclude accuracy and fatigue when assessing speed.Small fracture-detection experiment; not a national estimate of error.
Christensen, Liu, et al. (2026) Radiologist attrition cohortMeasure retention and clinical availability when planning capacity.Claims-derived attrition does not establish why radiologists leave.
Singh et al. 2009 Abnormal-result follow-up studyVerify action after an actionable result is reported.One integrated care setting; historical rates are not current national rates.
Batra et al. 2023 Retrospective AI workflow studySeparate worklist waiting from interpretation time.Task-specific before-and-after evidence; causal and generalizability limits.

Source details appear in the references. This is a selected evidence synthesis, not a systematic review.

A recent United States cohort study documents increasing radiologist attrition and provides empirical support for concerns about workforce sustainability (Christensen, Liu, et al., 2026). It cannot establish that fatigue caused the increase. For leadership, the implication is to assess retention, actual clinical availability, and recruitment alongside throughput. Staffing decisions that preserve output for one quarter while increasing the probability of subsequent departures may worsen future access.

Technology should address a defined source of delay

Artificial intelligence may help when its function matches the source of delay. In a retrospective study of pulmonary embolism, Batra et al. (2023) found that worklist reprioritization was associated with shorter turnaround times for positive examinations. The improvement was concentrated in waiting time, while interpretation time was essentially unchanged. This distinction matters because moving a case earlier in a queue and reducing the work needed to interpret it are different interventions. Evidence for one should not be presented as evidence for the other.

A local evaluation should include all relevant worklist groups. If urgent positive studies move forward, leaders should check whether lower-priority studies are delayed and whether missed or false-positive classifications affect care. Useful outcomes include time to an actionable report, classification performance, workload transferred to clinicians, and the age of the oldest outstanding examinations. Economic assessment should include integration, monitoring, correction, and maintenance work. A favorable result in one imaging task does not establish that the same technology can solve an entire department’s capacity problem.

More recent randomized evidence from the Swedish MASAI trial supports evaluating AI against clinical outcomes. Gommers et al. (2026) reported noninferior interval cancer rates with AI-supported mammography compared with standard double reading. That finding strengthens the case for prospective validation, while its screening context limits transfer to United States CT or MRI reporting. Local workflow benefit still requires local evidence.

Diagnostic completion requires responsibility after reporting

Timely interpretation becomes clinically useful when the result reaches an appropriate clinician and supports action. Singh et al. (2009) documented incomplete follow-up of abnormal imaging results despite electronic notification. This supports measuring what happens after a report becomes available. Receipt, acknowledgment, and an appropriate documented plan are distinct events; a visible alert alone does not complete the pathway.

Table 3. A conceptual matrix for interpreting operational performance

Reliable action on clinically important findingsUnreliable or unverified action
Reports meet urgency-based deadlinesMaintain performance and monitor diagnostic quality.Investigate routing, acknowledgment, and ownership of follow-up.
Reports miss urgency-based deadlinesInvestigate staffing, work distribution, and causes of delay.Coordinate reporting recovery with active review of unresolved findings.

Original conceptual synthesis informed by Singh et al. (2009). This matrix has not been empirically validated. Timeliness thresholds require local clinical governance.

Equity assessment should follow the same clinical pathway. Neighborhood disadvantage can identify communities requiring investigation, but it should not be treated as a direct measure of an individual patient’s resources. Leaders should examine access to appointments, reporting coverage, communication barriers, and completion of recommended care. Adjustments can help explain differences, while unadjusted results remain necessary to reflect the experiences of the populations being served.

An operating response for outpatient imaging

For an outpatient MRI service, the practical starting point is a shared definition of when each interval begins and ends. The medical director should define urgency categories and escalation expectations with referring clinicians. Operational leaders should verify that completed studies are added to the correct reading worklist and remain visible until responsibility is resolved. These are proposed management actions, not tested recommendations from the national study.

Table 4. Proposed measures and accountable functions

MeasureDefinition and purposeResponsible function
Access to examinationOrder received to completed scan; distinguish scheduling and preparation delays.Scheduling and modality operations
Interpretation intervalCompleted scan to clinically available report; report median and upper percentiles by urgency.Radiology medical leadership and analytics
Outstanding workCount and age of all completed examinations awaiting interpretation, including overdue cases.Daily worklist lead
Actionable findingsTime to acknowledgment and documented follow-up plan within clinically defined deadlines.Named reporting and referring clinicians
Quality and sustainabilityReview discrepancies, material amendments, overtime, fatigue indicators, and retention with throughput.Quality lead and practice leadership

Original recommendations. No numerical target in this framework should be interpreted as a national standard. Denominators, exclusions, and missing timestamps must be documented.

A daily review should identify the oldest unresolved studies and any urgent case lacking an accountable reader. Weekly reviews should examine recurring causes by modality, acquisition time, site, and reading group. Monthly leadership review should connect demand forecasts with staffing availability and quality findings. Where contractual reading coverage is used, reporting obligations should include escalation and coverage continuity, with clinical review of exceptions.

Implementation should begin with a documented baseline, one clearly defined intervention, and a prespecified assessment period. Outcomes should be reviewed for clinically meaningful improvement and unintended effects. For example, adding a focused evening reading session can be assessed against evening backlog and next-day delays while tracking diagnostic discrepancies and total work hours. This creates a testable local response without assuming that a national trend has already identified the local cause.

Priorities for the next generation of research

The next research step should validate the measurement and then explain variation. Multisite linkage between claims and clinical timestamps would establish when a calendar-date measure tracks actual waiting and when it diverges. Reporting the proportion of eligible examinations successfully linked, together with its stability across years and billing arrangements, would help readers judge the population represented by the indicator. These checks would support stronger surveillance while making its boundaries more transparent.

Subsequent studies should examine within-practice changes in staffed hours, incoming workload, examination complexity, and interpretation intervals. They should explicitly consider repeated observations within patients, radiologists, and organizations. A larger sample size reduces statistical uncertainty only for the model being estimated; it does not remove confounding, selection effects, or limitations in how the outcome was measured. Sensitivity analyses should therefore address plausible alternative explanations rather than rely on statistical significance alone.

Patient outcomes deserve direct measurement. Relevant endpoints include time to appropriate clinical action, missed follow-up, repeat examinations attributable to unavailable results, and avoidable disruption of care. The relation between turnaround and harm is likely to depend on indication and urgency, so a uniform time target may conceal clinically important variation. Prospective evaluation should also ask whether an intervention benefits all patients on the worklist and whether gains persist without increasing fatigue or turnover.

Generalizability requires deliberate extension. Findings from fee-for-service outpatient imaging should be compared with other payer populations and clinical settings before being used to characterize the entire United States diagnostic system. Changes in enrollment composition also deserve attention in long-term comparisons. Such analyses can establish whether observed patterns reflect a broad system constraint, particular service arrangements, or different combinations of causes across markets.

Conclusion

Christensen and colleagues provide a reason to investigate the deterioration in diagnostic responsiveness. The appropriate response is to strengthen the connection between the reported indicator, the mechanism of delay, and the clinical outcome that matters. Workforce expansion may be needed in some settings; in others, redistributing coverage or removing workflow barriers may be more effective. Decisions should follow measured local conditions and be evaluated against the standards of patient care and professional sustainability. The leadership obligation is to make completed imaging reliably available for clinical decisions, with sufficient time and expertise to preserve diagnostic accuracy.

Appraisal scope: the original article’s published abstract and investigators’ visual summary, supplemented by the peer-reviewed studies cited here. Detailed evaluation of claim-linkage rules, statistical models, and sensitivity analyses requires the complete article. This narrative commentary presents no new patient-level analysis.

References

Batra, K., Xi, Y., Bhagwat, S., Espino, A., & Peshock, R. M. (2023). Radiologist worklist reprioritization using artificial intelligence: Impact on report turnaround times for CTPA examinations positive for acute pulmonary embolism. American Journal of Roentgenology, 221(3), 324–333. https://doi.org/10.2214/AJR.22.28949

Christensen, E. W., Drake, A. R., Parikh, J. R., Rubin, E. M., & Rula, E. Y. (2025). Projected US imaging utilization, 2025 to 2055. Journal of the American College of Radiology, 22(2), 151–158. https://doi.org/10.1016/j.jacr.2024.10.017

Christensen, E. W., Drake, A. R., Rula, E. Y., Yuan, C. X., Wald, C., Johnson, M. H., & Nicola, G. N. (2026). National turnaround time trends for Medicare fee-for-service beneficiaries, 2014-2023. Journal of the American College of Radiology, 23(7), 1244-1252. https://doi.org/10.1016/j.jacr.2026.02.038

Christensen, E. W., Liu, C.-M., Rula, E. Y., & Parikh, J. R. (2026). Attrition of the national radiologist workforce: Associations with radiologist and practice characteristics. American Journal of Roentgenology, 226(1), e2533587. https://doi.org/10.2214/AJR.25.33587

Gommers, J., Hernström, V., Josefsson, V., Sartor, H., Schmidt, D., Hjelmgren, A., Larsson, A.-M., Hofvind, S., Andersson, I., Rosso, A., Hagberg, O., & Lång, K. (2026). Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study: A randomized, controlled, non-inferiority, single-blinded, population-based, screening-accuracy trial. The Lancet, 407(10527), 505–514. https://doi.org/10.1016/S0140-6736(25)02464-X

Krupinski, E. A., Berbaum, K. S., Caldwell, R. T., Schartz, K. M., & Kim, J. (2010). Long radiology workdays reduce detection and accommodation accuracy. Journal of the American College of Radiology, 7(9), 698–704. https://doi.org/10.1016/j.jacr.2010.03.004

MacDonald, S. L. S., Cowan, I. A., Floyd, R. A., & Graham, R. (2013). Measuring and managing radiologist workload: A method for quantifying radiologist activities and calculating the full-time equivalents required to operate a service. Journal of Medical Imaging and Radiation Oncology, 57(5), 551–557. https://doi.org/10.1111/1754-9485.12091

Singh, H., Thomas, E. J., Mani, S., Sittig, D., Arora, H., Espadas, D., Khan, M. M., & Petersen, L. A. (2009). Timely follow-up of abnormal diagnostic imaging test results in an outpatient setting: Are electronic medical records achieving their potential? Archives of Internal Medicine, 169(17), 1578–1586. https://doi.org/10.1001/archinternmed.2009.263

Use Next to open the book. Arrow keys turn pages; Home and End jump to the covers.

Original report preserved in nine pages, plus a digital back cover. Downloads work offline; external references require internet access.