LANGUAGE models accurately extracted smoking history from clinical notes, supporting faster lung cancer screening eligibility decisions.
Lung Cancer Screening Eligibility Tested Across 3,000 Notes
Researchers compared a low-cost, non-generative structured-judgment model with four general-purpose large language models for extracting smoking status, pack-years, and quit date from outpatient clinical notes. These variables were then used to determine lung cancer screening eligibility under 2021 U.S. Preventive Services Task Force and 2023 American Cancer Society criteria.
The benchmark included 3,000 synthetic notes across three increasingly difficult conditions. One thousand were template-generated, 1,000 were realistic “messy” notes created from structured facts, and 1,000 messy notes required more complex arithmetic to calculate pack-years and quit dates. Reference labels were generated before the notes were created.
Accuracy Remained High, but Complexity Exposed Differences
All five systems achieved high accuracy, although performance diverged when smoking histories became more complex. The structured-judgment model correctly determined eligibility for 99.1% of template notes and 99.9% of realistic messy notes, but accuracy fell to 94.4% in the complex condition.
The four general-purpose large language models correctly classified 98.1% to 99.8% of complex notes. The two highest-performing systems reached 99.8% and 99.6%, while the lowest-cost large language model achieved 98.5%.
Complexity also affected clinically important errors. In the difficult note set, the structured-judgment model produced 27 false-positive and 14 false-negative lung cancer screening flags per 1,000 notes, more than any other system evaluated.
Speed and Cost Shape Clinical Decision Support
The structured-judgment model was fastest, with median processing times of 0.45 to 1.21 seconds per note, compared with 1.71 to 3.50 seconds for the other models. Its cost ranged from $0.61 to $0.64 per 1,000 notes.
However, the least expensive large language model cost only $0.12 to $0.21 per 1,000 notes and was more accurate on complex notes, although slower.
For clinicians and health systems, the findings suggest that extracting smoking history from narrative documentation could support real-time lung cancer screening decision support. The authors cautioned that model capabilities, latency, and pricing are changing rapidly, making ongoing benchmarking important. The study is a preprint and has not yet undergone journal peer review.
Reference
Wright A et al. Extracting smoking history from clinical notes for lung cancer screening decision support: comparing a structured-judgment model with general-purpose large language models. medRxiv. 2026. doi:10.64898/2026.09.24.26363906.
Featured Image: Dragana Gordic on Adobe Stock.