Developing a named entity framework for thyroid cancer staging and risk level classification using large language models.

Fung, MM; Tang, EH; Wu, T; Luk, Y; Au, IC; Liu, X; Lee, VH; Wong, CK; Wei, Z; Cheng, WY; +8 more...Tai, IC; Ho, JW; Wong, JW; Lang, BH; Leung, KS; Wong, ZS; Wu, JT; Wong, CK

and (2025) Developing a named entity framework for thyroid cancer staging and risk level classification using large language models. NPJ digital medicine, 8 (1). 134-. ISSN 2398-6352 DOI: 10.1038/s41746-025-01528-y

Copy

We developed a named entity (NE) framework for information extraction from semi-structured clinical notes retrieved from The Cancer Genome Atlas-Thyroid Cancer (TCGA-THCA) database and examined Large Language Models (LLMs) strategies to classify the 8th edition of American Joint Committee on Cancer (AJCC) staging and American Thyroid Association (ATA) risk category for patients with well-differentiated thyroid cancer. The NE framework consisted of annotation guidelines development, ground truth labelling, prompting approaches, and evaluation codes. Four LLMs (Mistral-7B-Instruct, Llama-3.1-8B-Instruct, Gemma-2-9B-Instruct, and Qwen2.5-7B-Instruct) were offline utilised for information extraction, comparing with expert-curated ground truth. Our framework was developed using 50 TCGA-THCA pathology notes. 289 TCGA-THCA notes and 35 pseudo-clinical cases were used for validation. Taking an ensemble-like majority-vote strategy achieved satisfactory performance for AJCC and ATA in both development and validation sets. Our framework and ensemble classifier optimised efficiency and accuracy of classifying stage and risk category in thyroid cancer patients.

Item Type	Article
Elements ID	237751
Official URL	https://doi.org/10.1038/s41746-025-01528-y
Date Deposited	24 Mar 2025 10:35

Explore Further

Wong, Carlos King Ho

Dept of Infectious Disease Epidemiology & Dynamics (2023-)

NPJ digital medicine

picture_as_pdf

picture_as_pdf: Fung-etal-2025-Developing-a-named-entity-framework-for-thyroid-cancer-staging-and-risk-level-classification-using-large-language-models.pdf
subject: Published Version
: Available under Creative Commons: Attribution-NonCommercial-No Derivative Works 4.0

View

Download

Atom

BibTeX

OpenURL ContextObject in Span

Multiline CSV

OpenURL ContextObject

Dublin Core

MPEG-21 DIDL

Data Cite XML

EndNote

HTML Citation

JSON

MARC (ASCII)

MARC (ISO 2709)

METS

MODS

RDF+N3

RDF+N-Triples

RDF+XML

RIOXX2 XML

Reference Manager

Refer

Simple Metadata

ASCII Citation

EP3 XML

Export

Downloads