Validation of algorithms in studies based on routinely collected health data: general principles.

Vera Ehrenstein ORCID logo ; Maja Hellfritzsch ORCID logo ; Johnny Kahlert ; Sinéad M Langan ORCID logo ; Hisashi Urushihara ORCID logo ; Danica Marinac-Dabic ORCID logo ; Jennifer L Lund ORCID logo ; Henrik Toft Sørensen ORCID logo ; Eric I Benchimol ORCID logo ; (2024) Validation of algorithms in studies based on routinely collected health data: general principles. American journal of epidemiology, 193 (11). pp. 1612-1624. ISSN 0002-9262 DOI: 10.1093/aje/kwae071
Copy

Clinicians, researchers, regulators, and other decision-makers increasingly rely on evidence from real-world data (RWD), including data routinely accumulating in health and administrative databases. RWD studies often rely on algorithms to operationalize variable definitions. An algorithm is a combination of codes or concepts used to identify persons with a specific health condition or characteristic. Establishing the validity of algorithms is a prerequisite for generating valid study findings that can ultimately inform evidence-based health care. In this paper, we aim to systematize terminology, methods, and practical considerations relevant to the conduct of validation studies of RWD-based algorithms. We discuss measures of algorithm accuracy, gold/reference standards, study size, prioritization of accuracy measures, algorithm portability, and implications for interpretation. Information bias is common in epidemiologic studies, underscoring the importance of transparency in decisions regarding choice and prioritizing measures of algorithm validity. The validity of an algorithm should be judged in the context of a data source, and one size does not fit all. Prioritizing validity measures within a given data source depends on the role of a given variable in the analysis (eligibility criterion, exposure, outcome, or covariate). Validation work should be part of routine maintenance of RWD sources. This article is part of a Special Collection on Pharmacoepidemiology.


picture_as_pdf
Ehrenstein_etal_2024_ Validation of algorithms in studies.pdf
subject
Accepted Version
Available under Creative Commons: Attribution-NonCommercial-No Derivative Works 3.0

View Download

Atom BibTeX OpenURL ContextObject in Span Multiline CSV OpenURL ContextObject Dublin Core Dublin Core MPEG-21 DIDL Data Cite XML EndNote HTML Citation JSON MARC (ASCII) MARC (ISO 2709) METS MODS RDF+N3 RDF+N-Triples RDF+XML RIOXX2 XML Reference Manager Refer Simple Metadata ASCII Citation EP3 XML
Export

Downloads