ICD Code Embeddings using LLMs
Speaker: Michael Kane, PhD
Affiliation: Yale
Date: September 2024
Watch the recording on YouTube
Abstract
Dr. Michael Kane from Yale will be talking about their work on ICD Code Embeddings. This work is fascinating because it provides a pathway for the rich information encoded in the patient ICD codes to be included in predictive models.
There are tens of thousands of ICD codes, so including each code individually creates too many features. Using parent codes to reduce the number of classes throws away patient specific detail, e.g. diabetes with and without complications may simply roll up to diabetes, along with gestational diabetes. ICD Code Embeddings provide a sophisticated path to efficiently represent a patient’s clinical condition as documented in ICD codes in as low as a ten-dimensional space, perhaps making it possible to predict patient outcomes based on these ten columns.
Speaker bio
Michael Kane is an Assistant Professor in Yale University’s Biostatistics Department. He develops methods in statistical/machine learning in biomedicine to understand patient-level heterogeneity in clinical trials and to understand patterns of human mobility.
Rights and attribution
This recording and accompanying text are presented for educational access with attribution to the speaker. Speaker views are their own. No open license is implied unless explicitly stated.