NIH Translates 70 Years of Health Data Into Unified Language

Key Takeaways

  • NHLBI’s BioData Catalyst platform harnesses over 12 petabytes of multimodal health data.
  • The initiative includes a “converter box” approach for data interoperability across various health studies.
  • Collaboration between NHLBI and ODSS aims to integrate research standards into electronic medical records, enhancing data utility for AI applications.

Data Integration Challenges

NHLBI’s BioData Catalyst stands at the forefront of an extensive project that integrates a vast array of health data. Developed in collaboration with the National Library of Medicine (NLM) and the Office of Data Science Strategy (ODSS), it houses over 12 petabytes of multimodal data, including genomics, clinical imaging, and sensor data. Notably, this data encompasses long-term studies like the Trans-Omics for Precision Medicine (TOPMed) program, tracking around 180,000 individuals.

Access to this substantial dataset alone poses challenges for artificial intelligence (AI) readiness. A major hurdle is interoperability, which refers to ensuring consistency between data points across different studies. For instance, a cardiovascular measurement from the 1990s Framingham Heart Study must align with contemporary data from a pulmonary fibrosis study. To address this, NHLBI has implemented a linked data modeling language (LinkML) pipeline. This “converter box” effectively standardizes data inputs, enabling researchers to analyze them seamlessly.

Sweta Ladwa, chief of the Scientific Solutions Delivery Branch at NHLBI, emphasized the importance of clinical validation in automated data mapping. The initiative collaborates with pulmonologists to confirm that similar medications and health conditions are classified correctly, ensuring that the resulting data is accurate and reliable. The AI-assisted mapping process primarily utilizes publicly available metadata without compromising patient confidentiality.

Advancements in Clinical Systems

While NHLBI concentrates on optimizing existing research data, the ODSS is focused on integrating research-grade standards into daily clinical systems. Susan Gregurick, the NIH associate director for data science and director of ODSS, highlighted efforts to incorporate NIH research standards into the United States Core Data for Interoperability (USCDI). This standard functions as a basis for electronic medical records (EMR) used for system accreditation.

The ongoing work began in oncology and is set to expand to other areas, including cardiovascular health in collaboration with NHLBI. This integration means that when cardiovascular-related data arises in patient encounters, even if they are not part of formal studies, EMR systems can capture it in a manner that researchers can utilize.

Gregurick noted the profound implications of such cross-agency collaborations, suggesting that they play a critical role in driving future AI innovations in healthcare. With standards from research translating into actionable data in clinical systems, the potential for enhanced patient care and improved research outcomes is significant.

The content above is a summary. For more details, see the source article.

Leave a Comment

Your email address will not be published. Required fields are marked *

ADVERTISEMENT

Become a member

RELATED NEWS

Become a member

Scroll to Top