Clinical Language Models for Real-Time EHR Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing clinical predictive models rely heavily on structured inputs, leading to complexity in data processing and deployment, and are often not deployed in real-world clinical settings, while conventional large language models have not effectively integrated with clinical workflows to support medical decision-making.

Innovation Solution

A language-model based system, NYUTron, that utilizes self-supervised LLMs to process structured and unstructured clinical data, integrating with clinical workflows for real-time decision support by converting clinical notes into training data, finetuning a machine learning model, and generating medical predictions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If conventional large language models are used to process clinical data, then the model can potentially access comprehensive patient information from unstructured notes, but the model fails to integrate effectively with clinical workflows and does not provide practical decision support

Engineering Contradiction:
Improvecomprehensive patient information accessVSAvoidclinical workflow integration
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent introduces a specialized clinical language model as an intermediary between unstructured clinical notes and decision support systems. This model is specifically trained on medical literature and clinical data, acting as a mediator that translates unstructured text into structured clinical insights that can be effectively integrated into existing workflow systems

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The language model is designed to perform multiple functions within clinical workflows, including extracting patient information, generating clinical predictions, supporting decision-making, and interfacing with electronic health records. This multi-functionality enables seamless integration across different clinical tasks and systems

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of manufacture

If clinical predictive models rely on structured inputs, then the models can be developed and deployed with defined parameters, but the data processing complexity and deployment challenges increase significantly

Engineering Contradiction:
Improvemodel developmentVSAvoiddata processing complexity
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical data processing approaches with a neural network-based language model. Instead of manually structuring and processing clinical data through complex algorithms, the system uses deep learning to automatically extract and process information from unstructured notes, significantly reducing data processing complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The model transforms the nature of input parameters by accepting unstructured text data rather than requiring pre-structured clinical inputs. This parameter change simplifies the development process while maintaining the ability to generate accurate clinical predictions

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If clinical predictive models are developed with rigorous training and validation, then the models achieve high predictive accuracy, but the models are rarely deployed to assess their impact on real-world clinical care

Engineering Contradiction:
Improvepredictive accuracyVSAvoiddeployment rate
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The language model is designed to be self-sufficient in handling the full pipeline from data input to clinical prediction output. It automatically processes unstructured clinical notes, extracts relevant features, and generates predictions without requiring extensive manual intervention or complex integration infrastructure, enabling rapid deployment

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The model is pre-trained on extensive medical literature and clinical data before deployment. This preliminary training ensures high predictive accuracy is achieved beforehand, allowing the model to be deployed quickly to real-world settings without requiring extensive post-deployment tuning or validation

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250357007A1Systems, apparatus, methods and computer-accessible medium for providing health system scale language models which can include clinical prediction engines
Publication Date: 2025.11.20 NEW YORK UNIV
  • US20250357007A1 patent drawing
  • US20250357007A1 patent drawing
  • US20250357007A1 patent drawing

AI summary

Exemplary systems, methods, and computer-accessible medium are provided that that can implement and/or utilize clinical predictive models, which can assist physicians and administrators make decisions by forecasting clinical and operational events. Thus, the exemplary systems, methods, and computer-accessible medium are provided that convert clinical notes to training data using at least one natural language processing procedure, train a machine learning model using the training data finetune the trained machine learning model based on selected parameters, receive patient data, and generate at least one medical prediction on the received patient data with the trained finetuned machine learning model. Additional exemplary systems, methods, and computer-accessible medium are provided that can generate a table language by implementing an artificial intelligence model configured to generate code to create a structured database procedure. Further exemplary systems, methods, and computer-accessible medium are provided that can train an electronic health records (EHR) artificial intelligence model on a training data set that comprises a plurality of EHR records utilizing an under-sampling technique.