Automating Clinical Data Standards with NLP

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current clinical data collection methods, including electronic data capture systems and electronic case report forms, still require manual data entry from electronic medical records, which is time-consuming and inefficient, especially for observational trials, and struggle with processing unstructured clinical narratives that account for a significant portion of patient care information.

Innovation Solution

A clinical data standards automated system using a machine learning model and natural language processing engine that extracts metadata from raw datasets, predicts case report form annotations, maps data against a study data tabulation model, and generates necessary artifacts with minimal user intervention, thereby automating the clinical data standards and regulatory submission process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual data entry is used from electronic medical records to electronic case report forms, then data collection can be performed with existing systems, but the process becomes time-consuming and inefficient

Engineering Contradiction:
Improvedata collection efficiencyVSAvoidtime for manual data entry
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces the mechanical manual data entry process with an automated natural language processing system. The NLP engine automatically extracts information from unstructured clinical narratives in electronic medical records and populates structured electronic case report forms, eliminating the need for manual transcription and significantly improving data collection efficiency while reducing time loss.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by allowing the data extraction and population process to occur automatically without human intervention. The NLP engine independently processes clinical narratives, identifies relevant information, and fills eCRF fields autonomously, making the system self-sufficient for the data transfer task.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If predesigned patient information templates are used to structure medical records, then data interoperability is improved, but clinician freedom of expression and researcher usability are restricted

Engineering Contradiction:
Improvedata interoperabilityVSAvoidclinician freedom of expression
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The NLP engine acts as an intermediary between unstructured clinical narratives and structured data templates. It processes free-text clinical information and automatically maps it to standardized eCRF fields, enabling data interoperability without requiring clinicians to manually adapt their documentation style to template constraints.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the state of data from unstructured text to structured format through automated processing. By transforming clinical narratives into standardized data elements, the system maintains clinician freedom to document naturally while achieving the structural organization needed for interoperability and research usability.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If automated natural language processing is used to extract data from unstructured clinical narratives, then data entry efficiency is improved, but system complexity increases

Engineering Contradiction:
Improvedata extraction efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent extracts the complex NLP processing functionality as a separate, dedicated engine that interfaces with existing EDC systems. This modular approach isolates the complexity within the NLP component while maintaining simplicity in the overall system architecture and existing infrastructure.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20240428903A1Method and system for automating clinical data standards
Publication Date: 2024.12.26 SYMBIANCE LLC
  • US20240428903A1 patent drawing
  • US20240428903A1 patent drawing
  • US20240428903A1 patent drawing

AI summary

A method and a clinical data standards (CDS) automated system are provided for automating clinical data standards and generating study data tabulation model (SDTM) artifacts required for a regulatory submission process using a machine learning model and a natural language processing (NLP) engine with minimal user intervention. The CDS automated system extracts metadata from multiple raw datasets automatically using NLP and feeds the extracted metadata into the machine learning model; predicts automatic case report form (CRF) annotations on the extracted metadata and records new learnings onto the CDS automated system; maps one or more raw datasets against a target SDTM variable; generates an SDTM statistical analysis system (SAS) code, an SDTM specification, and one or more SDTM datasets; generates a define package; validates the generated define package and the SDTM artifacts generated throughout the entire cycle; and generates validation reports in real time.