Clinical NLP Data Augmentation Through Entity and Assertion Replacement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Clinical natural language processing models require extensive and accurately annotated training data, which is difficult to obtain due to the expertise needed, leading to inaccurate and inefficient model performance.

Innovation Solution

A system and method for generating augmented medical data by replacing words associated with assertion labels and/or entity labels in annotated medical data, allowing for the expansion of training data sets using various learning techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If extensive annotated training data is used to train clinical NLP models, then model accuracy is improved, but the time and resources required for data annotation increase significantly

Engineering Contradiction:
Improvemodel accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by automatically generating augmented training data through entity replacement and text transformation before the actual model training begins. This pre-processing step creates synthetic training samples that reduce the need for extensive manual annotation later, thereby improving model accuracy while minimizing annotation time investment.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates copies of existing annotated medical data by replacing entities with synonyms or related terms while preserving the annotation structure. This copying mechanism generates additional training samples without requiring new manual annotation, thus improving model accuracy while avoiding proportional increases in annotation time.

Inventive Principle:
Principle #26Copying

2Reliability

If domain experts annotate training data to ensure accuracy, then data quality is improved, but the cost and complexity of the annotation process increase

Engineering Contradiction:
Improvedata qualityVSAvoidannotation process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system enables self-service by automatically generating augmented training data through computational methods rather than relying on manual domain expert annotation. The automated entity replacement and text transformation processes maintain data quality while eliminating the complexity associated with coordinating and managing expert annotators.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes parameters of the training data by transforming text while preserving semantic meaning and annotation integrity. Through controlled entity replacement and text variation, the system generates diverse training samples that maintain high data quality without requiring complex human annotation processes.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If more training data is generated manually, then model generalization is improved, but the productivity of the annotation process decreases

Engineering Contradiction:
Improvemodel generalizationVSAvoidannotation productivity
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system performs preliminary data augmentation by automatically generating diverse training samples through entity replacement and text transformation before model training. This pre-computation approach improves model generalization capabilities while maintaining high productivity, as the automated process can generate large volumes of training data rapidly without manual intervention.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system enhances model generalization by changing text parameters through synonym replacement, entity substitution, and linguistic transformations. These parameter changes create diverse training samples that improve model adaptability to various input formulations while the automated process maintains high annotation productivity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250266136A1System and method for data augmentation for clinical natural language processing
Publication Date: 2025.08.21 GE PRECISION HEALTHCARE LLC
  • US20250266136A1 patent drawing
  • US20250266136A1 patent drawing
  • US20250266136A1 patent drawing

AI summary

Various systems and methods are provided for generating augmented medical data. Annotated medical data including text and annotations of the text including an entity label and an assertion label may be received. A first set of words related to the assertion label may be replaced with a second set of words. A first entity in the text may be replaced with a second entity related to the entity label. Augmented medical data may be generated based on replacing the first set of words and/or replacing the first entity with the second entity. A computer executed task may be performed using the augmented medical data.