Medical Training Data Linking for Clinical Report Annotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The generation of training data for AI algorithms in medical imaging is complex, time-consuming, and often does not reflect typical clinical use cases, leading to biased and insufficiently representative data sets that can introduce regulatory challenges.

Innovation Solution

Automatically link annotations from routinely generated findings reports and other clinical data to medical images, creating quality-assured training data that include segmentation, classification, and semantic features, allowing for adaptive and comprehensive algorithm training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If training data are recorded specifically for AI training purposes with manual annotations by trained personnel, then annotation quality and completeness are improved, but the complexity and time consumption of data generation increase significantly

Engineering Contradiction:
Improveannotation qualityVSAvoiddata generation complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system enables automated self-annotation by leveraging existing clinical workflow data. Findings reports and image data are automatically linked through unique identifiers without requiring manual annotation efforts, allowing the system to generate training data autonomously from routine clinical operations

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system creates training data by copying and linking existing clinical data structures. Findings report contents are copied and associated with image data sets through unique identifiers, reproducing real clinical annotation patterns without manual intervention

Inventive Principle:
Principle #26Copying

2Manufacturing precision

If training data are recorded specifically for AI training purposes with manual annotations, then annotation quality is improved, but the time consumption for data acquisition, annotation, review and supervision increases

Engineering Contradiction:
Improveannotation qualityVSAvoiddata generation time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by automatically generating training data during routine clinical workflows before AI training begins. Findings reports and image data are linked in advance through automated processes, eliminating the need for separate time-consuming annotation phases

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system enables automated self-annotation by leveraging existing clinical workflow data. Findings reports and image data are automatically linked through unique identifiers without requiring manual annotation efforts, allowing the system to generate training data autonomously from routine clinical operations

Inventive Principle:
Principle #25Self-service

3Productivity

If selected data sets are used for training, then the training process becomes manageable, but the representativeness of training data for typical clinical use cases and patient populations deteriorates

Engineering Contradiction:
Improvetraining efficiencyVSAvoidclinical representativeness
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system achieves universality by automatically generating training data that serves multiple purposes: it represents diverse patient populations, reflects typical clinical use cases, and maintains statistical representativeness. The automated linking of findings reports with image data creates universally applicable training datasets without manual selection bias

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Reliability

If large amounts of training data are required for AI algorithm training, then model performance is improved, but the challenge of providing sufficient quality-assured and statistically representative data increases

Engineering Contradiction:
Improvemodel performanceVSAvoiddata provision challenge
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system enables automated self-annotation by leveraging existing clinical workflow data. Findings reports and image data are automatically linked through unique identifiers without requiring manual annotation efforts, allowing the system to generate training data autonomously from routine clinical operations

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system ensures continuous generation of training data by integrating with ongoing clinical workflows. As long as clinical activities continue to generate findings reports and image data, the system continuously accumulates training data without interruption or manual intervention

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS12603176B2Automated generation of medical training data for training AI-algorithms for supporting clinical reporting and documentation
Publication Date: 2026.04.14 QMEDIFY
  • US12603176B2 patent drawing
  • US12603176B2 patent drawing
  • US12603176B2 patent drawing

AI summary

A computer-implemented method, computer-system and computer-program product for generating medical training data for training artificial-intelligence (AI) algorithms for supporting clinical reporting and documentation are described. To generate the training data medical image data of a patient comprising medical image data elements are received and a medical findings report is generated, edited and/or received that summarizes individual medical findings. It comprises machine-readable findings-report elements the contents of which comprise semantic features. The contents of the findings report elements are automatically assigned to unique identifiers, wherein each identifier uniquely represents the medical semantic content of exactly one individual medical finding. The medical image data are annotated by linking one or more medical image data elements to the unique identifiers of one or more contents of the findings-report elements, and the annotated received medical image data are stored as training data for AI algorithms for supporting clinical reporting and documentation.