Machine Learning Supervision for Medical Images Using Clinical Reports

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning systems face challenges in obtaining and labeling training data, particularly in medical imaging, due to the high cost and expertise required, leading to inaccurate results and failure to converge.

Innovation Solution

A supervised visual grounding method using cross-modal transformers with noise contrastive estimation, cross-modal feature-wise linear modulation, and recursive textual mining to generate annotations from clinical reports, enabling the use of unlabeled medical datasets for training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If expert-labeled annotations are used for training, then measurement precision is improved, but loss of time and productivity deteriorate due to high cost and expertise requirements

Engineering Contradiction:
Improvelocalization accuracyVSAvoidtime for data labeling
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by using a first neural network to generate predicted annotations from clinical reports before the main detection network training. These pre-generated annotations serve as preliminary labeled data that can be used to train the detection network, eliminating the need for time-consuming expert labeling while maintaining training quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The first neural network acts as an intermediary that translates clinical reports into predicted annotations. This intermediary component bridges the gap between unstructured clinical text and structured detection training data, enabling the system to leverage existing clinical reports without requiring direct expert annotation of images.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If expert-labeled training data is obtained, then manufacturing precision is improved, but productivity deteriorates due to expensive experiments and expertise requirements

Engineering Contradiction:
Improvedetection performanceVSAvoiddata labeling efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system implements self-service by enabling the detection network to train on its own using predicted annotations generated from clinical reports. The workflow allows the system to automatically create its training data from available clinical reports, eliminating the need for external expert annotation services and significantly improving data labeling productivity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The first neural network serves multiple functions: it generates predicted annotations for training the detection network, validates image-report pairs, and can potentially be used for other annotation tasks. This multi-functionality increases overall system productivity by eliminating the need for separate expert annotation processes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If more training data is collected, then reliability is improved, but loss of time increases due to data collection and labeling processes

Engineering Contradiction:
Improvesystem convergenceVSAvoiddata collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-processing and validating clinical reports to generate predicted annotations before they are used for training. This preliminary processing ensures that the training data is ready-to-use and of high quality, allowing the system to achieve reliable convergence without the time-consuming process of collecting and manually labeling additional training data.

Inventive Principle:
Principle #10Preliminary action

4Measurement precision

If expert annotation is used, then measurement precision is improved, but device complexity increases due to the need for expert systems

Engineering Contradiction:
Improveannotation accuracyVSAvoidsystem architecture
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The first neural network serves as an intermediary that automatically generates predicted annotations from clinical reports, replacing the need for complex expert annotation systems. This intermediary component simplifies the overall system architecture by automating the annotation process while maintaining measurement precision through the use of trained neural network models.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system replaces the mechanical process of expert human annotation with an automated neural network-based system. The first neural network automatically translates clinical reports into structured annotations, substituting the manual, expertise-dependent process with an automated computational process that reduces system complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12367261B1Extending supervision using machine learning
Publication Date: 2025.07.22 NVIDIA CORP
  • US12367261B1 patent drawing
  • US12367261B1 patent drawing
  • US12367261B1 patent drawing

AI summary

Apparatuses, systems, and techniques to generate labeled training data. In at least one embodiment, labeled training images are generated from medial images annotated with natural language text.