Machine Learning Supervision for Medical Images Using Clinical Reports
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning systems face challenges in obtaining and labeling training data, particularly in medical imaging, due to the high cost and expertise required, leading to inaccurate results and failure to converge.
Innovation Solution
A supervised visual grounding method using cross-modal transformers with noise contrastive estimation, cross-modal feature-wise linear modulation, and recursive textual mining to generate annotations from clinical reports, enabling the use of unlabeled medical datasets for training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If expert-labeled annotations are used for training, then measurement precision is improved, but loss of time and productivity deteriorate due to high cost and expertise requirements
Solution Approach 1:
The system performs preliminary action by using a first neural network to generate predicted annotations from clinical reports before the main detection network training. These pre-generated annotations serve as preliminary labeled data that can be used to train the detection network, eliminating the need for time-consuming expert labeling while maintaining training quality.
Solution Approach 2:
The first neural network acts as an intermediary that translates clinical reports into predicted annotations. This intermediary component bridges the gap between unstructured clinical text and structured detection training data, enabling the system to leverage existing clinical reports without requiring direct expert annotation of images.
2Manufacturing precision
If expert-labeled training data is obtained, then manufacturing precision is improved, but productivity deteriorates due to expensive experiments and expertise requirements
Solution Approach 1:
The system implements self-service by enabling the detection network to train on its own using predicted annotations generated from clinical reports. The workflow allows the system to automatically create its training data from available clinical reports, eliminating the need for external expert annotation services and significantly improving data labeling productivity.
Solution Approach 2:
The first neural network serves multiple functions: it generates predicted annotations for training the detection network, validates image-report pairs, and can potentially be used for other annotation tasks. This multi-functionality increases overall system productivity by eliminating the need for separate expert annotation processes.
3Reliability
If more training data is collected, then reliability is improved, but loss of time increases due to data collection and labeling processes
Solution Approach 1:
The system performs preliminary action by pre-processing and validating clinical reports to generate predicted annotations before they are used for training. This preliminary processing ensures that the training data is ready-to-use and of high quality, allowing the system to achieve reliable convergence without the time-consuming process of collecting and manually labeling additional training data.
4Measurement precision
If expert annotation is used, then measurement precision is improved, but device complexity increases due to the need for expert systems
Solution Approach 1:
The first neural network serves as an intermediary that automatically generates predicted annotations from clinical reports, replacing the need for complex expert annotation systems. This intermediary component simplifies the overall system architecture by automating the annotation process while maintaining measurement precision through the use of trained neural network models.
Solution Approach 2:
The system replaces the mechanical process of expert human annotation with an automated neural network-based system. The first neural network automatically translates clinical reports into structured annotations, substituting the manual, expertise-dependent process with an automated computational process that reduces system complexity.
Data Source
AI summary
Apparatuses, systems, and techniques to generate labeled training data. In at least one embodiment, labeled training images are generated from medial images annotated with natural language text.


