Automated Learning Data Generation for Medical Image Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In the medical field, acquiring large amounts of high-quality data necessary for deep learning is challenging due to the vast amount of data stored in medical image management systems, making it irrational to manually discriminate correct answer data, and existing systems lack efficiency in automating the process for feature recognition in medical images.
Innovation Solution
A learning data generation support apparatus and method that analyzes character strings in interpretation reports to extract lesion portion information and features, performs image analysis to acquire corresponding features, and determines matching features to register image data as correct answer data for deep learning, utilizing natural language processing and image recognition techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual discrimination of correct answer data is performed from large amount of medical data, then data quality can be ensured, but productivity is reduced and time consumption increases
Solution Approach 1:
The patent introduces an automated system as an intermediary between the large amount of medical image data and the deep learning model. This system includes: (1) a data extraction unit that automatically extracts image data and interpretation reports from the medical information system, (2) a feature extraction unit that automatically extracts lesion features from interpretation reports using NLP, and (3) a data generation unit that automatically generates learning data by combining extracted images and features. This intermediary automated system resolves the contradiction by replacing manual discrimination with automatic processing, thereby ensuring both data quality and high productivity.
Solution Approach 2:
The system enables self-service by allowing the medical information system to automatically generate its own learning data without external manual intervention. The automated workflow extracts data, processes interpretation reports, generates learning data, and stores it back in the system. This self-service capability resolves the contradiction by making the data generation process autonomous, highly productive, and reliable through systematic automation.
2Measurement precision
If deep learning is performed with large amount of data, then recognition accuracy is enhanced, but data processing time and computational resources increase
Solution Approach 1:
The patent applies preliminary action by pre-processing and organizing medical image data and interpretation reports into a structured learning data format before actual deep learning training. The system extracts and stores image data with associated lesion features in advance, creating ready-to-use learning datasets. This preliminary preparation reduces the time required for actual model training and iteration, while still enabling high recognition accuracy through comprehensive data availability.
Solution Approach 2:
The system changes parameters by transforming unstructured interpretation report text into structured lesion feature parameters that can be directly used for training. This parameter transformation includes extracting lesion type, location, size, and other characteristics from free-text reports and converting them into standardized formats. This parameter change enables efficient data processing while maintaining the rich information needed for high recognition accuracy.
3Productivity
If automated system is introduced for data generation, then productivity increases, but device complexity increases
Solution Approach 1:
The patent applies universality by designing an integrated system that performs multiple functions within a single framework: data extraction from medical information systems, NLP processing of interpretation reports, feature extraction, learning data generation, and storage. This multi-functional approach increases productivity by consolidating operations while managing complexity through unified system architecture rather than separate independent components.
Solution Approach 2:
The system replaces manual mechanical processes with automated computational processes. Instead of manual data collection, extraction, and preparation, the system uses automated algorithms including NLP for text processing, image processing techniques for feature extraction, and database operations for data management. This substitution dramatically increases productivity while the complexity is managed through software automation rather than manual procedures.
Data Source
AI summary
Extraction means analyzes a character string of an interpretation report to extract lesion portion information recorded in the interpretation report and a first lesion feature recorded in the interpretation report, and analysis means performs an image analysis process corresponding to the lesion portion information with respect to the image data corresponding to the interpretation report to acquire a second lesion feature based on a result of the image analysis process. Determination means determines whether the first lesion feature matches the second lesion feature, and registration means registers the image data corresponding to the interpretation report as learning correct answer data in a case where it is determined that the first lesion feature matches the second lesion feature.


