Pathology Slide Abnormal Region Detection Using Synthetic Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing pathological slide image analysis algorithms face degradation due to errors, making it difficult to accurately extract or predict patient information, and manual error detection is challenging due to the rarity of such errors.
Innovation Solution
A method and system for training a machine learning model to detect abnormal regions in pathological slide images by generating training data that includes normal and abnormal regions through image processing, using a machine learning model trained on these sets to identify abnormality scores and regions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual error detection is performed on pathological slide images, then detection accuracy may be maintained, but the process becomes extremely time-consuming and impractical due to the rarity of errors
Solution Approach 1:
The patent applies preliminary action by pre-processing pathological slide images to generate synthetic error data before actual error detection. The system creates training datasets with artificially introduced errors (blur, noise, artifacts) so that the machine learning model can learn error patterns in advance, enabling efficient real-time detection without manual inspection of every image.
Solution Approach 2:
The patent uses copying by creating synthetic copies of normal pathological images and introducing artificial errors to these copies. This generates a large volume of training data with known error types, allowing the model to learn from these copied examples without requiring manual annotation of rare actual errors, thus maintaining detection accuracy while reducing time investment.
2Productivity
If machine learning models are trained on limited actual error data, then training time is reduced, but detection performance degrades due to insufficient learning samples
Solution Approach 1:
The system applies self-service by automatically generating its own training data through synthetic error introduction. The machine learning model's training process is self-sufficient as it creates its own labeled error examples by processing normal images and adding artificial errors, eliminating the need for external manual annotation and enabling continuous model improvement without additional human resources.
Solution Approach 2:
The patent applies parameter changes by systematically varying error parameters (blur radius, noise intensity, artifact type and position) when generating synthetic training data. This creates diverse training examples that cover a wide range of error conditions, enabling the model to learn robust error detection across different scenarios while maintaining efficient training through automated parameter manipulation.
3Loss of information
If analysis algorithms process all regions of pathological slide images, then comprehensive analysis is achieved, but computation time increases significantly due to error regions degrading algorithm performance
Solution Approach 1:
The patent applies the extraction principle by using the trained machine learning model to identify and extract error regions from pathological slide images. Once error regions are detected, the system can exclude these regions from subsequent analysis algorithms, processing only the clean, error-free regions. This maintains information completeness for valid regions while reducing computation time by avoiding processing of degraded error areas.
Data Source
AI summary
A method, performed by at least one processor, for training a machine learning model for detecting an abnormal region in a pathological slide image is disclosed. The method including receiving one or more first pathological slide images, determining, from the received one or more first pathological slide images, a normal region based on an abnormality condition indicative of a condition of an abnormal region, generating a first set of training data including the determined normal region, generating the abnormal region by performing image processing corresponding to the abnormality condition with respect to at least partial region in the received one or more first pathological slide images, and generating a second set of training data including the generated abnormal region.


