Machine Learning Model for Pathology Image Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for analyzing pathology images using artificial intelligence algorithms face challenges due to the need for large amounts of labeled training data, which is costly and time-consuming to obtain, especially for new biomarkers and low-prevalence cancer types.
Innovation Solution
A method and system that utilize a machine learning model trained on a training data set generated from heterogeneous pathology data sets, allowing for accurate analysis of various types of pathology images without the need for extensive retraining for new image types.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If artificial intelligence algorithms are trained using pathology images labeled with medical knowledge, then accurate prediction results are achieved, but cost and time increase due to the need for medical experts to build training data
Solution Approach 1:
The training data is segmented into two distinct domains: first domain data (e.g., PD-L1 IHC stained images) and second domain data (e.g., H&E stained images or new biomarker images). The machine learning model is trained to learn domain-specific features separately, then integrated to achieve accurate predictions while reducing the need for extensive labeled data in each individual domain.
Solution Approach 2:
The machine learning model is designed with multi-functionality to handle multiple types of pathology images from different domains. By training on heterogeneous data sets simultaneously, the model becomes universal and can accurately predict results for various cancer types and staining methods without requiring separate training processes for each domain.
2Adaptability or versatility
If artificial intelligence algorithms are trained using pathology images from new biomarkers, then analysis capability for new targets is achieved, but sufficient training data cannot be ensured in a short period due to limited clinical data
Solution Approach 1:
The patent merges data from multiple domains into a unified training data set. By combining first domain pathology data (which may have sufficient clinical data) with second domain pathology data (new biomarkers with limited data), the model learns transferable features that enable accurate analysis of new biomarkers even when training data for those specific biomarkers is scarce.
Solution Approach 2:
The machine learning model performs preliminary learning on well-established biomarkers with abundant clinical data (first domain), acquiring general pathology image recognition capabilities. This preliminary training enables the model to quickly adapt to new biomarkers with limited data, as the foundational skills are already established.
3Productivity
If artificial intelligence model is trained using small amount of data for low-prevalence cancer types, then model training is completed faster, but the model may not be properly trained or becomes biased toward specific training data set
Solution Approach 1:
The first domain pathology data acts as an intermediary that bridges the gap for low-prevalence cancer types. By training on this larger, more diverse data set first, the model develops robust generalization capabilities that prevent overfitting to small data sets, while still enabling faster training compared to traditional methods that would require extensive data collection for rare cancers.
Data Source
AI summary
Provided is a method for analysing a pathology image, which is performed by at least one processor and includes acquiring a pathology image, inputting the acquired pathology image into a machine learning model and acquiring an analysis result for the pathology image from the machine learning model, and outputting the acquired analysis result, in which the machine learning model is a model trained by using a training data set generated based on a first pathology data set associated with a first domain and a second pathology data set associated with a second domain different from the first domain.


