Feature Extraction for Privacy-Preserving Medical Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Federated learning is challenging to implement in medical institutions that prioritize personal information protection, as it requires data sharing and synchronization across different centers, which is difficult due to security and environmental restrictions.
Innovation Solution
A machine learning method and apparatus that receives and processes feature data from medical images using a basic model, performing lossy compression to remove personal information and generate a final machine learning model, allowing for improved performance without sharing sensitive data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If federated learning is implemented to improve machine learning performance using data from multiple medical centers, then model performance is improved, but implementation difficulty increases due to security restrictions and inability to share data across institutions
Solution Approach 1:
The patent extracts only the essential feature data from medical images using a basic model, rather than sharing the complete original medical images. This extracted feature data is then used for training the final machine learning model, enabling multi-center collaboration while maintaining data security and privacy protection within each institution.
Solution Approach 2:
The patent introduces an intermediary processing step where a basic model transforms original medical images into feature data before sharing. This intermediary representation serves as a bridge that allows information exchange between institutions without directly sharing sensitive original data, thus resolving the contradiction between collaboration needs and security restrictions.
2Measurement precision
If original medical images are shared across institutions to improve model training, then machine learning performance is improved, but personal information security deteriorates
Solution Approach 1:
The patent extracts only the essential feature data from medical images using a basic model, rather than sharing the complete original medical images. This extracted feature data is then used for training the final machine learning model, enabling multi-center collaboration while maintaining data security and privacy protection within each institution.
Solution Approach 2:
The patent applies different processing quality levels to different data representations. Original medical images maintain their full quality and detail for local use only, while extracted feature data provides a lower-quality, anonymized representation suitable for sharing. This local quality differentiation allows simultaneous achievement of high-performance local analysis and secure collaborative training.
3Object-affected harmful factors
If data is compressed to remove personal information, then information security is improved, but data quality deteriorates
Solution Approach 1:
The patent introduces an intermediary processing step where a basic model transforms original medical images into feature data before sharing. This intermediary representation serves as a bridge that allows information exchange between institutions without directly sharing sensitive original data, thus resolving the contradiction between collaboration needs and security restrictions.
Solution Approach 2:
The patent changes the parameter representation of medical data from pixel-based images to feature-based representations. This parameter transformation maintains the essential diagnostic information needed for machine learning while removing personally identifiable information, thus achieving both security protection and data quality preservation.
Data Source
AI summary
There is provided a method and apparatus that collects feature points of data and performs machine learning. A machine learning method comprises receiving first feature data obtained by applying a basic model to first analysis target data, receiving second feature data obtained by applying the basic model to second analysis target data, and obtaining a final machine learning model through performing machine learning on a correlation between the first feature data and first analysis result data and a correlation between the second feature data and second analysis result data.


