Smart Data Selection for Deep Learning Retraining
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current deep learning models in medical imaging face challenges in efficiently prioritizing training data sets for retraining, leading to suboptimal performance due to resource-intensive data preparation and annotation processes, and the need to identify the most impactful data sets for improving model accuracy.
Innovation Solution
A method and system for smart data selection that uses clinically driven evaluation metrics to flag and prioritize data sets based on their impact on clinical outcomes, focusing on the most challenging data sets for retraining, and assigning them to expert annotators for efficient preprocessing and annotation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all collected data sets are used for retraining deep learning models, then model performance may be improved, but resource consumption and time cost increase significantly
Solution Approach 1:
The patent extracts and identifies only the most valuable data sets for retraining by computing importance scores based on model performance degradation. Instead of using all collected data, the system selectively extracts the subset of data sets that contribute most to model improvement, thereby reducing retraining time while maintaining performance gains.
Solution Approach 2:
The patent applies local quality by differentiating between data sets based on their individual importance to model performance. Each data set is evaluated and assigned a different priority level, allowing the system to focus computational resources on high-impact data sets rather than treating all data uniformly.
2Manufacturing precision
If all collected data sets are annotated by expert annotators, then data quality improves, but annotation cost and time increase significantly
Solution Approach 1:
The patent extracts only the most important data sets that require expert annotation by computing importance scores. Data sets with low importance scores are excluded from the annotation queue, allowing expert annotators to focus their time and effort on data sets that will have the greatest impact on model performance.
Solution Approach 2:
The patent applies partial action by annotating only a selective subset of data sets rather than all collected data. This partial annotation approach is sufficient to maintain model performance while significantly reducing the time and resource investment required for full data annotation.
3Productivity
If data sets are prioritized based on clinical outcome metrics, then retraining efficiency improves, but the complexity of data selection increases
Solution Approach 1:
The patent introduces an intermediary mechanism - an automated importance scoring system that computes data set priorities based on model performance degradation metrics. This intermediary automatically ranks data sets, replacing complex manual selection processes and reducing the perceived complexity for users while maintaining scientific rigor.
Solution Approach 2:
The patent implements feedback by computing importance scores based on observed model performance degradation when specific data sets are excluded. This feedback loop automatically identifies which data sets are most critical, allowing the system to adaptively prioritize data selection without requiring complex predefined criteria.
Data Source
AI summary
Systems and methods for smart selection of training data sets by using clinically driven application dependent evaluation metrics to assess the performance of deep learning models after deployment in the field. A machine trained model is deployed to a clinical environment. An evaluation metric is acquired that correlates with a clinical outcome for each instance of the machine trained model performing the task for a medical procedure. Data sets are flagged that are challenging for the machine trained model based on the evaluation metrics. The flagged data sets are prioritized during retraining of the machine trained model.


