Smart Data Selection for Deep Learning Retraining

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current deep learning models in medical imaging face challenges in efficiently prioritizing training data sets for retraining, leading to suboptimal performance due to resource-intensive data preparation and annotation processes, and the need to identify the most impactful data sets for improving model accuracy.

Innovation Solution

A method and system for smart data selection that uses clinically driven evaluation metrics to flag and prioritize data sets based on their impact on clinical outcomes, focusing on the most challenging data sets for retraining, and assigning them to expert annotators for efficient preprocessing and annotation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If all collected data sets are used for retraining deep learning models, then model performance may be improved, but resource consumption and time cost increase significantly

Engineering Contradiction:
Improvemodel performanceVSAvoidretraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent extracts and identifies only the most valuable data sets for retraining by computing importance scores based on model performance degradation. Instead of using all collected data, the system selectively extracts the subset of data sets that contribute most to model improvement, thereby reducing retraining time while maintaining performance gains.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies local quality by differentiating between data sets based on their individual importance to model performance. Each data set is evaluated and assigned a different priority level, allowing the system to focus computational resources on high-impact data sets rather than treating all data uniformly.

Inventive Principle:
Principle #3Local quality

2Manufacturing precision

If all collected data sets are annotated by expert annotators, then data quality improves, but annotation cost and time increase significantly

Engineering Contradiction:
Improvedata qualityVSAvoidannotation time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent extracts only the most important data sets that require expert annotation by computing importance scores. Data sets with low importance scores are excluded from the annotation queue, allowing expert annotators to focus their time and effort on data sets that will have the greatest impact on model performance.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by annotating only a selective subset of data sets rather than all collected data. This partial annotation approach is sufficient to maintain model performance while significantly reducing the time and resource investment required for full data annotation.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If data sets are prioritized based on clinical outcome metrics, then retraining efficiency improves, but the complexity of data selection increases

Engineering Contradiction:
Improveretraining efficiencyVSAvoiddata selection complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary mechanism - an automated importance scoring system that computes data set priorities based on model performance degradation metrics. This intermediary automatically ranks data sets, replacing complex manual selection processes and reducing the perceived complexity for users while maintaining scientific rigor.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements feedback by computing importance scores based on observed model performance degradation when specific data sets are excluded. This feedback loop automatically identifies which data sets are most critical, allowing the system to adaptively prioritize data selection without requiring complex predefined criteria.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20230259820A1Smart selection to prioritize data collection and annotation based on clinical metrics
Publication Date: 2023.08.17 SIEMENS HEALTHINEERS AG
  • US20230259820A1 patent drawing
  • US20230259820A1 patent drawing
  • US20230259820A1 patent drawing

AI summary

Systems and methods for smart selection of training data sets by using clinically driven application dependent evaluation metrics to assess the performance of deep learning models after deployment in the field. A machine trained model is deployed to a clinical environment. An evaluation metric is acquired that correlates with a clinical outcome for each instance of the machine trained model performing the task for a medical procedure. Data sets are flagged that are challenging for the machine trained model based on the evaluation metrics. The flagged data sets are prioritized during retraining of the machine trained model.