Image Embedding Clustering for Training Data Quality Assessment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The preparation and curation of high-quality training data for machine learning models are time-consuming and challenging, especially in assessing the effectiveness of training data before use.

Innovation Solution

A computing device and method that generate image embeddings for medical scene images, cluster them based on embeddings, and evaluate the quality of the image set using a parameter space trajectory, allowing for adaptation of the image set to improve its quality for training machine learning models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If training data are offered and sold in large bulks, then the quantity of training data increases, but the true effectiveness is difficult to assess before actually using the training data

Engineering Contradiction:
Improvequantity of training dataVSAvoidassessment of training data effectiveness
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent performs preliminary assessment of training data quality before the actual training process by generating image embeddings and analyzing clustering trajectories. This allows users to evaluate the effectiveness of training data in advance, without needing to actually use it for training a machine learning model first.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces image embeddings as an intermediary representation that bridges the gap between raw images and training effectiveness assessment. By transforming images into embedding space and analyzing clustering patterns, the system provides an indirect but reliable measure of training data quality.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If high quality training data are prepared and curated by human personnel, then the quality of training data improves, but the time required for preparation increases

Engineering Contradiction:
Improvequality of training dataVSAvoidtime for data preparation
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces manual human assessment and curation of training data with an automated computational system. The system uses machine learning models to generate image embeddings and algorithmically analyze clustering trajectories, substituting human time and effort with automated processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The training data assessment system is self-service in nature, automatically evaluating the quality of training data without requiring human intervention. The system performs embedding generation, clustering analysis, and trajectory evaluation autonomously, allowing training data to be assessed independently of human resources.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250046073A1Computing device, method and computer program
Publication Date: 2025.02.06 KARL STORZ SE & CO KG
  • US20250046073A1 patent drawing
  • US20250046073A1 patent drawing
  • US20250046073A1 patent drawing

AI summary

A computing device includes an image embeddings generating module configured to generate a data array as an image embedding for each image received; a clustering module configured to determine, a plurality of clustering parameter values within the images based on the generated image embeddings; an evaluation module configured to construct a trajectory in a parameter space, wherein one dimension of the parameter space represents the plurality of clustering parameter values and another dimension of the parameter space is based on the number of clusters determined by the clustering module; wherein the evaluation module is further configured to determine a measure of the parameter space between the origin of the parameter space and the trajectory; and a user interface configured to to indicate changes and/or effects of the user input on/in the measure.