Image Embedding Clustering for Training Data Quality Assessment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The preparation and curation of high-quality training data for machine learning models are time-consuming and challenging, especially in assessing the effectiveness of training data before use.
Innovation Solution
A computing device and method that generate image embeddings for medical scene images, cluster them based on embeddings, and evaluate the quality of the image set using a parameter space trajectory, allowing for adaptation of the image set to improve its quality for training machine learning models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If training data are offered and sold in large bulks, then the quantity of training data increases, but the true effectiveness is difficult to assess before actually using the training data
Solution Approach 1:
The patent performs preliminary assessment of training data quality before the actual training process by generating image embeddings and analyzing clustering trajectories. This allows users to evaluate the effectiveness of training data in advance, without needing to actually use it for training a machine learning model first.
Solution Approach 2:
The patent introduces image embeddings as an intermediary representation that bridges the gap between raw images and training effectiveness assessment. By transforming images into embedding space and analyzing clustering patterns, the system provides an indirect but reliable measure of training data quality.
2Reliability
If high quality training data are prepared and curated by human personnel, then the quality of training data improves, but the time required for preparation increases
Solution Approach 1:
The patent replaces manual human assessment and curation of training data with an automated computational system. The system uses machine learning models to generate image embeddings and algorithmically analyze clustering trajectories, substituting human time and effort with automated processing.
Solution Approach 2:
The training data assessment system is self-service in nature, automatically evaluating the quality of training data without requiring human intervention. The system performs embedding generation, clustering analysis, and trajectory evaluation autonomously, allowing training data to be assessed independently of human resources.
Data Source
AI summary
A computing device includes an image embeddings generating module configured to generate a data array as an image embedding for each image received; a clustering module configured to determine, a plurality of clustering parameter values within the images based on the generated image embeddings; an evaluation module configured to construct a trajectory in a parameter space, wherein one dimension of the parameter space represents the plurality of clustering parameter values and another dimension of the parameter space is based on the number of clusters determined by the clustering module; wherein the evaluation module is further configured to determine a measure of the parameter space between the origin of the parameter space and the trajectory; and a user interface configured to to indicate changes and/or effects of the user input on/in the measure.


