Embedding-Space Data Diagnosis for Synthetic Training Data Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for evaluating the quality of data for deep learning models are limited, particularly for unstructured data, and there is a need for a comprehensive solution that can accurately assess and modify data quality to enhance training efficiency.
Innovation Solution
A computing device and method for identifying and modifying data points in an embedding space to generate high-quality synthetic data, providing a Modified Image of Data (MIOD) through data clinic techniques, including data imaging, modification, and evaluation algorithms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If synthetic data is generated indiscriminately to increase data quantity, then the amount of training data increases, but the quality of training data deteriorates
Solution Approach 1:
The patent changes the parameters of data points in the embedding space by adjusting distribution properties (e.g., density, clustering characteristics) to generate modified data point sets. This allows systematic modification of data characteristics while maintaining quality standards, resolving the contradiction between quantity and quality of training data.
Solution Approach 2:
The patent replaces traditional mechanical data collection and annotation methods with an AI-based system that automatically generates synthetic data by manipulating embedding space representations. This substitution enables scalable data generation with controlled quality through algorithmic processes rather than manual methods.
2Reliability
If commercial data quality verification methods are used, then structured data integrity can be verified, but unstructured data quality assessment is insufficient
Solution Approach 1:
The patent creates a universal data quality assessment system that handles both structured and unstructured data through a common embedding space framework. By representing different data types in a unified mathematical space, the system achieves versatility across data types while maintaining reliable quality verification through distribution analysis.
3Quantity of substance
If data points are densely distributed in embedding space, then data coverage is improved, but data discrimination capability deteriorates
Solution Approach 1:
The patent applies local quality modification by adjusting the distribution properties of data points in specific regions of the embedding space. Different areas of the embedding space can have different density characteristics, allowing dense coverage in some regions while maintaining discrimination capability in others through localized distribution control.
Data Source
AI summary
According to an embodiments of the present disclosure, a method comprising: at an electronic device with one or more processors, obtaining a data set; identifying, based on the data set, a first data point set on a first embedding space, wherein each data point included in the first data point set corresponds to each data included in the data set; identifying a modified first data point set on the first embedding space based on the first data point set by adjusting a property associated with a distribution of the first data point set, wherein the modified first data point set includes at least one modified data point which is not included in the first data point set; and providing a Modified Image of Data (MIOD) by representing the modified first data point set on an imaging space may be provided.


