Embedding-Space Data Diagnosis for Synthetic Training Data Quality

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for evaluating the quality of data for deep learning models are limited, particularly for unstructured data, and there is a need for a comprehensive solution that can accurately assess and modify data quality to enhance training efficiency.

Innovation Solution

A computing device and method for identifying and modifying data points in an embedding space to generate high-quality synthetic data, providing a Modified Image of Data (MIOD) through data clinic techniques, including data imaging, modification, and evaluation algorithms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If synthetic data is generated indiscriminately to increase data quantity, then the amount of training data increases, but the quality of training data deteriorates

Engineering Contradiction:
Improveamount of training dataVSAvoidquality of training data
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent changes the parameters of data points in the embedding space by adjusting distribution properties (e.g., density, clustering characteristics) to generate modified data point sets. This allows systematic modification of data characteristics while maintaining quality standards, resolving the contradiction between quantity and quality of training data.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional mechanical data collection and annotation methods with an AI-based system that automatically generates synthetic data by manipulating embedding space representations. This substitution enables scalable data generation with controlled quality through algorithmic processes rather than manual methods.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If commercial data quality verification methods are used, then structured data integrity can be verified, but unstructured data quality assessment is insufficient

Engineering Contradiction:
Improvedata integrityVSAvoidapplicability to various data types
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal data quality assessment system that handles both structured and unstructured data through a common embedding space framework. By representing different data types in a unified mathematical space, the system achieves versatility across data types while maintaining reliable quality verification through distribution analysis.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Quantity of substance

If data points are densely distributed in embedding space, then data coverage is improved, but data discrimination capability deteriorates

Engineering Contradiction:
Improvedata coverageVSAvoiddata discrimination capability
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies local quality modification by adjusting the distribution properties of data points in specific regions of the embedding space. Different areas of the embedding space can have different density characteristics, allowing dense coverage in some regions while maintaining discrimination capability in others through localized distribution control.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12481720B2Computing device that performs a method for diagnosing properties of data and a system comprising the computing device
Publication Date: 2025.11.25 PEBBLOUS INC
  • US12481720B2 patent drawing
  • US12481720B2 patent drawing
  • US12481720B2 patent drawing

AI summary

According to an embodiments of the present disclosure, a method comprising: at an electronic device with one or more processors, obtaining a data set; identifying, based on the data set, a first data point set on a first embedding space, wherein each data point included in the first data point set corresponds to each data included in the data set; identifying a modified first data point set on the first embedding space based on the first data point set by adjusting a property associated with a distribution of the first data point set, wherein the modified first data point set includes at least one modified data point which is not included in the first data point set; and providing a Modified Image of Data (MIOD) by representing the modified first data point set on an imaging space may be provided.