Dataset Diagnostic System for Synthetic Data Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for evaluating the quality of data used in deep learning models are limited, as they primarily verify the integrity of structured data and lack a comprehensive approach applicable across various technical fields, necessitating a solution that can assess the quality of unstructured data effectively.

Innovation Solution

A computing device and method that identify essential characteristics of a data set by mapping it to a latent space, adjusting data point distributions, and generating a Modified Image of Data (MIOD) to provide diagnostic insights and improve data quality for deep learning model training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If commercial data quality verification methods are used, then structured data integrity can be verified, but they cannot effectively assess unstructured data quality

Engineering Contradiction:
Improvedata quality assessment applicabilityVSAvoiddata quality evaluation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent transforms data quality assessment from traditional structured data verification parameters to embedding space distribution parameters. By mapping data to embedding spaces and analyzing distribution characteristics (density, clustering, uniformity), the system adapts quality assessment to work effectively across both structured and unstructured data types while maintaining precise evaluation capabilities.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces a new dimension for data quality assessment by utilizing embedding spaces. Instead of verifying data quality in the original data structure, the system projects data into embedding spaces where quality can be evaluated through distribution patterns, providing a versatile framework that works for various data types while preserving assessment precision.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Quantity of substance

If synthetic data is generated indiscriminately, then data quantity increases, but data quality deteriorates

Engineering Contradiction:
Improvedata quantityVSAvoiddata quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent implements a feedback mechanism where the quality assessment system continuously evaluates synthetic data in embedding spaces and provides guidance for generating higher quality synthetic data. By measuring distribution characteristics and comparing against quality standards, the system ensures that data quantity increases while maintaining or improving data quality through iterative refinement.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system uses parameter changes in embedding space distributions to control synthetic data generation quality. By adjusting and optimizing distribution parameters (density, clustering patterns, uniformity) during synthetic data generation, the system maintains high data quality while increasing data quantity, avoiding the deterioration that occurs with indiscriminate generation.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If data distribution is not optimized, then data processing is simpler, but deep learning model performance suffers

Engineering Contradiction:
Improvedeep learning model training effectivenessVSAvoiddata processing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by optimizing data distribution in embedding spaces before deep learning model training. The system performs embedding mapping and distribution optimization as pre-processing steps, arranging data in optimal configurations that enhance model training effectiveness. This preliminary optimization of data layout and distribution simplifies the actual training process while significantly improving productivity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240086493A1Method for diagnosing a dataset to generate synthetic data, and a computing device and system for performing such a method
Publication Date: 2024.03.14 PEBBLOUS INC
  • US20240086493A1 patent drawing
  • US20240086493A1 patent drawing
  • US20240086493A1 patent drawing

AI summary

According to an embodiments of the present disclosure, a computer-implemented method comprising: obtaining, by one or more processors, a first data set; identifying, by one or more processors, a first data point set by determining at least one feature of the first data set from at least one layer of a first trained model, wherein the first data point set corresponding to the first data set is associated with a first embedding space of a first dimension; obtaining, by one or more processors, a first diagnostic data corresponding to the first data set based on the first data point set by analyzing at least one property of the first data set; and generating, by one or more processors, a first set of synthetic data, wherein the generating the first set of synthetic data comprises: inputting a prompt data associated with the at least one property of the first data set into a second trained model; and obtaining the first set of synthetic data from at least one layer of the second trained model may be provided.