Dataset Diagnostic System for Synthetic Data Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for evaluating the quality of data used in deep learning models are limited, as they primarily verify the integrity of structured data and lack a comprehensive approach applicable across various technical fields, necessitating a solution that can assess the quality of unstructured data effectively.
Innovation Solution
A computing device and method that identify essential characteristics of a data set by mapping it to a latent space, adjusting data point distributions, and generating a Modified Image of Data (MIOD) to provide diagnostic insights and improve data quality for deep learning model training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If commercial data quality verification methods are used, then structured data integrity can be verified, but they cannot effectively assess unstructured data quality
Solution Approach 1:
The patent transforms data quality assessment from traditional structured data verification parameters to embedding space distribution parameters. By mapping data to embedding spaces and analyzing distribution characteristics (density, clustering, uniformity), the system adapts quality assessment to work effectively across both structured and unstructured data types while maintaining precise evaluation capabilities.
Solution Approach 2:
The patent introduces a new dimension for data quality assessment by utilizing embedding spaces. Instead of verifying data quality in the original data structure, the system projects data into embedding spaces where quality can be evaluated through distribution patterns, providing a versatile framework that works for various data types while preserving assessment precision.
2Quantity of substance
If synthetic data is generated indiscriminately, then data quantity increases, but data quality deteriorates
Solution Approach 1:
The patent implements a feedback mechanism where the quality assessment system continuously evaluates synthetic data in embedding spaces and provides guidance for generating higher quality synthetic data. By measuring distribution characteristics and comparing against quality standards, the system ensures that data quantity increases while maintaining or improving data quality through iterative refinement.
Solution Approach 2:
The system uses parameter changes in embedding space distributions to control synthetic data generation quality. By adjusting and optimizing distribution parameters (density, clustering patterns, uniformity) during synthetic data generation, the system maintains high data quality while increasing data quantity, avoiding the deterioration that occurs with indiscriminate generation.
3Productivity
If data distribution is not optimized, then data processing is simpler, but deep learning model performance suffers
Solution Approach 1:
The patent applies preliminary action by optimizing data distribution in embedding spaces before deep learning model training. The system performs embedding mapping and distribution optimization as pre-processing steps, arranging data in optimal configurations that enhance model training effectiveness. This preliminary optimization of data layout and distribution simplifies the actual training process while significantly improving productivity.
Data Source
AI summary
According to an embodiments of the present disclosure, a computer-implemented method comprising: obtaining, by one or more processors, a first data set; identifying, by one or more processors, a first data point set by determining at least one feature of the first data set from at least one layer of a first trained model, wherein the first data point set corresponding to the first data set is associated with a first embedding space of a first dimension; obtaining, by one or more processors, a first diagnostic data corresponding to the first data set based on the first data point set by analyzing at least one property of the first data set; and generating, by one or more processors, a first set of synthetic data, wherein the generating the first set of synthetic data comprises: inputting a prompt data associated with the at least one property of the first data set into a second trained model; and obtaining the first set of synthetic data from at least one layer of the second trained model may be provided.


