Synthetic Data Generation for Unsupervised Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing unsupervised learning techniques rely heavily on large-scale labeled datasets, which are costly and time-consuming to create, and are hindered by data privacy and usage rights concerns.

Innovation Solution

The proposed solution involves using model-generated data for unsupervised training, where a generative model produces sample images and attention maps based on text prompts, and a target model is trained using these synthetic data and attention maps for image processing tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If large-scale datasets are used for training, then model performance is improved, but data acquisition cost and time consumption increase

Engineering Contradiction:
Improvemodel performanceVSAvoiddata acquisition time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent uses a pre-trained generative model to synthesize training data that copies the essential statistical properties and patterns of real-world data. This synthetic data serves as a substitute for collecting and annotating large-scale real datasets, dramatically reducing data acquisition time while maintaining model training effectiveness

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The generative model is pre-trained on real data beforehand to learn data distributions and patterns. This preliminary action enables the model to subsequently generate synthetic training data without requiring time-consuming manual data collection and annotation for each training iteration

Inventive Principle:
Principle #10Preliminary action

2Reliability

If large-scale datasets are used for training, then model performance is improved, but data cost increases

Engineering Contradiction:
Improvemodel performanceVSAvoiddata cost
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

Synthetic data generated by the pre-trained generative model serves as a cost-effective substitute for expensive real-world datasets. The generative model captures data distributions and generates unlimited synthetic samples without incurring additional data collection, storage, or annotation costs

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system uses the pre-trained generative model to self-generate training data autonomously without requiring external data sources, human annotators, or expensive data procurement processes, making the training process self-sufficient and cost-effective

Inventive Principle:
Principle #25Self-service

3Reliability

If real-world data is used for training, then model performance is improved, but data privacy concerns increase

Engineering Contradiction:
Improvemodel performanceVSAvoiddata privacy concerns
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent generates synthetic data that replicates the statistical properties and patterns of real-world data without containing actual sensitive information. This copying approach maintains model training effectiveness while eliminating privacy risks associated with using real personal or sensitive data

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The pre-trained generative model acts as an intermediary between real-world data and the training process. It transforms real data patterns into synthetic representations that preserve learning value while removing privacy-sensitive information, enabling safe model training

Inventive Principle:
Principle #24Intermediary (Mediator)

4Quantity of substance

If synthetic data is used for training, then data acquisition cost is reduced, but data richness and diversity may decrease

Engineering Contradiction:
Improvedata acquisition costVSAvoiddata diversity
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The generative model uses text prompts with varying parameters (object types, attributes, relationships, scenarios) to generate diverse synthetic data samples. By systematically changing prompt parameters, the system ensures rich and varied training data coverage without requiring manual data collection

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250029289A1Unsupervised learning with synthetic data and attention masks
Publication Date: 2025.01.23 LEMON INC(GB)
  • US20250029289A1 patent drawing
  • US20250029289A1 patent drawing
  • US20250029289A1 patent drawing

AI summary

Embodiments of the disclosure relate to unsupervised learning with synthetic data and attention masks. According to example embodiments of the disclosure, a plurality of sample images are generated by providing a plurality of text prompts into a trained generative model, respectively. For a sample image of the plurality of sample images, at least one attention map is obtained from a generative model, the at least one attention map being determined by the generative model for generating the sample image, an attention map indicating visual elements of an object within the sample image. Training of a target model is performed according to unsupervised learning at least based on the plurality of sample images and attention maps for the plurality of sample images, the target model being configured to perform an image processing task.