Style-Matched Image Generation for Small Dataset Expansion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computing systems face inefficiencies and inaccuracies in generating and managing image datasets due to small data problems, where insufficient training images hinder the accurate training of image-based machine-learning models, and combining small datasets with larger ones often results in models fitting to the larger datasets while discounting the original small dataset.

Innovation Solution

The style matching system utilizes a generative machine-learning model to expand a small set of input images by selecting images from a catalog of stored image sets based on style distribution, generating a larger dataset of synthesized images that match the style and content of the initial small set, and conditionally sampling images to ensure they align with the original input images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a small image dataset is used for training, then training time and computational resources are reduced, but the model cannot be accurately trained due to insufficient training images

Engineering Contradiction:
Improvetraining efficiencyVSAvoidmodel accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent uses a generative machine-learning model to create synthetic copies of images from the small input dataset. These generated images replicate the style and content characteristics of the original images, effectively multiplying the training data without requiring additional real images or manual data collection

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system modifies the parameters of the generative model during training to optimize the generation process. By adjusting model parameters and using conditional sampling based on the original input images, the system ensures that generated images maintain the desired style and content distribution while providing sufficient variety for robust model training

Inventive Principle:
Principle #35Parameter changes

2Loss of time

If a small image dataset is used for training, then data processing time is reduced, but the model lacks image diversity needed for robust training

Engineering Contradiction:
Improvedata processing timeVSAvoidmodel robustness
Core Design Contradiction:
Loss of timeVSAdaptability or versatility

Solution Approach 1:

The generative model creates diverse synthetic images that preserve the stylistic and content characteristics of the original dataset. This copying approach generates unlimited variations of training images without requiring additional data collection or processing time

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system employs dynamic conditional sampling during the generation process, where the generative model is conditioned on the original input images to produce diverse outputs. This dynamic approach ensures that the generated dataset maintains the essential characteristics of the original while providing the diversity needed for robust model training

Inventive Principle:
Principle #15Dynamics

3Quantity of substance

If small datasets are combined with larger datasets, then training data quantity increases, but models produce results fitted to the larger datasets while discounting the original small dataset

Engineering Contradiction:
Improvetraining data quantityVSAvoidstyle matching accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

Instead of combining with external larger datasets, the system copies and generates images that are specifically tailored to match the style and content of the original small dataset. This ensures the generated images preserve the unique characteristics of the original data without introducing external style biases

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The generative machine-learning model acts as an intermediary between the small input dataset and the training process. It transforms the limited original images into a expanded training dataset that maintains fidelity to the original style and content, serving as a bridge that preserves the original dataset's characteristics while providing sufficient training data

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250292539A1Generating large datasets of style-specific and content-specific images using generative machine-learning models to match a small set of sample images
Publication Date: 2025.09.18 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250292539A1 patent drawing
  • US20250292539A1 patent drawing
  • US20250292539A1 patent drawing

AI summary

The present disclosure relates to utilizing a style-matching image generation system to generate large datasets of style-matching images having matching styles and content to an initial small sample set of input images. For example, the style-matching image generation system utilizes a selection of style-mixed stored images with a generative machine-learning model to produce large datasets of synthesized images. Further, the style-matching image generation system utilizes the generative machine-learning model to conditionally sample synthesized images that accurately match the style, content, characteristics, and patterns of the initial small sample set and that also provide added variety and diversity to the large image dataset.