Diffusion Model Training for Style-Aligned Synthetic Images

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models require labeled training examples, which are manually labeled and time-consuming, and synthetically generated images often do not belong to the same domain as physically captured images, causing a domain shift.

Innovation Solution

A diffusion model is trained to generate synthetic images that conform to a specified style by iteratively applying noise and optimizing parameters using a cost function, allowing the inclusion of synthetically generated images that match the domain of physically captured images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If synthetic images are generated to supplement physically captured images, then the quantity of training examples is improved, but the domain shift between synthetic and physical images increases

Engineering Contradiction:
Improvequantity of training examplesVSAvoiddomain shift
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent applies parameter changes by modifying the noise level during diffusion model training. By controlling the noise level parameter, the model learns to generate images that balance synthetic versatility with physical domain characteristics, reducing domain shift while maintaining adequate training example quantity

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements feedback mechanisms through cost functions that evaluate the agreement between predicted and actual noisy versions of training images. This feedback loop enables the diffusion model to iteratively improve its parameter optimization, ensuring generated images remain within the physical domain while providing sufficient training data

Inventive Principle:
Principle #23Feedback

2Measurement precision

If manual labeling of training examples is performed, then the quality of training data is improved, but the time and cost increase

Engineering Contradiction:
Improvequality of training dataVSAvoidtime and cost
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies self-service by enabling the diffusion model to automatically generate and optimize training images without requiring manual labeling. The model uses its own predictions and the cost function to self-optimize parameters, eliminating the time-consuming manual labeling process while maintaining training data quality through automated evaluation

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent uses copying by creating synthetic training images that replicate the characteristics of physical images. The diffusion model copies the essential features and domain characteristics from physical training images to generate synthetic counterparts, providing adequate training data without manual intervention

Inventive Principle:
Principle #26Copying

Data Source

PatentEP4651071A1Generation of synthetic images for machine learning model training
Publication Date: 2025.11.19 ROBERT BOSCH GMBH
  • EP4651071A1 patent drawingFigure 1
  • EP4651071A1 patent drawingFigure 2
  • EP4651071A1 patent drawingFigure 3

AI summary

Method (100) for training a diffusion model (1) that can be used to iteratively generate a synthetic image (4) from noise (2) in conjunction with a given conditioning (3), which is consistent with this conditioning (3), comprising the steps: • a style (5) is specified (110) that the synthetically generated images (4) should have; • a set of training images x0 is provided (120) that correspond to the specified style (5) to varying degrees; • the training images x0 are successively subjected to noise (2) in a given number T of iterations (130) so that noisy versions x1,...,xT are produced; • samples xt are drawn from the noisy versions x1,...,xT (140); • the drawn samples xt are processed in conjunction with the specified conditioning (3) by the diffusion model (1) to make predictions x̂t-1 for the respective previous noisy version xt-1 (150);• The agreement of these predictions x̂t-1 with the respective actual noisy versions xt-1 is evaluated using a predefined cost function (7) (160); and • Parameters (1a) that characterize the behavior of the diffusion model (1) are optimized (170) to improve the evaluation (7a) by the cost function during further processing of training images x0 and samples xt generated therefrom, wherein when drawing (140) the samples xt, and/or when evaluating (160) predictions x̂t-1 generated therefrom by the cost function (7), samples xt that still reveal the style of the respective training image x0 are represented more strongly the more the respective training image x0 corresponds to the predefined style (5).