Methods for generating synthetic images
By optimizing the image generation process through a two-stage noise adjustment process of the diffusion model, the problem of low image generation efficiency in existing technologies is solved, and efficient generation of images that conform to specific content is achieved, thereby improving the training effect of machine learning systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ROBERT BOSCH GMBH
- Filing Date
- 2026-01-28
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies are inefficient in generating image data and struggle to efficiently generate images that match specific content, especially when training machine learning systems, particularly for image data in traffic scenarios involving vehicles and robots and industrial facilities.
Image generation is performed using a diffusion model, which involves a two-stage process: in the first iteration, noise is repeatedly reduced and added to an intermediate generated image; in the second iteration, noise is repeatedly reduced and added to multiple intermediate generated images. The noise adjustment process is optimized by combining text input and random noise information.
It improves the efficiency and quality of image generation, better presents detailed structures and patterns, and generates more realistic images. It is suitable for training machine learning systems to improve pixel-level object detection and classification, especially in the fields of vehicles and robotics.
Smart Images

Figure CN122492861A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for generating composite images. Furthermore, the invention also relates to computer programs, apparatus, and storage media used for this purpose. Background Technology
[0002] Neural networks are now the preferred method for processing image data. The quality of a network depends on the data used to train it.
[0003] However, compiling training data is increasingly proving challenging. Therefore, generative neural networks are often employed, which can generate images based on text input or similar input. Since a certain amount of data is typically required, efficiently generating such data is crucial. Summary of the Invention
[0004] The subject of this invention is a method, a computer program, an apparatus, and a computer-readable storage medium. Other features and details of the invention derive from their respective dependent claims, description, and drawings. Hereinafter, the features and details described in association with the method according to the invention also apply to the computer program according to the invention, the apparatus according to the invention, and the computer-readable storage medium according to the invention, and vice versa; therefore, cross-referencing with respect to the disclosure of this invention is always possible.
[0005] The subject of this invention is, in particular, a method for generating synthetic images, hereinafter also referred to as image generation or synthesis.
[0006] Here, image synthesis can be provided using generative models (particularly diffusion models). In other words, image synthesis can be based on machine learning and / or on a stepwise transformation from an initial distribution, where data is transformed into a noisy representation and subsequently iteratively reconstructed. Diffusion models here preferably employ a two-stage process: in the forward stage, random noise is added to the data, gradually blurring the original structure; in the backward stage, noise is progressively removed using a trained neural network to reconstruct the original data or generate new synthetic content. This process can be part of a noise conditioning process, which will be described in more detail below.
[0007] Furthermore, the method according to the invention may include the following steps, which are preferably performed automatically and / or sequentially and / or repeatedly: - Provide input, - The provided input is processed at least in the first iteration by a noise remover (also known as a denoiser), where noise is repeatedly reduced and added in an intermediate generated image during the first noise conditioning process. - The intermediate output of the first iteration is processed in the second iteration by at least the noise remover and / or another noise remover, wherein, in the second noise conditioning process, noise is repeatedly reduced and added in multiple intermediate generated images. - Provide the output of the second iteration process, preferably in the form of the generated synthetic image.
[0008] This invention describes a more efficient way to generate images—preferably using a diffusion model. In particular, this invention provides an improved diffusion process that follows a single data point process in the initial steps, branching out to generate multiple images only at later time points.
[0009] Optionally, in the first iteration, the noise conditioning process is performed only once on the single intermediate generated image, while in the second iteration, it is performed multiple times on the multiple intermediate generated images. In other words, while the noise conditioning process in the first iteration can be repeated, it is performed only once on a single intermediate generated image, meaning noise is generated only once per iteration. In contrast, in the second iteration, noise can be generated multiple times per processing step, and correspondingly, multiple intermediate generated images are provided per processing step (opposite to the first iteration). Overall, image quality is improved because repeated addition and removal of noise within a single image achieves higher detail fidelity. The multiple applications of noise conditioning in the second iteration also allow for a more refined adjustment of the image to match the text description. Consequently, the detailed structure and patterns in the generated image are better presented, resulting in a more realistic outcome.
[0010] According to an advantageous improvement of the invention, the input may be specified to include at least one of the following: - Text input, which describes the images to be generated, preferably defining the goal of image generation, i.e., generating multiple images whose content conforms to the text input, and preferably adapted to the technical application purpose of the machine learning system to be trained using the generated images. - Random input, which is preferably generated by means of a random generator, to provide initial noise information.
[0011] In other words, the input can include both structured content (text description) and unstructured content (random input). Text description enables the model to generate images that conform to specific content. Random number generation provides initial noise information, which the model can use for image generation.
[0012] Advantageously, in this invention, the input can be specified to include at least initial noise information, which corresponds to an initial noisy image, and particularly to random noise. The first iterative process can then perform a noise reduction process based on this initial noisy image to obtain an intermediate generated image. The advantage of this is more efficient image generation. Using the initial noise information allows for a more direct entry into the denoising process and reduces the necessary number of iterations, thereby saving resources and shortening generation time.
[0013] Within the scope of this invention, it is preferable to specify that, in each processing step of the first iteration process, noise is removed from the intermediate generated image and then the noise is added back to the intermediate generated image. Preferably, no other intermediate generated images and / or other noise are used or generated in this processing step.
[0014] It can also be specified that noise is generated multiple times after the processing steps of the first iteration process, thereby adding noise to the intermediate generated image multiple times, and generating a noise image in this way, i.e., multiple noise images, as intermediate outputs of the first iteration process.
[0015] Furthermore, it is possible that in the processing steps of the second iteration, the plurality of intermediate generated images are generated from these noisy images, and noise is repeatedly reduced and added to these intermediate generated images multiple times. This contrasts with the first iteration, where noise is reduced and added only once per processing step, thus allowing for multiple reductions and additions of noise in the second iteration. In other words, the denoiser in the first iteration can be used to generate a noisy image at each processing step, and then the second iteration generates and processes multiple noisy images from this image at each processing step. This can improve the efficiency of the entire process and save resources.
[0016] Furthermore, the image processed during the process can be referred to as an intermediate generated image to explicitly indicate that it is, for example, the internal processing result or intermediate product of the diffusion model (and therefore not necessarily the actual image output). Conversely, the output provided by the second iteration process—especially the final generated image—can be considered the preferred output. In addition, the intermediate generated image, the initial input, and the output can also correspond to an abstract representation of the image (e.g., an encoded image) or other types of (abstract) signals (e.g., (encoded) audio, video).
[0017] Furthermore, it is conceivable that the at least one noise remover (particularly a denoiser) is implemented as an artificial neural network of a generative model (preferably a diffusion model) that reduces noise by progressively removing noise from the corresponding image, while the generative model iteratively performs the transition from an initial highly noisy image to a clear, detailed image. This has the advantage that the denoiser, through its iterative structure and progressive noise suppression, can achieve high-quality generated images.
[0018] Within the scope of this invention, the provided output (particularly synthetic images) can be used as training and / or validation and / or testing data for a machine learning system for the application of a technical system (preferably a vehicle and / or robot). This particularly enables the machine learning system to be trained to perform classification, preferably pixel-based object detection, so as to control the technical system, preferably based on the classification results. The scene is preferably implemented as a traffic scene and / or a scene in a robotic industrial facility. Therefore, the generated images can be used to improve the machine learning system. Using synthetic images as training data enables efficient training of the classification system, particularly in the fields of vehicles and robots. This results in improved pixel-based object detection, which can be used in scenes such as traffic situations or industrial facilities to control the technical system.
[0019] It is possible that the methods and / or machine learning systems according to the invention are applied in vehicles. Vehicles may be configured, for example, as motor vehicles and / or passenger vehicles and / or at least partially automated / autonomous vehicles. Vehicles may have vehicle devices, for example, for providing autonomous driving functions and / or driving assistance systems. These vehicle devices may be configured for at least partially automatic control of the vehicle and / or acceleration and / or braking and / or steering.
[0020] The method according to the invention can be used to enrich datasets containing synthetic images, which are then used as training datasets. Machine learning systems (particularly in the form of machine learning models) are trained using synthetic (i.e., generated) images, particularly for classification, especially for object detection. Training can be used to train the machine learning system or model using a training dataset for classification, particularly image classification, based on pixels and / or image points (particularly pixel values), preferably edges or pixel attributes (image data) of image data (such as digital images). The image data or digital images can, for example, originate from recordings by at least one sensor (preferably at least one camera, particularly a vehicle camera), especially preferably vehicle cameras and / or recordings of the vehicle environment during driving (vehicle travel). Recording can be performed by at least one camera of the vehicle. Classification can be used to identify objects in the environment depicted by the image data or digital images and / or to capture traffic scenes.
[0021] Classification can be used in various technological applications. One example is its application in vehicles and / or robots. Based on classification, particularly at least one classification result, at least one control action can be initiated and / or executed, which is particularly useful in vehicles, robots, or other technological systems.
[0022] The classification results may include and / or correspond to at least one of the following: object category, object identification, location of object and / or obstacle (e.g., within or beside the direction of travel), presence of obstacle, description of traffic scene, hazard warning, number of objects, type and / or location of lane markings and / or lane boundaries, location and / or status of traffic lights, lane location, or similar information.
[0023] At least one control action for the vehicle can be initiated and / or executed based on the classification results. The control action may include at least one of the following: braking, steering, acceleration, overtaking maneuver, emergency braking, activation of an alarm system, activation of hazard warning lights, activation of turn signals, light control, or similar actions.
[0024] Classification can, for example, identify obstacles, whether they are directly in the direction of travel or beside it. Based on their location (e.g., depending on the expected vehicle trajectory), appropriate control actions, such as braking or avoidance, can be initiated.
[0025] For example, if the classification indicates the presence of an obstacle and / or a potential collision in the direction of travel, braking can be initiated. Similarly, it is conceivable to identify lanes and / or lane boundaries based on classification, in order to move the vehicle within the lane at least partially automatically through control actions.
[0026] The terms "classification" and "image classification" can also include "object detection" or "object detection in an image." This specifically refers to classifying whether an object exists in certain regions of an image. Furthermore, the terms "classification" and "image classification" can also refer to "semantic segmentation," particularly in the form of pixel-level classification.
[0027] Accordingly, training can produce at least one trained machine learning model that can be used for classification and / or object detection. Use and inference can be provided, for example, in a vehicle. The data points of the input data can be, for example, pixels of image data, or based on this, classification and / or object detection can be performed based on the pixels. The input data can include sensor data and / or image data, which are at least partially derived from acquisitions by sensors (preferably camera sensors) and / or at least partially synthesized, i.e., particularly simulating real data from sensors. Specifically, it can be specified that the environment and / or traffic scene of the sensor and / or vehicle is represented by the values of image data points (preferably pixels). Classification, preferably image classification and / or object detection, can be specified based on these values. This allows, for example, the detection of objects in a traffic scene. The image data can be images from radar sensors and / or ultrasonic sensors and / or lidar sensors and / or thermal imaging cameras. Accordingly, the images can also be implemented as radar images and / or ultrasonic images and / or thermal images and / or lidar images.
[0028] The subject of this invention is also a computer program, and more particularly a computer program product, containing instructions that, when executed by at least one computer, cause the computer to perform the method according to the invention. Therefore, the computer program according to the invention provides the same advantages as those described in detail with reference to the method according to the invention.
[0029] The subject of this invention is also an apparatus for data processing configured to perform the method according to the invention. As such an apparatus, for example, at least one computer may be provided, which executes a computer program according to the invention. The computer may have at least one processor for executing the computer program. A non-volatile data memory may also be provided, in which the computer program is stored, and from which the processor may read the computer program for execution.
[0030] The subject of this invention can also be a computer-readable storage medium having a computer program and / or instructions according to the invention, which, when executed by at least one computer, causes the computer to perform the method according to the invention. The storage medium may be configured, for example, as a data storage device, such as a hard disk and / or non-volatile memory and / or a memory card. The storage medium may be integrated into a computer.
[0031] Furthermore, the method according to the invention can also be executed as a computer-implemented method. Alternatively or additionally, at least one of the disclosed method steps can be computer-implemented and / or automatically executed. Attached Figure Description
[0032] Other advantages, features, and details of the invention will become apparent from the following description, in which embodiments of the invention are described in detail with reference to the accompanying drawings. Here, the features mentioned in the claims and specification may each contribute individually or in any combination to the essence of the invention. In the drawings: Figure 1 A schematic visualization of a method, apparatus, storage medium, and computer program according to embodiments of the present invention is shown.
[0033] Figure 2 Another illustrative representation of an embodiment according to the present invention is shown. Detailed Implementation
[0034] exist Figure 1 The present invention, according to an embodiment, schematically illustrates a method 100, an apparatus 10, a storage medium 15, and a computer program 20. Furthermore, Figure 1 The application of method 100 in generating synthetic images using a generative model (particularly a diffusion model) is shown in the embodiments of the present invention.
[0035] Therefore, in the first method step 101, input is provided, which may include, for example, text input and random input.
[0036] Then, in step 102 of the second method, the provided input can be processed during the first iteration. At least one noise remover (also called a denoiser) can be used here. Here, during the first noise conditioning process, noise is repeatedly reduced and added in an intermediate generated image.
[0037] Then, in the third method step 103, it can be specified that the intermediate output of the first iteration is processed during the second iteration. At least the noise remover or another noise remover can also be used here. During the second noise conditioning process, noise is repeatedly reduced and added in multiple intermediate generated images.
[0038] According to step 104 of the fourth method, the output of the second iteration process is finally specified.
[0039] One of the most common types of models used to generate image data as training data is diffusion models, see, for example, "Ho, Jonathan, Ajay Jain, and Pieter Abbeel. 'Denoising diffusionprobabilistic models.' Advances in neural information processing systems 33(2020): 6840-6851". These models have the property of being able to reverse a given diffusion process.
[0040] The diffusion process can be formalized as follows: 1. Input: A data point (e.g., in an image).
[0041] 2. Repeat N times: Add noise to the data points with a certain intensity (usually based on a normal distribution). This intensity can vary in each step.
[0042] 3. Output: A new data point that, when N is large, should be indistinguishable from noise.
[0043] The diffusion model learns the inverse process based on this process. Therefore, given the data points in step t, the noise added from step t-1 to step t can be estimated. Following the literature, the inverse process can be formulated as follows—see [link to literature]. Figure 2 a) in: 1. Input: A data point consisting only of noise.
[0044] 2. Repeat N times: Estimate the noise using a diffusion model and subtract it from the data points. Then, add noise again (depending on the current step t).
[0045] 3. Output: A new data point that should be similar to the data point used to train the diffusion model.
[0046] This process can be very time-consuming, depending on the number of steps N.
[0047] In practice, it's often desirable to generate multiple data points based on the same input. A typical approach is to generate multiple points in parallel. Therefore, instead of starting with just one data point, one begins with a list containing B data points—see [link to relevant documentation]. Figure 2 (b)
[0048] On the other hand, there are indications that similar noise is often estimated at the start of the reverse process. Therefore, it is unnecessary to estimate each of the B data points individually.
[0049] Therefore, embodiments of the present invention propose an improved diffusion process that follows a single data point process in the initial steps, branching out to generate multiple images only at later time points (see also...). Figure 2 ).
[0050] Due to its later branching, this process is more efficient than existing diffusion processes.
[0051] Embodiments of the present invention include the following aspects—see also Figure 2 c): According to the first aspect, the denoiser (such as a neural network from the U-Net family) receives text input and input generated by a random generator. The goal is to output B images that match the text input. This begins at step T.
[0052] According to the second aspect, the denoiser follows a process consisting of T steps, in which the denoiser removes noise from the image in each step and then adds the noise back to the image.
[0053] According to the third aspect, after n steps, after removing noise by a denoiser, B-order noise is generated to add B-order noise to the image.
[0054] According to the fourth aspect, in the steps from n to 0, each of the B images is denoised separately and then mixed with noise again.
[0055] According to the fifth aspect, the output is B images that conform to the text input, where the efficiency of the first n steps is the same as when only one image is generated.
[0056] Therefore, a denoiser (also known as a noise remover) is able to generate at least one image based on at least one text input and a random input vector.
[0057] exist Figure 2 In the diagram, a) shows a conventional denoising process, b) shows the processing of multiple images according to the conventional process, and c) shows a modified process according to an embodiment of the present invention. The inputs are cue 201 and random noise 202 or random noise value batch 205. The input data is fed to denoiser 210 to obtain denoised image 215 or denoised image batch 216. Noise 220 or B-order noise 221 can then be added. A noisy image 230 or noisy image batch 231 can be generated. By applying denoiser 210, a denoised image 240 or denoised image batch 241 can be obtained.
[0058] The foregoing description of the embodiments illustrates the invention by way of example only. Of course, features of the various embodiments can be freely combined with each other as long as it is technically reasonable, without departing from the scope of the invention.
Claims
1. A method (100) for generating synthetic images using a generative model, particularly a diffusion model, the method comprising the following steps: - Provides (101) inputs, - The input provided (102) is processed at least in the first iteration by a noise remover, wherein, in the first noise conditioning process, noise is repeatedly reduced and added in an intermediate generated image. - The intermediate output of the first iteration process is processed (103) at least by the noise remover and / or another noise remover during the second iteration process, wherein, during the second noise conditioning process, noise is repeatedly reduced and added in multiple intermediate generated images. - Provide the output of the second iterative process (104).
2. The method (100) according to claim 1, Its features are, In each processing step of the first iteration process, the noise conditioning process is performed once on the one intermediate generated image, while in each processing step of the second iteration process, the noise conditioning process is performed multiple times on the multiple intermediate generated images.
3. The method (100) according to any one of the preceding claims, Its features are, The input includes at least one of the following: - Text input, which describes the image to be generated in terms of content, preferably defining the goal of image generation, namely, generating multiple images whose content conforms to the text input. - Random input, which is preferably generated by means of a random generator to provide initial noise information.
4. The method (100) according to any one of the preceding claims, Its features are, The input includes at least initial noise information, wherein the initial noise information corresponds to an initial noise image, and in particular to random noise, wherein the first iterative process performs the noise conditioning process based on the initial noise image to obtain the intermediate generated image.
5. The method (100) according to any one of the preceding claims, Its features are, In the processing steps of the first iteration, noise is removed from each of the intermediate generated images, and then the noise is added back to the intermediate generated image. Following the processing step of the first iteration, noise is generated multiple times, thereby adding noise to the intermediate generated image multiple times, and noise images are generated in this way, i.e., multiple noise images, as the intermediate output of the first iteration. In the processing step of the second iteration process, the plurality of intermediate generated images are generated from the noisy image, and noise is repeatedly reduced and added to the intermediate generated images.
6. The method (100) according to any one of the preceding claims, Its features are, The at least one noise remover, particularly the denoiser, is implemented as an artificial neural network of a generative model, preferably a diffusion model, which reduces noise by progressively removing noise from the corresponding image, while the generative model iteratively performs a transition from an initial highly noisy image to a clear, detailed image.
7. The method (100) according to any one of the preceding claims, Its features are, The synthesized images are used as training and / or validation and / or test data for a machine learning system for a technological system, preferably a vehicle and / or robot application, to train the machine learning system to perform classification, preferably pixel-based target detection, so as to control the technological system, preferably based on the results of the classification, wherein the scene is preferably implemented as a traffic scene and / or a scene in a robotic industrial facility.
8. A computer program (20) comprising instructions which, when executed by at least one computer (10), cause the computer to perform the method (100) according to any one of the preceding claims.
9. An apparatus (10) for data processing, the apparatus being configured to perform the method (100) according to any one of claims 1 to 7.
10. A computer-readable storage medium (15) including instructions that, when executed by at least one computer (10), cause the computer to perform the steps of the method (100) according to any one of claims 1 to 7.