Method of generating data set for training and / or testing machine learning system

By selecting and combining multiple conditions and using conditional masking, the problem of image synthesis quality degradation in existing technologies is solved, enabling image generation with multiple conditions applied in specific areas, thus improving the training and testing effects of autonomous driving systems.

CN121600341APending Publication Date: 2026-03-03ROBERT BOSCH GMBH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510980256.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-08-16
Filing Date
2025-07-16
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing image synthesis methods suffer from image quality degradation when using various conditions, especially combinations of different image regions, making it difficult to achieve flexible and adaptable application.

Method used

By selecting at least two distinct conditions and specifying their respective effects when generating the dataset, combined with conditional masking, and using a single control model such as ControlNet for image generation, the influence of conditions can be excluded in specific regions, enabling the combination and flexible application of conditions.

Benefits of technology

The generated dataset can be applied under multiple conditions in a specific area, providing precise details of the scene, improving the flexibility and accuracy of image generation, and is suitable for training and testing of autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600341A_ABST
    Figure CN121600341A_ABST
Patent Text Reader

Abstract

The invention relates to a method for generating a data set for training and / or testing a machine learning system, in particular to a method (100) for generating at least one data set for training and / or testing a machine learning system (55), said generation being carried out by controlling a model (50).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for generating at least one dataset. Furthermore, the invention also relates to a machine learning model, a computer program, an apparatus, and a storage medium for this purpose. Background Technology

[0002] Image synthesis using generative artificial intelligence (AI) refers to generating images through a generative model trained to produce visual content. In this process, specific conditions or inputs can be pre-defined. For example, the model can respond to a variety of input formats, from simple text descriptions to complex datasets.

[0003] Image synthesis enables a wide range of applications, particularly for training and / or testing machine learning systems for driver assistance or autonomous driving. Flexible and adaptable application conditions and inputs are especially important in this context.

[0004] As is known in the art, existing image synthesis methods are primarily based on a limited number of conditions. With the help of techniques such as ControlNet, it is possible to control the generation of synthesized images based on Stable Diffusion through conditional inputs that go beyond simple text input. These conditions can be derived from real images or from simulated environments.

[0005] However, these solutions can encounter problems when using multiple conditions (such as Canny edges, depth maps, or semantic label maps), such as a decrease in image synthesis quality. Therefore, combining multiple conditions, especially for different image regions, remains a challenge. Summary of the Invention

[0006] This invention provides a method, a machine learning model, a computer program, an apparatus, and a computer-readable storage medium. Further features and details of the invention can be derived from the corresponding dependent claims, the specification, and the drawings. The features and details described in relation to the method according to the invention are naturally also related to the machine learning model according to the invention, the computer program according to the invention, the apparatus according to the invention, and the computer-readable storage medium according to the invention, and therefore can always be referred to in conjunction with each other within the disclosure of this invention.

[0007] The subject of this invention is, in particular, a method for generating at least one dataset for training and / or testing machine learning systems, especially for applications in vehicle control systems. The generation can be achieved through a control model, particularly a generative one, which is also based on machine learning and can therefore be executed as a learning system or part of it.

[0008] The method according to the invention may include training and / or testing a control model to train the control model for application of combined conditions. Alternatively or additionally, the method according to the invention may also include inference of the control model, wherein combined conditions can be selected. It is also possible that the method according to the invention includes using a dataset and / or learning the inference of the learning system.

[0009] The method according to the invention may include selecting at least two different conditions that are chosen for use when generating a dataset. These conditions may each provide different controls for generating the dataset. Here, the effect of each condition on the generation can be specified. In this context, control over the generation or the influence of control over the generation may also be involved.

[0010] Furthermore, the method according to the invention may include selecting regions within conditions in which the application of the conditions is excluded, preferably all conditions are excluded. This allows specific regions of the dataset to be generated without being restricted by (selected / combined) conditions.

[0011] The conditions can be spatially defined, and are therefore preferably implemented as at least two-dimensional conditions. This is the case, for example, when the conditions are implemented as masks or graphs, which spatially define different specifications for generation such that each specification applies only to a specific region. Alternatively, the conditions can also be referred to as masks / graphs or conditional masks, such that these regions are correspondingly regions within the mask / graph. The selected regions can be at least two-dimensional regions in the graph and are specified, for example, by coordinates. To name just a few examples, spatially defined conditions can be implemented as semantic segmentation label graphs or Canny edges.

[0012] Furthermore, the conditions can be, for example, specific algorithms, filters, and / or processing rules defined by at least one at least two-dimensional graph, and therefore should be applied to specific regions within the graph. For example, the conditions can provide pixel-based processing rules (specifically, pixel value processing rules) as a specification for generation.

[0013] These regions can represent different areas or segments within the graph, each with different attributes or features. For example, in a semantic segmentation labeling graph, one region can be labeled with a specific object, such as a building or road, while another region can contain information about vegetation, etc. In regions highlighted by Canny edges, special filters or algorithms can be applied as appropriate to enhance edges or highlight object boundaries.

[0014] Furthermore, the method according to the invention may include combining selected conditions. For example, this can be achieved through a corresponding specification used by a control model.

[0015] Subsequently, according to the method of the present invention, a dataset can be generated using a control model (i.e., particularly through and / or controlled by the control model) while applying combinatorial conditions and considering selected regions. Here, the control model can be a machine learning model, such as the ControlNet diffusion model. The control model can be implemented as a single control model and, in particular, trained (end-to-end) for applying combinatorial conditions. Depending on the architecture, the control model itself can be implemented as a generative model or implemented to control another generative model (such as Stable Diffusion).

[0016] The method according to the present invention provides a flexible possibility for automatically generating datasets, particularly image synthesis, enabling automatic...

[0017] a) Considering multiple conditions simultaneously, and / or

[0018] b) Using different combinations of these conditions in reasoning, and / or

[0019] c) It also completely excludes the region from the influence of one or more conditions.

[0020] The method according to the invention can generate datasets, particularly image data, for applications such as autonomous driving, providing precise details of scenes, especially traffic scenes. Simultaneously, diverse scenes can be represented by applying conditions only in specific spatial regions. For example, in certain regions, Canny edges can be given more weight than semantic labels to represent accurate contours and shapes. However, in other regions, the weights can be reversed to emphasize semantic labels. Regions can also be provided where no conditions are applied, thus allowing for unrestricted synthesis of diffusion models.

[0021] The conditions can be combined by scene and / or by image, while areas can be partially or completely excluded from the influence of the conditions. In other words, the method according to the invention can select the conditions and the areas by image / scene.

[0022] Similarly, it is conceivable that when generating datasets, the generation process can be controlled by controlling the model based on selected conditions and / or limiting it to regions that have not been excluded.

[0023] According to the method of the invention, regions can also be selected in which the conditions are not applied at all, and therefore no conditions are defined and / or exerted, and the generation process is carried out uncontrolled by the control model in terms of conditions. This makes it possible for the existence of regions where multiple (combined) conditions simultaneously exert a controlling effect and / or the effects of multiple conditions are simultaneously defined; and also for the existence of regions where no conditions exert a controlling effect.

[0024] The selection can be, for example, masking elements in an image to determine which conditions apply to that element. In this context, "masking conditions" can also be mentioned. Specifically, "masking" conditions mean masking elements / regions for different conditions (e.g., semantic labels or label maps, Canny edges) to modulate their impact on image synthesis. Masking can weaken or even completely eliminate the influence of a particular condition, enabling control models such as ControlNet to correctly handle different conditions and control the image synthesis of the generative model in different ways according to the needs of a specific problem.

[0025] To control the generation process through combined conditions, control models such as ControlNet can be trained using joint masked conditions. Alternatively, the already known training process can be used unchanged by simply adding a new component that provides a joint combination of multiple conditions.

[0026] Multiple conditions can be extracted from the base image (also known as the original image) to train a control model such as ControlNet, which learns how to reconstruct the image based on the conditions. This can also include extracting Canny edges and semantic label maps using a pre-trained semantic segmentation model. Human annotations can also be used if available. Preferably, there can be two or more conditions or combinations of other conditions (e.g., depth map and semantic label map, etc.).

[0027] In principle, for each condition (Canny edge, semantic label, etc.), there are three possible operations in ControlNet training based on the mask condition: 1) completely preserve; 2) partially preserve (by category or region); or 3) remove.

[0028] During training, the masking process for determining how to apply different conditions / canonicals can be chosen randomly. In this way, the control model (especially ControlNet) learns to rely on a single condition (e.g., only Canny edges or only semantic labels), a complete combination of both, or a partial combination of both during inference. This gives users the freedom to decide which conditions to use and how to combine them during inference.

[0029] Furthermore, it is advantageous if the method includes the following:

[0030] - Collect first user input, which specifies manual selection of the conditions, and / or

[0031] - Collect a second user input that specifies a manual selection of the area.

[0032] The selection of conditions can be based on a first user input and / or the selection of regions can be based on a second user input, allowing the user to decide which conditions should be combined, and / or allowing the user to mask the conditions, which modulates the influence of the conditions on dataset generation (particularly image synthesis). This ensures that the control model can handle different user inputs and outputs to control the combination of conditions.

[0033] The machine learning system can be implemented as a model for image synthesis, preferably an image diffusion model, and / or a control model can be implemented as a model for controlling the image diffusion model to perform image synthesis.

[0034] Furthermore, it is conceivable that the dataset includes multiple synthetic images representing objects in the environment, used for training and / or testing machine learning systems. The environment could be, for example, a camera environment and / or a vehicle environment.

[0035] Furthermore, the design and / or arrangement of objects can be influenced by the application of conditions. This has the advantage of allowing machine learning systems to be trained and tested based on multiple conditions, resulting in greater flexibility and accuracy in image generation. Combinations of different conditions allow for the generation of high-quality images that depend on some of those conditions.

[0036] According to an advantageous improvement of the invention, the application of combined conditions can be specified as being provided through a single control model. Therefore, the generation of the dataset can be performed solely through a single control model, particularly an end-to-end trained model such as ControlNet. Thus, the combined conditions can be specified as being provided through a single control model containing all the necessary information for controlling and / or generating the dataset. In particular, this can reduce the disadvantages caused by the cascading application of control models with different conditions.

[0037] Alternatively, it is conceivable that the method may further include:

[0038] - Provide raw images that are used to train the control model and / or the image diffusion model controlled by the control model.

[0039] In addition, the method may also include:

[0040] - Select regions in the form of pixels and / or points and / or two-dimensional regions, especially in the original image, selecting regions where conditions (especially combinations) are not intended to be used - i.e., excluding these conditions.

[0041] The advantage of this is that the control model can consider multiple conditions simultaneously, thus providing greater flexibility and accuracy in image generation. Furthermore, by combining multiple conditions, precise detailing of the scene can be achieved.

[0042] Another advantage is that the generated and / or original images can be specified to represent traffic scenes, allowing the dataset to be used to train and / or test machine learning systems for controlling vehicles in at least partially automated and / or driver assistance systems. This offers the advantage that, through the use of combined and masked conditions, the images can be optimized for training machine learning models in the field of vehicle control, as traffic scenes typically represent highly complex environments with numerous objects and interactions.

[0043] It is also conceivable that training aims to train a machine learning system using the generated dataset to classify digital images, particularly image classification, based on image points and / or pixels, especially pixel values, preferably edges or pixel attributes. These digital images can be, for example, digital images recorded by at least one sensor, preferably at least one camera, particularly a vehicle camera, and especially preferably by the camera and / or the vehicle's environment during movement. For example, recording can be done by at least one camera of a vehicle. The classification can be used to identify objects in the environment depicted by the digital images and / or to capture traffic scenes.

[0044] The classification can be used in various technical applications. One example is its application in vehicles. Based on the classification, and in particular at least one classification result, at least one control action can be initiated and / or executed, preferably in vehicles or other technical systems.

[0045] The classification results may include at least one of the following results and / or be specific to at least one of the following results: object category, object identification, object and / or obstacle location (e.g., in or next to the direction of travel), obstacle presence, description of traffic scene, hazard warning, number of objects, type and / or location of lane markings and / or lane boundaries, location and / or status of traffic signal equipment, lane location, etc.

[0046] Based on the classification results, at least one control action for the vehicle can be initiated and / or executed. The control action may include at least one of the following: braking, steering, acceleration, overtaking maneuver, emergency braking, activation of an alarm system, activation of hazard warning lights, activation of a turn indicator, and light control.

[0047] The classification process can, for example, identify obstacles, whether they are directly in or beside the direction of travel. Based on their location (e.g., depending on the expected vehicle trajectory), appropriate control actions, such as braking or avoidance, can be initiated.

[0048] For example, if the classification indicates the presence of an obstacle and / or a potential collision in the direction of travel, braking can be initiated. It is also conceivable to identify lanes and / or lane boundaries based on the classification so that the vehicle can be kept within the lane at least partially automatically through control actions.

[0049] The vehicle may be configured as a motor vehicle and / or a passenger car and / or at least a partially automated vehicle.

[0050] The terms "classification" and "image classification" may also include "object detection" or "object detection in an image." This should be understood in particular as classifying whether an object exists in certain regions of an image. Furthermore, the terms "classification" and "image classification" may also refer to "semantic segmentation," especially in the form of pixel-level classification.

[0051] In this invention, it is advantageous to specify that, by selecting conditions and regions, the influence of conditions can be dynamically preserved, partially preserved, and / or removed during the generation of the dataset. This has the advantage of enabling flexible control over image generation. Different regions can be considered without affecting the model's overall generated output in the affected regions. This ensures that control is applied to specific regions while other regions continue to utilize the free creative potential of the diffusion model.

[0052] Alternatively, it may be conceivable that the condition includes at least two of the following elements and / or other similar elements:

[0053] -Canny Edge, especially for edge and structure recognition.

[0054] - Semantic tags, especially for object classification and annotation.

[0055] - A color palette, especially for visual differentiation and categorization.

[0056] - Depth maps, especially used for collecting and analyzing spatial information.

[0057] This allows for targeted influence on different features of an image or scene.

[0058] The subject matter of this invention also includes a machine learning model, particularly the aforementioned machine learning system, which has been trained using at least one dataset obtained by the method according to the invention. Therefore, the machine learning model according to the invention has the same advantages as that described in detail with reference to the method according to the invention.

[0059] Within the scope of this invention, it can be specified that the machine learning model according to the invention has been trained for use in at least partially automated driving and / or driver assistance systems. This enables more precise vehicle control even in complex driving scenarios, thereby contributing to vehicle safety. Furthermore, training using a dataset generated according to the method of the invention can achieve high accuracy in the recognition and interpretation of environmental information, thereby improving the ability to react to traffic conditions.

[0060] More generally, machine learning systems (particularly the machine learning model according to the invention) can be used in vehicles. Vehicles can be configured, for example, as motor vehicles and / or passenger cars and / or autonomous vehicles. Vehicles can have vehicle devices, for example, for providing autonomous driving functions and / or driving assistance systems. Vehicle devices can be configured to control the vehicle and / or accelerate and / or brake and / or steer the vehicle at least partially automatically based on the output of the learning system.

[0061] The subject matter of this invention also includes a computer program, particularly a computer program product, comprising instructions that, when executed by a computer, cause the computer to perform the method according to the invention. Therefore, the computer program according to the invention has the same advantages as those described in detail with reference to the method according to the invention.

[0062] The subject matter of the invention also includes an apparatus for data processing configured to perform the method according to the invention. As an apparatus, for example, a computer, such as a control unit of a vehicle, may be provided, which executes a computer program according to the invention. The computer may have at least one processor for executing the computer program. A non-volatile data memory may also be provided, in which the computer program is stored, and from which the processor may read the computer program for execution.

[0063] The subject of this invention can also be a computer-readable storage medium having a computer program and / or instructions according to the invention, which, when executed by a computer, cause the computer to perform the method according to the invention. The storage medium is, for example, configured as a data storage device, such as a hard disk and / or non-volatile memory and / or a memory card. The storage medium can, for example, be integrated into a computer.

[0064] Furthermore, the method according to the invention can also be executed as a computer-implemented method. Alternatively or additionally, at least one disclosed method step can be computer-implemented and / or automatically executed. Attached Figure Description

[0065] Other advantages, features, and details of the invention will become apparent from the following description, in which embodiments of the invention will be described in detail with reference to the accompanying drawings. Here, the features mentioned in the claims and specification may be essential to the invention individually or in any combination. In the drawings:

[0066] Figure 1 A schematic visualization of a method, apparatus, storage medium, and computer program according to embodiments of the present invention is shown.

[0067] Figure 2 An exemplary problem is shown when using only ControlNet for Canny edge condition control of stable diffusion.

[0068] Figure 3 This paper illustrates a method for training ControlNet using combined mask conditions, based on an embodiment. Detailed Implementation

[0069] Figure 1 The method 100, apparatus 10, storage medium 15, vehicle 60, control model 50, learning system 55, and computer program 20 according to embodiments of the present invention are illustrated schematically.

[0070] Method 100 can be used to generate at least one dataset for training and / or testing the machine learning system 55. The generation of the at least one dataset can be achieved by controlling model 50.

[0071] According to step 101 of the first method, at least two different conditions can be selected for use when generating the dataset. These conditions can each provide different controls for the generation of the dataset, specifying the impact of each condition on the generation process.

[0072] According to step 102 of the second method, select regions 307 within the conditions and exclude the application of the conditions in these regions.

[0073] In step 103 of the third method, the selected conditions are combined.

[0074] Subsequently, according to step 104 of the fourth method, a dataset can be generated with the help of the control model 50, taking into account the combined conditions and the selected region 307.

[0075] Method 100 may further include providing raw images 301 for training control model 50 and / or an image diffusion model controlled by control model 50. Regions 307 may then be selected in the raw images 301 in the form of pixels and / or points and / or two-dimensional regions, in which conditions (especially combined conditions) are not intended to be applied. In other words, regions that should be unaffected by conditions are masked. For example, in these regions, control model 50 will not perform condition-based control.

[0076] Figure 3 The original image 301 shown in the example may represent a traffic scene so that the dataset can be used to train and / or test a machine learning system 55, for example, to control a vehicle 60 in at least a partially automated driving and / or driver assistance system.

[0077] The control model 50 can be implemented, for example, as ControlNet. ControlNet (see [1], references listed at the end of the specification) is an extension of the Stable Diffusion model that can be finely tuned using a variety of conditions. These conditions may include semantic segmentation label maps, Canny edges, pose estimation, and / or other parameters. During the inference phase, these conditions help to regulate the image generation process. This allows the generated data to conform to a specific scene layout while maintaining the powerful textual conditional control capabilities of the pre-trained Stable Diffusion model.

[0078] Despite the immense potential of this technology, several limitations remain. When using Canny edges extracted from another image for conditional control, some objects may not provide enough information, causing the edges to fail to be correctly identified as "pedestrians." For example, objects located far away in the scene—such as pedestrians at the other end of a street. Therefore, the diffusion process may treat this part of the input as interference, generating mismatched objects. Furthermore, issues have been found when using Semantic Label Maps (SLMs) for conditional control of stable diffusion. For instance, when extracting SLM or Canny edges from an image template, labeled objects (such as cars) in the generated image often exhibit unexpected orientation and scaling.

[0079] Figure 2 This illustrates an exemplary problem when ControlNet only performs Canny edge conditional control on stable diffusion. The figure shows the original image 201 from which edges are extracted, with the generated image 202 below it. When edges are noisy and not recognized by the model, objects may disappear, such as the marked pedestrian in the example.

[0080] However, this problem can be mitigated by adding category information to pedestrians. Therefore, according to embodiments of the present invention, it is proposed to employ multiple concurrent conditions and assign different weights (masks) to each condition during the training phase. This enriches the information input to the diffusion model. This expanded information can be used during inference to at least partially address the aforementioned problem.

[0081] By simultaneously using conditions such as Canny edges and semantic label maps, the semantic label map can provide the necessary information to ensure accurate representation of pedestrians during image generation, even when the edges cannot identify distant pedestrians. Therefore, this method has many advantages over existing technologies.

[0082] In embodiments of the present invention, known methods, such as those disclosed in [1], [2], or [3], may also be used. Specifically, these methods combine text-image diffusion models (such as Stable Diffusion) with external control signals (such as Canny edges, semantic tag maps, sketches, etc.). This enables fine-grained control over the image generation process, going beyond simple text prompts. Through this conditional control, a high degree of control can be exercised over multiple aspects of the image. For example, attributes such as the position, orientation, and pose of an object can be defined using Canny edges. These attributes cannot be defined solely by text prompts. Therefore, a high degree of control can be provided for text-image models such as Stable Diffusion.

[0083] Each conditional control method has its advantages and disadvantages, and the decision should be made at the level of the specific application scenario. For example, Canny edges can define pose and orientation, but for objects that are far away in the background or a large number of densely distributed objects, Canny edges may not provide the diffusion model with the correct information about which objects should be drawn. This problem is preferably solved by using a semantic label map (also referred to as a semantic tag map in this invention) as a condition, which defines the category of the object to be drawn for each pixel in the image. However, this does not allow the user to control the position and orientation of objects like Canny edges do.

[0084] Therefore, the embodiments of the present invention provide a solution that combines the advantages of multiple conditional inputs while mitigating their limitations.

[0085]

[0086] One way to combine multiple conditions is to use multiple ControlNets connected to the same Stable Diffusion backbone, also known as "multi-control networks." ControlNets work by guiding the model to generate images in the direction of the control input by adding or subtracting the outputs of nodes in different intermediate layers of the Stable Diffusion model. During training, only ControlNet models are updated; therefore, they actually learn how to control the diffusion model rather than directly learning how to generate images.

[0087] Since these control networks simply add or subtract values ​​from the diffusion model at each node, multiple ControlNet models can be arbitrarily superimposed on the diffusion model. Each model guides the diffusion model in its corresponding direction at each step, ultimately resulting in a cumulative effect of "following all control instructions" in the generated results.

[0088] However, in practice, this strategy is not ideal for generating high-quality images because multiple chained conditions can interfere with each other (since they are not jointly trained) and it cannot apply different conditions to different image regions. For example, in a driving scene, if the car is defined using Canny edges and the background using semantic label maps, this configuration could achieve precise control over the car's appearance while allowing the model creative freedom in background design, thus balancing control and diversity across different image regions. The above strategy fails to achieve this. This key issue will be analyzed in depth below. Related background information can be found at:

[0089] [1] "Adding conditional control totext-to-image diffusion models." arXiv preprint arXiv:2302.05543 (2023) by Zhang, Lvmin and Maneesh Agrawala;

[0090] [2] "T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models." arXiv preprintarXiv:2302.08453 (2023) by Mou, Chong et al.;

[0091] [3] "Composer: Creative and controllable images synthesis with composable conditions." arXiv preprint arXiv:2302.09778 (2023) by Huang, Lianghua et al.

[0092] This invention proposes a novel training variant of a single ControlNet that can simultaneously incorporate multiple conditions. To address the aforementioned issues, co-training these conditions is crucial. Furthermore, each condition should be optional during inference, allowing the diffusion process to generate high-quality images using only some conditions while also considering conditions that exist only in specific spatial regions. This flexibility enables precise scene detailing where needed (e.g., object inpainting) while preserving unconstrained areas for more diverse generated results.

[0093] Figure 3 This illustrates a method for training ControlNet using combined mask conditions, as proposed in an embodiment. In the final conditional tensor 203 (top right), some regions contain only labels, some regions contain only Canny edges, and some regions contain both or neither.

[0094] First, conditions are extracted from the original image 301 (see 311) to obtain semantic labels 304 and Canny edges 305, which can be used as conditions 302 for generative image synthesis. Through masking 312, results 306 and 307 are obtained, which can be combined into conditions (see 313 and 303).

[0095] In principle, for each condition (Canny edges, semantic labels, color palettes, depth maps, etc.), there are at least or exactly three possible operations: complete preservation, partial preservation (by category or region), and removal. By determining the optimal combination of these options for different condition maps during training, high-quality images can be generated using any combination of conditions during the inference phase.

[0096] This allows multiple conditions to be combined to jointly control the image generation of the diffusion model. These conditions can be masked and applied to the entire image or only specific regions, providing greater flexibility in image generation. Although multiple conditions are used during training, their use is not mandatory during inference; it is still possible to choose to use only one condition.

[0097] Because multiple conditions are trained together, using multiple conditions during inference ensures higher image quality. However, training multiple conditions separately (as is currently possible) leads to contradictory image generation processes, making it impossible to easily integrate the various conditions.

[0098] Therefore, one objective of an embodiment of the present invention is to generate photorealistic images with predefined driving scenarios, wherein the conditions for synthesizing images using diffusion models are extracted from existing real or simulated images.

[0099] One possible application of the method according to embodiments of the present invention is that these synthesized images can be used to enhance (extend) the training and validation of autonomous driving systems. However, the proposed method is not limited to this application scenario in principle and can be used in any scenario that requires synthesized images.

[0100] The above description of the embodiments illustrates the invention by way of example only. Of course, individual features of the embodiments can be freely combined without departing from the scope of the invention, provided it is technically feasible.

[0101] Cross-references to related applications

[0102] This application claims priority to European patent application filed on August 16, 2024 (previous application number: 24195032.8). All disclosures of the above application are incorporated herein by reference.

Claims

1. A method (100) for generating at least one dataset for training and / or testing a machine learning system (55), wherein, The generation is achieved through a control model (50), and the method includes: - Select (101) at least two different conditions to use when generating the dataset, each condition providing a different control over the generation of the dataset, wherein the influence of each condition on the generation is specified. - Select the region (312) within the conditions described in (102) to exclude the application of the conditions in the region. -Combined (103) selected conditions, - The dataset (104) is generated by means of the control model (50) under the application of combined conditions and taking into account the selected region (312).

2. The method (100) according to claim 1, Its features are, In generating the dataset (104), the control model (50) controls the generation process according to selected conditions and limited to the unexcluded regions, and Regions where the application of the conditions is completely excluded are also selected, in which the effects of the conditions are not specified and / or the conditions are not applied, and thus the generation is performed uncontrolled by the control model (50) in terms of the conditions.

3. The method (100) according to any one of the preceding claims, Its features are, The method (100) further includes: - Collect input from a first user, whereby the first user input specifies a manual selection of the conditions. - Collect input from a second user, who specifies manual selection of the area. The conditions are selected based on the first user input selection (101) and the region is selected based on the second user input selection (102), allowing the user to decide which conditions should be combined and allowing the user to mask the conditions, the masking process modulating the influence of the conditions on the generation of the dataset, particularly on image synthesis.

4. The method (100) according to any one of the preceding claims, Its features are, The machine learning system (55) is implemented as a model for image synthesis, preferably as an image diffusion model, and / or The control model (50) is implemented as a model for controlling the image diffusion model to perform image synthesis. Furthermore, the dataset includes multiple synthetic images representing objects in the environment, used for training and / or testing the machine learning system (55). The design and / or arrangement of the object are influenced by the application of the conditions.

5. The method (100) according to any one of the preceding claims, Its features are, The application of the combined conditions is provided by a single control model (50), therefore, the generation of the dataset is preferably performed only by a single control model (50), particularly by an end-to-end trained ControlNet.

6. The method (100) according to any one of the preceding claims, Its features are, The method (100) further includes: - Provide (101) an original image (301) for training the control model (50) and / or an image diffusion model controlled by the control model (50). - The selection (102) of the region (312) in the original image is performed in the form of pixels and / or points and / or two-dimensional regions, in which the conditions, in particular combined conditions, should not be applied, wherein the conditions are preferably implemented as spatially defined, preferably at least two-dimensional conditions, in particular in the form of masks or graphs.

7. The method (100) according to claim 6, Its features are, The original image (301) represents a traffic scene so that the dataset can be used to train and / or test the machine learning system (55) for controlling the vehicle (60) in at least a partially automated driving and / or driver assistance system.

8. The method (100) according to any one of the preceding claims, Its features are, The training aims to train the machine learning system (55) based on the generated dataset to classify digital images based on image points and / or pixels, particularly digital images obtained from environmental records of the vehicle (60) and / or from cameras during vehicle (60) operation, wherein the vehicle (60) is preferably controlled based on the classification.

9. The method (100) according to any one of the preceding claims, Its features are, By selecting the conditions (101, 102) and the region (312), the effects of the conditions can be dynamically preserved, partially preserved, and / or removed during the generation of the dataset.

10. The method (100) according to any one of the preceding claims, Its features are, The condition includes at least two of the following elements: - Canny edges for edge and structure recognition - Semantic tags used for object classification and annotation. - A color palette for visual differentiation and classification. - Depth maps used for collecting and analyzing spatial information.

11. A machine learning model (55) that has been trained using at least one dataset obtained by the method (100) according to any one of the preceding claims.

12. The machine learning model (55) according to claim 11, Its features are, The machine learning model (55) is trained for use in at least partially automated driving and / or driver assistance systems.

13. A computer program (20) comprising instructions that, when executed by at least one computer (10), cause the computer to perform the method (100) according to any one of claims 1 to 10.

14. An apparatus (10) for data processing, the apparatus being configured to perform the method (100) according to any one of claims 1 to 10.

15. A computer-readable storage medium (15) comprising instructions that, when executed by at least one computer (10), cause the computer to perform the steps of the method (100) according to any one of claims 1 to 10.