Methods for training and / or testing machine learning systems

By classifying pixels into relevance groups and using layout specifications, the method improves training data diversity and accuracy for machine learning systems, particularly in autonomous driving, focusing on critical elements while allowing flexibility in less important areas.

JP2026082674APending Publication Date: 2026-05-19ROBERT BOSCH GMBH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ROBERT BOSCH GMBH
Filing Date
2025-09-12
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing methods for training machine learning systems, particularly in autonomous driving, fail to balance image diversity and accuracy, often prioritizing global constraints that limit the effectiveness in applications requiring subtle differences.

Method used

A method that spatially constrains image generation by classifying pixels into groups of relevance, allowing more flexibility in less critical areas while maintaining focus on essential elements, using layout specifications and semantic label maps to enhance training data.

Benefits of technology

Enhances the accuracy and diversity of training data, improving the machine learning system's ability to identify critical information, leading to better decision-making in specific applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026082674000001_ABST
    Figure 2026082674000001_ABST
Patent Text Reader

Abstract

The present invention provides training and / or testing methods for machine learning systems for specific technical applications, machine learning systems, computer programs, devices, and storage media. [Solution] The method generates a composite image comprising the steps of: providing at least one instruction 310 for an image generation process; providing at least one layout specification 330 that identifies spatial constraints for a generation process 340; providing a classification specification 350 that provides a plurality of different classes for a represented scene; dividing the plurality of different classes of the classification specification into at least two groups that represent different levels of relevance to an application; determining at least one modification to the layout specification based on the divided classes; and starting a generation process to generate an image 320 based on at least one instruction and at least one modified layout specification 330.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for training and / or testing a machine learning system. Furthermore, the present invention relates to a machine learning system, a computer program, a device, and a storage medium for this purpose.

Background Art

[0002] Generative diffusion models such as Stable Diffusion have become the forerunners of a new era of applications with controllable spatial layouts when combined with ControlNet. These models are fine-tuned in their own image datasets and can convert images from driving simulators into photo-realistic outputs that closely resemble film footage from in-vehicle cameras. Additionally, these images can be dynamically changed using text prompts.

[0003] A general solution for image synthesis is disclosed in "High-resolution image synthesis with latent diffusion models" by Rombach, Robin et al., Proceedings of the IEEE / CVF conference on computer vision and pattern recognition (2022, Non-Patent Document 1). <0​​​​​​​​​​​​​​​​[Means for solving the problem]

[0005] According to aspects of the present invention, a method having the features of claim 1, a machine learning system having the features of claim 9, a computer program having the features of claim 10, a data processing device having the features of claim 11, and a computer-readable storage medium having the features of claim 12 are provided. Further features and details of the present invention are disclosed in the respective dependent claims, specifications, and drawings. Features and details described in relation to the method according to the present invention also apply to the machine learning system according to the present invention, the computer program according to the present invention, the data processing device according to the present invention, and the computer-readable storage medium according to the present invention, and vice versa.

[0006] According to one aspect of the present invention, a method, in particular a method for training and / or testing a machine learning system for a particular—particularly technical—application, (preferably as a process carried out automatically) - A step of providing at least one instruction to an image generation process in order to generate a composite image that represents a scene specific to the application or a particular technical application, - A step of providing at least one layout specification that identifies spatial constraints on the generation process, wherein the spatial constraints relate in particular to multiple different pixels of an image, - A process of providing a classification specification that provides multiple different classes for the represented scene, particularly for images, - A step of dividing multiple different classes of a classification specification into at least (or exactly) two groups representing different levels of relevance to an application, wherein the groups and / or levels of relevance are preferably manually predefined, and preferably each pixel of a layout specification is subsequently mapped to one of the groups. - A step of determining at least one modification to the layout specification based on the divided classes, preferably by removing constraints on pixels mapped to a specific group, - A step of initiating the generation process to generate an image based on at least one instruction and at least one modified layout specification. Includes.

[0007] The method can improve the training and / or testing of machine learning systems for specific technical applications by generating synthetic images that accurately reflect the target environment. Using multiple different groups, the method can restrict the generation process in essential areas of the image while allowing flexibility in less critical areas. This results in more diverse and representative training data, improving the accuracy and performance of the machine learning system. By focusing on relevant classes, the system can learn how to identify and interpret critical information within the synthetic image, leading to improved decision-making in specific applications.

[0008] Each of the method steps described above may be performed automatically. For example, instructions and / or at least one layout specification and / or classification specification may be provided as digital data, for example, based on user input. Segmentation and / or determination may be performed by a computer program using a predefined set of rules. The generation process may be initiated using a digital interface to an image generation model that uses at least one instruction and at least one modified layout specification as digital inputs. The at least one instruction may include a text prompt and / or at least one initial image, particularly from a simulator such as a driving simulator and / or from a camera, etc.

[0009] Generative models, particularly generative diffusion models such as Stable Diffusion, can transform images from simulators, such as driving simulators, into photorealistic outputs that closely resemble images from in-vehicle cameras. In other words, a more photorealistic generative model can generate a composite image based on images from a driving simulator. In addition, the generated image may be dynamically modified with the help of text input, especially prompting.

[0010] The generated images may then be used to train a machine learning system. This may be done using two basic inputs: a layout specification, particularly in the form of a conditional spatial layout, using an edge detection method such as Canny or HED to construct the layout; and instructions, such as descriptive prompts, to define overall image features, such as a snowy road.

[0011] This invention addresses the needs of autonomous driving in particular, and presents an approach that balances diversity and accuracy in image generation. One objective is to create edge representations that highlight critical objects such as roads and road users, while less important areas can still be freely interpreted by the generative model, increasing the diversity of the output image without compromising the accuracy of the main elements.

[0012] A combination of a base image, semantic label map, text description, and layout information (e.g., from HED or Canny edges) may be used. A mask can be created from the semantic label map to isolate important classes. This mask can then be used to filter layout conditions, excluding edge information for non-critical classes. The image can then be generated using this layout information. This approach aims to balance image diversity and fidelity.

[0013] The method may further include providing training and / or evaluation data for training and / or testing a machine learning system based on the generated images. That is, in particular, the training and / or evaluation data may include the generated images and / or further modified generated images (optionally by using further processing steps). This makes it possible to obtain high-quality training and / or test data for any particular application using the method according to the present invention. This is particularly advantageous for applications such as object and / or scene detection based on images recorded by vehicles.

[0014] The method may further include, in particular for training and / or testing a machine learning system for a specific technical application, in particular for training and / or testing a machine learning system for object and / or scene detection based on images recorded by a vehicle, using the generated images as training and / or evaluation data.

[0015] The generation process can be spatially different and constrained, and in particular, controlled by at least one modified layout specification. This allows the generation process to be more constrained in spatial regions of the image where pixels are classified into at least one first group with high application relevance, and less constrained in spatial regions of the image where pixels are classified into at least one second group with low application relevance. Therefore, the spatial layout of the generated image can be dynamically adjusted based on the classification of its elements. This means that in areas containing pixels classified as highly relevant to the application, such as roads or vehicles in an autonomous driving scenario, the spatial layout will be more clearly defined and constrained. Conversely, in areas of pixels classified as less relevant, such as background scenery, a more relaxed and flexible layout will be shown. The more constrained parts of the image may be controlled in the generation process in a way predefined by the application. This targeted control across the image composition increases the accuracy and realism of critical elements, while allowing more variation and diversity in less critical areas.

[0016] Layout specifications can also identify spatial constraints in relation to multiple different classes. Furthermore, at least one modification decision may include excluding edge information related to spatial constraints, particularly those relating to Canny edges, that are related to at least one group representing non-critical classes for the application (e.g., a second group of at least one of the groups mentioned above). This makes it possible to achieve more focused image generation by modifying spatial layout constraints based on class relevance. This means, in particular, that layout specifications can prioritize edge information for critical classes while not highlighting edges related to less important classes, such as background elements.

[0017] In particular, a simulator can be used to initially provide synthetic and / or sensor images representing the scene, and the generated images can be created based on the initially provided images and used as training and / or evaluation data, especially for machine learning systems. In other words, the method can leverage existing synthetic or sensor images as a starting point. This initial dataset can serve as a foundation for generating new synthetic images and can enhance the training data for machine learning systems. By utilizing existing images, the training process can be accelerated and the accuracy of the model can be improved.

[0018] The provided classification specifications offer multiple distinct classes in the form of categories, making it possible to classify images, particularly multiple distinct objects represented in each of the images. Classification may be performed based on the pixels of the image and the provided categories. This pixel-based classification enables accurate detection and / or identification and / or segmentation of objects in a composite image.

[0019] It is also possible to provide a semantic label map for the represented scene. Furthermore, the division of multiple different classes may include creating a mask from the semantic label map to isolate the classes relevant to the application, thereby dividing multiple different classes into groups, i.e., critical and non-critical classes. This makes it possible to create a mask based on a semantic label map that highlights the classes that are extremely important to a particular application. This mask can then be used to improve the layout specification, highlighting the spatial placement of important elements while allowing for greater flexibility in representing less critical areas.

[0020] The scene is a traffic scene, and the machine learning system can be further trained and / or tested for use in a driving system such as a driver assistance system and / or an autonomous driving system. The generated synthetic images can thus depict realistic traffic scenarios and enhance the training data for driver assistance and autonomous driving systems. With this approach, the machine learning system can learn, for example, specific traffic situations, object interactions, and road conditions related to a particular application. By improving the accuracy and diversity of the training data, the performance in real-world driving situations can be improved.

[0021] In another aspect of the invention, a machine learning system is provided and may be trained and / or tested using the images generated by the method according to the invention as training and / or evaluation data. Thus, the machine learning system according to the invention can have the same advantages as those described in detail in relation to the method according to the invention.

[0022] In another aspect of the invention, a computer program, particularly a computer program product, may be provided that includes instructions for causing a (at least one) computer to implement the method according to the invention when executed by the at least one computer. Thus, the computer program according to the invention can have the same advantages as those described in detail in relation to the method according to the invention.

[0023] In another aspect of the invention, an apparatus for data processing configured to execute the method according to the invention may be provided. As the apparatus, for example, a computer that executes the computer program according to the invention may be provided. The computer may include at least one processor that can be used to execute the computer program. Also, a non-volatile data memory in which the computer program is stored and which is read by the processor to implement the computer program may be provided.

[0024] According to another aspect of the present invention, there may be provided a computer program according to the present invention and / or a computer-readable storage medium including instructions for causing a computer to perform the steps of the method according to the present invention when executed by the computer. The storage medium may be configured as a data storage device such as a hard disk and / or a non-volatile memory and / or a memory card and / or a solid state drive. The storage medium may be incorporated into a computer, for example.

[0025] Furthermore, the method according to the present invention may be implemented as a computer-implemented method. Alternatively or additionally, at least one of the disclosed method steps may be implemented and / or automated by a computer.

[0026] According to the present invention, it is possible to use a trained and / or tested machine learning system in a vehicle. The vehicle may be designed as, for example, an automobile and / or a passenger car and / or a vehicle that is at least partially automated / autonomous. The vehicle may have, for example, an in-vehicle device for providing an autonomous driving function and / or a driver assistance system. The in-vehicle device may be designed to at least partially automatically control and / or accelerate and / or brake and / or steer the vehicle.

[0027] A machine learning system, particularly a machine learning system in the form of a machine learning model, is preferably trained for classification, particularly for object detection. The training may aim to train the machine learning system / machine learning model using a training data set for classification, particularly image classification, of image data such as generated images based on pixels and / or pixel values, preferably edges or pixel attributes (of the image data). The initial image used for the generation process may be obtained, for example, from the recording of at least one sensor, preferably at least one camera, preferably of the vehicle environment, particularly preferably the camera environment during (vehicle) driving and / or the vehicle environment.

[0028] Specific technical applications may include at least one of the following: classification, preferably detection, of objects in images received from a vehicle, particularly from a camera in the vehicle's driving system; scene recognition based on images received from the camera in the driving system; vehicle control based on the output of a machine learning system; classification; classification of objects in images; classification for identifying road signs, vehicles, or pedestrians in the vehicle's environment; or assistance or optimization of a vehicle control system based on image recognition. Furthermore, applications may include at least one of monitoring and / or analyzing sensor data for driver assistance systems, autonomous driving functions, or real-time traffic assessment. Additionally or alternatively, applications may include tasks in autonomous industrial systems, such as defect detection in a manufacturing process, quality control through visual inspection, or categorization of products and materials based on image analysis. Classification may be performed based on images received from at least one sensor, particularly from a camera in a vehicle, preferably an autonomous vehicle.

[0029] Since the application is used to recognize traffic scenes and / or detect objects within traffic scenes, the scene specific to the application may be a traffic scene. In this case, multiple different classes for the represented scene may include multiple different specific traffic-related scenarios that may be considered critical to the application, such as classes for other vehicles and pedestrians. The classes may also include less critical classes, such as those for background or multiple different weather conditions. Additionally or alternatively, the scene specific to the application may be an industrial scene, such as in a manufacturing process.

[0030] Classifications may be provided for various technical applications, particularly for at least one of the applications described above. One example is the use of classifications in vehicles. Based on the classifications, and particularly based on at least one classification result, at least one control operation may be initiated and / or performed, preferably for a vehicle or another technical system.

[0031] The classification results may include, and may be specific to, at least one of the following: the category of an object, the identification of an object, the location of an object and / or obstacle (e.g., in or adjacent to the direction of travel), the presence of an obstacle, a description of the traffic scene, a danger or warning message, the number of objects, the type and / or location of lane markings and / or lane boundaries, the location and / or status of a traffic signal system, the location of a lane, etc.

[0032] At least one control action for the vehicle may be initiated and / or performed based on the classification result. The control action may include at least one of the following: braking, steering, acceleration, overtaking, emergency braking, activation of an alarm system, activation of a hazard warning system, activation of a turn signal or light control.

[0033] Classification can be used, for example, to recognize obstacles regardless of whether they are directly in the direction of travel or adjacent to it. Depending on the location (for example, by the expected trajectory of the vehicle), corresponding control actions such as braking or swerving may be initiated.

[0034] "Classification" and "image classification" may also include "object detection" or "object detection in an image." In particular, this means classifying whether or not there is an object in a particular area of ​​an image. In addition, the terms "classification" and "image classification" can also refer to "semantic segmentation," especially in the form of pixel-level classification.

[0035] Therefore, training can result in at least one trained machine learning model that can be used for classification and / or object detection. Specific (technical) applications and thus inferences may be provided, for example, in vehicles.

[0036] Further advantages, features, and details of the present invention will become apparent from the following description, and embodiments of the present invention will be described in detail with reference to the drawings. In this regard, each of the features described in the claims and detailed description may be essential to the present invention, individually or in any combination. [Brief explanation of the drawing]

[0037] [Figure 1] This figure shows a method, computer program, storage medium, and apparatus according to embodiments of the present invention. [Figure 2] This figure shows an example of a composite image generated using a diffusion model. [Figure 3] This figure shows another visualization of an embodiment of the present invention. [Modes for carrying out the invention]

[0038] Figure 1 shows a computer program 20 having instructions to cause the computer 10 to perform a method 100 according to an embodiment of the present invention when the computer program 20 is executed by the computer 10. Furthermore, Figure 1 shows a computer-readable storage medium 15 according to an embodiment of the present invention, and a method 100 for training and / or testing a machine learning system 50 for a particular technical application, in particular for a vehicle 60.

[0039] Method 100 will be described below illustratively with reference to Figures 1 and 3.

[0040] According to the first method step 101, at least one instruction 310 (such as a text description or another image) is provided to the image generation process 340, and a composite image 320 representing a scene specific to the application is generated. The scene may be a traffic scene if the application targets a vehicle 60.

[0041] According to the second method step 102, at least one, in particular conditional, layout specification 330 is provided that identifies spatial constraints on the generation process 340.

[0042] Next, according to the third method step 103, a classification specification 350 is provided that provides several different classes for the represented scene. For this purpose, the classification specification 350 may include a semantic label map. The classes may include categories such as "lane," "car," "wheel," "vehicle," or "house."

[0043] In the fourth method step 104, the multiple different classes of the classification specification may be divided into at least two groups representing different levels of relevance to the application. This could be, for example, a highly relevant group and another less relevant group. The first highly relevant group may include classes such as "lane" or "vehicle". The second less relevant group may include "background" or "snow".

[0044] Next, according to the fifth method step 105, at least one modification to the layout specification 330 is determined based on the divided classes. This allows the generation process 340 to be started in the sixth method step 106 to generate the image 320 based on at least one instruction 310 and at least one modified layout specification 330. The modification may be implemented in such a way that the spatial regions of the image corresponding to at least one less relevant group of the group are given more freedom in the generation process compared to the spatial regions corresponding to other groups. In other words, the generation process 340 may be more constrained in the spatial regions of the image 320 where the pixels of the image 320 belong to at least one first group of the group that is more relevant to the application, and less constrained in the spatial regions of the image 320 where the pixels of the image 320 belong to at least one second group of the group that is less relevant to the application.

[0045] The methods described in the embodiments of the present invention can therefore be carried out by spatially differently constraining the generation process 340 and controlling it by at least one modified layout specification 330.

[0046] The synergistic effect of Stable Diffusion and ControlNet is known to be used to create images that mimic images taken by in-vehicle cameras. Figure 2 shows an example of a composite image generated by a diffusion model. These composite images can be created using two basic inputs: a conditional spatial layout that primarily uses edge detection techniques such as Canny and HED for layout construction, and a descriptive prompt to define overall image characteristics such as a snowy road. Typically, the layout conditions are based on Canny edges, which are computationally efficient and can be generated from both real and simulated images.

[0047] Embodiments of the present invention address the needs of autonomous driving technology and introduce an approach that balances diversity and fidelity in image generation. This particularly focuses on manipulating layout conditions dedicated to critical classes in the domain, such as Canny and HED edges, by utilizing class information available from semantic label maps generated from driving simulators or other sources.

[0048] In autonomous driving, not all objects are equally relevant; for example, background elements such as trees can excessively restrict image generation. Embodiments of the present invention therefore utilize the possibility of creating edge representations that highlight critical objects such as roads and road users, while allowing the generative model to freely interpret less important areas, thereby increasing the diversity of the output image without compromising the accuracy of key elements.

[0049] Current methods, such as "Adding conditional control to text-to-image diffusion models" by Zhang, Lvmin, Anyi Rao, and Maneesh Agrawala, as cited in the IEEE / CVF International Conference Proceedings on Computer Vision (2023), involve assigning global weights to spatial layout constraints related to text prompts. However, this approach lacks the specificity to focus on essential image areas, thereby limiting its effectiveness in applications with subtle differences, such as autonomous driving.

[0050] Figure 2 shows an example of a composite image that can be generated by the proposed approach. Since the background is not constrained by layout conditions (Canny edges), it differs across different images, resulting in increased layout diversity.

[0051] Figure 3 visualizes the synthetic image generation approach without selective filtering of critical edge areas. It shows how buildings (and their layout, such as windows) are constrained by Canny edges. However, in the proposed approach according to embodiments of the present invention, the network would be able to "imagine" any suitable background according to the training data distribution without being constrained to the extent illustrated here.

[0052] The methods according to embodiments of the present invention involve distinguishing “semantic classes,” particularly “important classes,” within the context of autonomous driving. Each pixel in an image belongs to a specific class and can be identified through semantic segmentation or provided by a driving simulator such as Nvidia® DriveSim. At the core of the approach is to selectively focus on important classes, such as roads and vehicles, and not emphasize other classes during image generation. This process is outlined in two phases, as also shown in Figure 1.

[0053] Phase I: Training: In this phase, any generative model trained with Canny edges for the approach described in Phase II can be applied; however, for optimal results, it is recommended to apply similar class filtering during training.

[0054] Phase II: Inference: In this phase, the procedure involves using a combination of a base image, a semantic label map, a text description, and layout information (e.g., from HED or Canny edges). A mask is created from the semantic label map to isolate important classes. This mask is then used to filter the layout conditions, excluding edge information related to non-critical classes. The image is then generated using this adapted layout information. This approach aims to strike a good balance between image diversity and fidelity, where features are not generally addressed in the current spatial layout conditions of the generative model.

[0055] The main advantage of this strategy lies in its ability to increase the diversity of generated images without excessive constraints, especially in areas that are not critical to autonomous driving. This allows for control over the subtle differences between fidelity and diversity in image generation.

[0056] The above description of embodiments relates to the present invention in reference to examples. Of course, individual features of the embodiments can be freely combined with each other, as long as they are technically reasonable and do not depart from the scope of the present invention.

Claims

1. A method (100) for training and / or testing a machine learning system (50) for a specific technical application, - A step (101) of providing at least one instruction (310) to an image generation process (340) in order to generate a composite image (320) that represents a scene specific to the aforementioned technical application, - A step (102) of providing at least one layout specification (330) that identifies spatial constraints for the image generation process (340), - A step (103) of providing a classification specification (350) that provides multiple different classes for the represented scene, - A step (104) of dividing the plurality of different classes of the classification specification into at least two groups representing different levels of relevance to the technical application, - A step (105) of determining at least one modification to the layout specification (330) based on the divided classes, - Step (106) to initiate the image generation process (340) for generating the composite image (320) based on the at least one instruction (310) and the at least one modified layout specification (330) and A method (100) including the following.

2. - A step of providing training and / or evaluation data for the training and / or testing of the machine learning system (50) based on the generated composite image (320), wherein the training and / or evaluation data includes, in particular, the generated composite image (320) and / or a further modified generated composite image (320), and - In particular, the process of training and / or testing the machine learning system (50) for the training and / or testing of the machine learning system (50) for the training and / or testing of the machine learning system (50) using the generated composite image (320) as training and / or evaluation data for the training and / or testing of the machine learning system (50) The method according to claim 1 (100), further comprising at least one of the following.

3. The method (100) according to claim 1 or 2, wherein the image generation process (340) is spatially differently constrained and controlled by the at least one modified layout specification (330), so as to be more constrained in the spatial region of the composite image (320) where the pixels of the composite image (320) are classified into at least one first group of the group that is highly relevant to the technical application, and less constrained in the spatial region of the composite image (320) where the pixels of the composite image (320) are classified into at least one second group of the group that is less relevant to the technical application.

4. The method according to any one of claims 1 to 3 (100), wherein the layout specification (330) identifies the spatial constraints in relation to the plurality of different classes, and the step of determining the at least one modification (105) includes excluding the spatial constraints relating to at least one group of the groups representing noncritical classes for the technical application, in particular edge information relating to canny edges.

5. The method according to any one of claims 1 to 4 (100), wherein a composite and / or sensor image representing the scene is first provided, in particular generated, and the generated composite image (320) is generated based on the initially provided composite image and is used in particular as training or evaluation data for the machine learning system (50).

6. The method (100) according to any one of claims 1 to 5, wherein the provided classification specification (350) provides the plurality of different classes in the form of categories to classify the plurality of different objects represented in the composite image, in particular each of the composite image, and the classification is carried out based on the pixels of the composite image (320) and the provided categories.

7. The method according to any one of claims 1 to 6 (100), wherein a semantic label map is provided for the represented scene, and the step of dividing the plurality of different classes (104) includes creating a mask from the semantic label map to isolate the classes related to the technical application, thereby dividing the plurality of different classes into groups, i.e., critical classes and non-critical classes.

8. The method (100) according to any one of claims 1 to 7, wherein the scene is a traffic scene, the machine learning system (50) is trained and / or tested for use in a driver assistance and / or autonomous driving system, and the technical application includes, in particular, at least one of classifying, preferably detecting, an object in an image received from a camera of the autonomous driving system, scene recognition based on the image, and controlling a vehicle based on the output of the machine learning system (50).

9. A machine learning system (50) that is trained and / or tested using an image (320) generated by the method (100) of any one of claims 1 to 8 as training and / or evaluation data.

10. A computer program (20) which, when the computer program (20) is executed by at least one computer (10), includes an instruction to cause the computer (10) to carry out the method (100) according to any one of claims 1 to 8.

11. A data processing device (10) comprising means for carrying out the method (100) described in any one of claims 1 to 8.

12. A computer-readable storage medium (15) that, when executed by a computer (10), includes instructions causing the computer (10) to carry out the steps of the method (100) according to any one of claims 1 to 8.