Method of generating data set for training and / or testing machine learning system

By leveraging metadata to generate detailed text prompts and image datasets, this approach addresses the issue of insufficient image diversity in existing generative text-image models, improving the reliability of training and testing machine learning systems. In particular, it enables more efficient dataset generation and model training in vehicle control applications.

CN121600342APending Publication Date: 2026-03-03ROBERT BOSCH GMBH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511016397.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-08-14
Filing Date
2025-07-23
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing generative text-image models generate images that are too general and lack diversity, making them unsuitable for effectively training and testing machine learning systems. In particular, existing text prompting methods are inefficient and rely on manual annotation in vehicle control applications.

Method used

By providing metadata to generate detailed text prompts and combining them with image data, generative models are used to generate datasets, particularly image datasets of vehicle environment scenes, for training and testing machine learning systems, including autonomous driving and driver assistance systems.

Benefits of technology

The generated dataset is highly variable, which improves the reliability of training and testing, enhances the generalization ability of machine learning systems, improves the accuracy and diversity of image generation, reduces the workload and cost of data collection, and enhances the adaptability and flexibility of training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600342A_ABST
    Figure CN121600342A_ABST
Patent Text Reader

Abstract

The invention relates to a method of generating a data set for training and / or testing a machine learning system, comprising: providing (101) image data specifically for presenting images of different environmental scenes; providing (102) metadata (65), the metadata being specifically used to describe the different environmental scenarios; generating (103) a text hint (70) based on the provided metadata (65), wherein information contained in the metadata (65) is used for the text hint (70); a dataset (60) is generated (104) based on the generated textual cues (70) and preferably based on the provided image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for generating datasets for training and / or testing machine learning systems. Furthermore, the invention also relates to a machine learning model, computer program, apparatus, and storage medium for this purpose. Background Technology

[0002] Significant progress has been made in the generation of synthetic images, for example through generative text-image models like Stable Diffusion. These models are capable of generating images based on text input.

[0003] However, the generated images remain too general and lack diversity for many applications. This is often related to the fact that the text prompts (i.e., prompts) that form the basis of image synthesis typically contain only very brief and insufficient descriptions of the scene.

[0004] In addition, other image synthesis-related schemes are known in the prior art, such as the use of VLP (Vision-Language Pretraining) systems.

[0005] Other known traditional solutions come from:

[0006] -Rombach et al.'s High-Resolution Image Synthesis with Latent Diffusion Models (arXiv:2112.10752), and

[0007] -Li et al.’s BLIP-2, Bootstrapping Language-Image Pre-training with FrozenImage Encoders and Large Language Models (arXiv:2301.12597). Summary of the Invention

[0008] This invention provides a method, a machine learning model, a computer program, an apparatus, and a computer-readable storage medium. Other features and details of the invention can be derived from the corresponding dependent claims, the specification, and the drawings. Features and details described in connection with the method of this invention are naturally also associated with the machine learning model, the computer program, the apparatus, and the computer-readable storage medium of this invention, and vice versa; therefore, reference should always be made to the disclosure of this invention.

[0009] The subject matter of this invention relates in particular to a method for generating datasets for training (especially fine-tuning) and / or testing machine learning systems (preferably the machine learning models of this invention).

[0010] The machine learning system may include at least one machine learning model (preferably having at least one neural network), which can be tested and / or trained on the dataset. The machine learning model to be trained and / or tested may be used, for example, for image classification.

[0011] Machine learning systems may also include generative models configured to generate synthetic image data. In particular, such generative models are generative text-image models such as Stable Diffusion.

[0012] This method can be used in conjunction with vehicle control. Specifically, it is used here to generate datasets for the development and / or improvement of vehicle control software. This can include machine learning, allowing the generated datasets to be used for training and / or testing in machine learning.

[0013] The method of the present invention may include providing metadata, particularly digital metadata. The metadata may be specifically used to describe different environmental scenarios, particularly an environment (preferably a vehicle environment). The description may include, for example, information about location and / or weather and / or driving conditions and / or similar content. The provision of metadata can be achieved, in particular, by retrieving it from a data storage device or data interface. Environmental scenarios may include, for example, the arrangement of different objects and / or road layout and / or the spatial structure of obstacles and / or road markings.

[0014] The metadata may exist alongside the associated image data. Therefore, the method of the present invention may also include providing image data specifically designed to present images of different environmental scenes. For example, this could be a single environmental image for each different environmental scene (preferably a vehicle environment image, more preferably an image during driving). The metadata is generated, for example, during image data recording, such as automatically generated and / or manually generated (e.g., by driver input). The metadata may, for example, indicate the location and / or time of image data recording. Furthermore, the metadata may be automatically and / or manually supplemented, for example, by adding more information about the environmental scene at the time the image data was recorded. For example, this could be determining (e.g., retrieving from a database) weather information at the recording location at the time of recording.

[0015] Furthermore, the method of the present invention may include generating a text prompt based on the provided metadata. Information contained in the metadata (such as weather information) may be used for the text prompt, particularly being automatically converted into a text prompt and / or adjusting existing text prompts.

[0016] Subsequently, in the method of the present invention, a dataset can be generated based on the generated text prompts and preferably also based on the provided image data, wherein the generation is preferably performed by a machine learning system or another generative machine learning system for data (especially images) synthesis. In other words, automatic text generation can be achieved for the creation of prompts, which can generate datasets with higher efficiency and / or quality and / or diversity. Preferably, text generation is performed directly using metadata to automatically generate detailed and diverse descriptions. The dataset can be generated, for example, based on the generated text prompts using a generative model (preferably a generative text-image model, such as Stable Diffusion). The provided image data can optionally also be used as training and / or testing data, or for training and / or testing, for example, to further control the images generated based on the text prompts.

[0017] The dataset can be used, for example, to fine-tune and / or train machine learning models for synthetic image data. It is also possible that the dataset itself already includes synthetic image data. The dataset and / or synthetic image data can be used, for example, to train machine learning models for autonomous vehicle driving and / or for scene interpretation and / or perception during autonomous driving. Therefore, the trained model can be applied, for example, to an artificial intelligence-controlled visual recognition system.

[0018] Optionally, the generated dataset may be specified to include multiple training data sets for training and / or testing the machine learning system. The training data may be image data, preferably synthetic image data. Thus, the generated dataset or training data can provide a representation of environments in different and / or newly generated environmental scenes. In this way, multiple synthetic images can be generated, which can then be used for training and fine-tuning of image models such as Stable Diffusion.

[0019] It is also conceivable that the training is used to train a machine learning system with the generated dataset to classify digital images based on pixels and / or pixel attributes (preferably edges or pixel properties), particularly image classification. The digital images may be, for example, digital images recorded by at least one sensor (preferably at least one camera, more preferably a vehicle camera), particularly recordings of the camera and / or the vehicle's environment during vehicle movement. The recording may be, for example, achieved through at least one camera of the vehicle. The classification can be used to identify objects in the environment presented by the digital images and / or capture traffic scenes.

[0020] The classification can be used in various technical applications. One example is its application in vehicles. Based on the classification (especially at least one classification result), for example, at least one control action (preferably a control action targeting the vehicle or other technical systems) can be initiated and / or executed.

[0021] The classification results may include at least one of the following results and / or specifically address at least one of the following results: object category, object identification, location of object and / or obstacle (e.g., in or beside the direction of travel), presence of obstacle, description of traffic scene, hazard warning, number of objects, type and / or location of road markings and / or road boundaries, location and / or status of traffic signal equipment, location of road, or similar content.

[0022] Based on the classification results, at least one control action for the vehicle can be initiated and / or executed. The control action may include at least one of the following: braking, steering, acceleration, overtaking, emergency braking, activation of an alarm system, activation of hazard warning lights, activation of turn signals, lighting control, or similar actions.

[0023] The classification process can, for example, identify obstacles, whether they are located directly in front of or beside the vehicle in the direction of travel. Based on the location (e.g., depending on the expected vehicle trajectory), appropriate control actions, such as braking or avoidance, can be initiated.

[0024] For example, if the classification indicates the presence of an obstacle and / or a potential collision in the direction of travel, braking can be initiated. It is also conceivable to identify roads and / or road boundaries based on classification, so that the vehicle can be driven at least partially automatically on the road through controlled actions.

[0025] A vehicle may be configured as a motor vehicle and / or a passenger car and / or a vehicle that is at least partially autonomous.

[0026] The method of this invention has the advantage of generating highly variable training and / or testing data (particularly in terms of the representation of different environmental scenarios). This improves the reliability of training and / or testing, and consequently, the reliability of the trained learning system for classification tasks. For example, testing can be performed during training, i.e., by dividing the generated dataset into test and training data, and using the test data to check training progress. In particular, the high variability in the dataset improves the generalization ability of the learning system, making the classification applicable, for example, to new situations in vehicle control.

[0027] "Classification" and "image classification" can also include "object detection" or "object detection in an image." This should be understood in particular as a classification process, that is, determining whether an object exists in certain regions of an image. Furthermore, the terms "classification" and "image classification" can also refer to "semantic segmentation," especially in the form of pixel-level classification.

[0028] Furthermore, it is conceivable that image data and / or metadata originate from sensing acquisition of the vehicle environment and / or cameras (preferably image acquisition), and / or from acquisition by at least one sensor (particularly a camera, more preferably a vehicle camera), during which the metadata is preferably defined to describe the environmental scene. This offers the advantage that the method for generating text prompts can consider a large number of different environmental scenes for image synthesis. Moreover, the method can be more reliable and accurate because the description is based on precise data about the environment. More precise text descriptions for synthesized image generation improve the accuracy of representing the realistic environment and objects in the generated image data.

[0029] Furthermore, it is also possible that the machine learning system includes (at least) a generative model (preferably for generating synthetic images), and / or (at least) a machine learning model for use in at least partially autonomous driving. The synthetic image material in the generated image data can then be used, for example, as a training or testing dataset for the machine learning model. This approach can significantly improve the training of the learning system by minimizing the workload and cost of data acquisition while increasing the quality and quantity of available data. Another advantage is the flexibility in adapting training data to specific needs.

[0030] It is conceivable that providing image data could also include providing image data derived from sensor acquisition and supplemented with metadata, preferably supplemented manually by the driver during driving in the corresponding environmental scenario. Therefore, additional information can be added to the image data to enrich it. This supplementation can also be selectively based on data sources such as vehicle cameras or other sensors that record image data during driving. Furthermore, data sources can also be used to identify metadata not directly used for sensor acquisition, such as cloud servers or GPS systems used to determine weather data. Drivers can also manually provide this or other metadata. By accessing these extended data sources, the environmental scenario can be presented more comprehensively, thereby further improving the training and testing of machine learning systems. This can also improve the accuracy and quality of text prompts.

[0031] Optionally, generating text prompts may include: converting information contained in the metadata into text prompts so that, for the dataset to be generated (preferably including multiple images to be generated), at least one of the following pieces of information in the metadata is taken into account and preferably presented in the images:

[0032] -Environmental conditions, especially the weather or time of day.

[0033] - Contextual details of the environment or corresponding environmental scenario, preferably the road type or traffic conditions.

[0034] -Location information, preferably collected from GPS.

[0035] The advantage of doing this is that the generated text prompts will be more detailed, thus providing more context for the image to be generated.

[0036] Within the scope of this invention, generating text prompts may also include transforming metadata, converting the metadata into structured text prompts during the transformation process. In this process, the metadata used for transformation may be randomly selected. Employing a probabilistic method ensures a high degree of diversity in the text prompts. The advantage of this approach is that it also achieves higher variability in the dataset.

[0037] The dataset can preferably be generated using a generative machine learning model. Alternatively or supplementarily, it is conceivable to supplement the generated text prompts with a conditional space layout to take into account the conditions applied by the machine learning system when generating the dataset. Therefore, by combining the generated text prompts with a conditional space layout, a realistic environmental representation can be ensured. This enables improved and more realistic dataset generation. Edge detection methods are preferably employed in this process.

[0038] The subject of this invention also includes a machine learning model obtained by training through (at least) the following steps:

[0039] - Generate a dataset using the method of this invention.

[0040] - At least train the machine learning model using the generated dataset.

[0041] Therefore, the machine learning model of the present invention has the same advantages as the method described in detail with reference to the present invention.

[0042] It is conceivable that the method of the present invention can be used to manufacture at least a partial autonomous driving system and / or driver assistance system for a vehicle. The vehicle may be configured, for example, as a motor vehicle and / or passenger car and / or autonomous vehicle. The vehicle may have vehicle devices, for example, for providing autonomous driving functions and / or driver assistance systems. The vehicle devices may be configured to control the vehicle at least partially automatically, particularly acceleration and / or braking and / or steering. For this purpose, control may be based on the output of a machine learning system (preferably the machine learning model of the present invention).

[0043] The subject matter of this invention also includes a computer program, and more particularly a computer program product, comprising instructions that, when executed by a computer, cause the computer to perform the method of the invention. Therefore, the computer program of the present invention has the same advantages as the method described in detail with reference to the invention.

[0044] The subject matter of this invention also includes an apparatus for data processing configured to perform the methods of this invention. For example, the apparatus may be a computer executing a computer program of this invention. The computer may have at least one processor for executing the computer program. A non-volatile data memory may also be provided, in which the computer program is stored, and from which the processor may read the computer program for execution.

[0045] The subject matter of this invention may also include a computer-readable storage medium having the computer program of this invention and / or including instructions that, when executed by a computer, cause the computer to perform the methods of this invention. The storage medium may be configured as a data storage device, such as a hard disk and / or non-volatile memory and / or a memory card. The storage medium may, for example, be integrated into a computer.

[0046] Furthermore, the method of the present invention can also be executed as a computer-implemented method. Alternatively or supplementarily, at least one of the disclosed method steps can be computer-implemented and / or automatically executed. Attached Figure Description

[0047] Other advantages, features, and details of the invention will become apparent from the following description, in which embodiments of the invention will be described in detail with reference to the accompanying drawings. The features mentioned in the claims and specification may be significant to the invention individually or in any combination. In the drawings:

[0048] Figure 1 A schematic visualization of a method, apparatus, machine learning model, storage medium, and computer program according to embodiments of the present invention is shown.

[0049] Figure 2 A schematic representation of a method according to an embodiment of the present invention is shown.

[0050] Figure 3 Another illustrative representation of a method according to an embodiment of the present invention is shown. Detailed Implementation

[0051] exist Figure 1 The diagram schematically illustrates a method 100, apparatus 10, machine learning model 51, storage medium 15, and computer program 20 according to an embodiment of the present invention.

[0052] The method 100 shown is used to generate a dataset 60 for training and / or testing the machine learning system 50.

[0053] According to step 101 of the first method, image data specifically designed to present images of different environmental scenes is provided.

[0054] According to step 102 of the second method, metadata 65 is provided specifically to describe different environmental scenarios. Metadata 65 may include specific information about the environment presented by the recorded image data. Metadata 65 may also specify the conditions at the time of recording. This specifically refers to information about the actual recording conditions collected during the recording process.

[0055] Then, according to step 103 of the third method, a text prompt 70 can be generated based on the provided metadata 65 so that the information contained in the metadata 65 can be used in the text prompt 70. For example, if it is known from the metadata 65 that it was snowing or raining when the record was made, this content can be included in the corresponding text prompt 70.

[0056] Furthermore, according to step 104 of the fourth method, a dataset 60 can be generated based on the generated text prompt 70, and preferably based on the provided image data.

[0057] Figure 2 A graphical description of embodiments of the invention is also shown. Within the scope of the embodiments, the synergistic effect of models such as StableDiffusion and ControlNet can be used to create synthetic images similar to those captured by vehicle cameras.

[0058] The composite image 80 can be created using the following main inputs:

[0059] - Conditional spatial layout, which is mainly constructed using edge detection methods (such as Canny) (see...) Figure 3 The edge diagram 75 is shown.

[0060] - Descriptive text prompts are used to define general image features, such as snow-covered roads.

[0061] - Optional additional inputs 71, such as random noise.

[0062] Text prompts can be automatically generated when necessary using automatic annotation tools such as BLIP (see prior art references). However, these annotations are often very brief and lack detail, which leads to a decrease in image generation quality. Therefore, embodiments of the present invention contribute to this field by creating improved and more detailed text descriptions. To this end, available metadata can be utilized to generate images better (more diverse and more accurate) based on descriptive and detailed text prompts.

[0063] The embodiments of this invention aim to address a problem inherent in traditional solutions: the generation of synthetic images often lacks diversity and specificity due to limitations in current text prompting methods. Images generated by traditional text-image models are typically too general or lack diversity to effectively train advanced AI perception models. Furthermore, relying on manual annotation to create text descriptions is often too inefficient for large datasets.

[0064] Embodiments of the present invention can at least partially solve these problems, providing a solution such as Figure 1 The example visualization illustrates a method 100 for creating text prompts that utilizes available metadata from a recorded dataset. This method ensures that the generated text prompts are both diverse and specific, resulting in synthetic images that are more representative of real-world scenes and sufficiently diverse to train robust AI models.

[0065] By leveraging detailed metadata, the tooltip creation process can be automated, improving efficiency and reducing reliance on manual operations or automation methods that do not meet the aforementioned standards. This metadata-based tooltip development approach at least partially addresses the limitations of traditional methods in terms of diversity and specificity, as well as the inefficiency of manual annotation.

[0066] Embodiments of the present invention also provide a tooltip engineering method that utilizes only metadata associated with an image. In other words, in the method according to embodiments of the present invention, the provided metadata can be the sole data basis for generating text tooltips.

[0067] The method according to embodiments of the present invention further includes extracting metadata information (such as time of day, road type, and weather conditions) and converting it into structured text prompts. This method ensures that the text prompts are not only accurate but also contain a wide range of image features, thereby improving the diversity and specificity of the generated images.

[0068] I. Metadata Collection

[0069] According to the first possibility, a metadata-based prompting engineering approach can leverage a comprehensive metadata database to provide data, which is typically well-organized and updated in such applications. This is usually accomplished through manual input from the driver, who collects this data during driving under varying environmental conditions. This metadata includes a wide range of information that may be image-related, such as environmental conditions, geographic location, time of day, and specific events or situations that occurred during data collection.

[0070] The method 100 according to the embodiment can be... Figure 1The step 101 shown extracts metadata to serve as the basis for a text prompt. The metadata can then be carefully analyzed to ensure its relevance and accuracy in describing the image context. For example, if the metadata indicates that an image was taken in the suburbs on a rainy evening, then method 100 can create a text prompt that accurately reflects these conditions. This process includes, for example, classifying and prioritizing metadata elements to create a coherent and detailed description.

[0071] II. Probability Randomization

[0072] To increase diversity and avoid redundancy in text prompts, probabilistic models can be used. These models can select and combine different metadata elements based on random principles to ensure that each prompt is as unique as possible and reflects a wide range of possible scenarios. This randomness is crucial for generating a large number of image descriptions, which facilitates efficient training and fine-tuning of the generative model.

[0073] By using metadata collected from drivers, the above method not only provides high accuracy and relevance in text prompts but also helps create a comprehensive and diverse dataset. This dataset is highly valuable for improving the capabilities of text-to-image models and enabling the generation of more accurate and diverse images that better reflect the complexity of real-world scenes.

[0074] This metadata-based approach to suggestion engineering has several significant advantages:

[0075] - The accuracy of textual descriptions can be improved by directly using metadata. Metadata can reflect specific and objective information about an image.

[0076] - This can significantly increase the diversity of text prompts because metadata typically contains a wide range of image features, from environmental conditions to context-related details.

[0077] - It can help automate and rationalize the process of creating detailed and diverse text descriptions for image datasets.

[0078] A possible application of embodiments of the present invention is to improve the training and fine-tuning of text-image models (such as Stable Diffusion). By providing a method for generating diverse and detailed text-image pairs, not only can the quality and variability of generated images be improved, but it also contributes to the further development of data augmentation techniques and the development of more complex generative models.

[0079] The above explanation of the embodiments describes the present invention by way of example only. Of course, the various features of the embodiments can be freely combined as long as they are technically reasonable, without departing from the scope of the present invention.

[0080] Cross-references to related applications

[0081] This application claims priority to European patent application filed on August 14, 2024 (previous application number: 24194564.1). All disclosures of the above application are incorporated herein by reference.

Claims

1. A method (100) for generating a dataset (60) for training and / or testing a machine learning system (50), the method comprising: - Provide (101) image data, which is specifically designed to present images of different environmental scenes; - Provide (102) metadata (65) specifically for describing the different environmental scenarios; - Generate (103) a text prompt (70) based on the provided metadata (65), wherein the information contained in the metadata (65) is used for the text prompt (70); - The dataset (60) is generated based on the generated text prompts (70) and preferably based on the provided image data (104).

2. The method (100) according to claim 1, Its features are, The generated dataset (60) is a training dataset that includes multiple synthetic image data used to train the machine learning system (50) in order to provide different and / or newly generated environmental scenes for the training.

3. The method (100) according to claim 2, Its features are, The training is used to train the machine learning system (50) with the generated dataset (60) to classify digital images based on the pixels and / or pixels of the images, particularly digital images generated by recordings of the vehicle environment during vehicle operation and / or generated by cameras, wherein the vehicle is preferably controlled based on the classification.

4. The method (100) according to any one of the preceding claims, Its features are, The metadata (65) comes from sensing acquisition by at least one sensor, preferably image acquisition, wherein the at least one sensor is particularly at least one camera, preferably a camera of a vehicle, and during the acquisition process, the metadata (65) is defined to describe the environmental scene.

5. The method (100) according to any one of the preceding claims, Its features are, The machine learning system (50) includes a generative model, preferably for generating synthetic images, and / or a machine learning model for use in at least partially autonomous driving.

6. The method (100) according to any one of the preceding claims, Its features are, Providing the image data (101) further includes: - Provide image data derived from sensor acquisition and supplemented by the metadata (65), preferably supplemented by the driver during driving in the corresponding environmental scenario.

7. The method (100) according to any one of the preceding claims, Its features are, Generating (103) the text prompt (70) further includes: converting the information contained in the metadata (65) into text prompts respectively, so that for a dataset to be generated, preferably including multiple images to be generated, at least one of the following information in the metadata (65) is taken into account and preferably presented in the images: -Environmental conditions, especially the weather or time of day; - Contextual details of the corresponding environmental scenario, preferably road type or traffic conditions; -Location information, preferably collected from GPS.

8. The method (100) according to any one of the preceding claims, Its features are, Generating (103) the text prompt (70) further includes converting the metadata (65) in which the metadata (65) is converted into a structured text prompt, wherein the metadata (65) is randomly selected for the conversion.

9. The method (100) according to any one of the preceding claims, Its features are, To optimize the generation of the dataset (104) by the generative machine learning model, the generated text prompts (70) are supplemented with a conditional space layout so that the conditions for applying the machine learning system (50) are taken into account when generating the dataset.

10. A machine learning model (51), said machine learning model being trained through the following steps: - Generate the dataset (60) by the method (100) according to any one of the preceding claims; - The machine learning model (51) is trained using at least the generated dataset (60).

11. A computer program (20) comprising instructions that, when executed by a computer (10), cause the computer to perform the method (100) according to any one of claims 1 to 9.

12. An apparatus (10) for data processing, the apparatus being configured to perform the method (100) according to any one of claims 1 to 9.

13. A computer-readable storage medium (15) comprising instructions that, when executed by a computer (10), cause the computer to perform the steps of the method (100) according to any one of claims 1 to 9.