Method for generating a dataset for training and / or testing a machine learning system
By using metadata to generate detailed text prompts, the method addresses the lack of diversity in existing image synthesis methods, creating high-quality datasets that enhance the training of AI models for autonomous vehicle navigation.
Patent Information
- Application Number
- JP2025105623
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-14
- Filing Date
- 2025-06-23
- Publication Date
- 2026-02-27
AI Technical Summary
Existing image synthesis methods, such as those using generative text-image models like Stable Diffusion, generate images that are too generic and lack diversity due to insufficient and short text descriptions, limiting their effectiveness in training advanced AI recognition models.
A method that leverages metadata associated with image data to generate detailed and diverse text prompts, which are then used to create synthetic images, incorporating environmental scenario details like weather and location, thereby improving the quality and variety of training datasets for machine learning systems.
The approach generates high-quality, diverse datasets that enhance the training and testing of machine learning models, particularly for autonomous vehicle navigation, by providing accurate and varied representations of real-world scenarios, thus improving the reliability and generalization of AI systems.
Smart Images

Figure 2026034366000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for generating a dataset for training and / or testing a machine learning system, and further to a machine learning model, a computer program, an apparatus, and a storage medium for this purpose. [Background technology]
[0002] It is known from the prior art that great progress has been made in creating synthetic images through generative text-image models, such as Stable Diffusion, which are able to generate images based on text input.
[0003] However, the generated images are still too generic and not diverse enough for many applications, which is often due to the fact that the text instructions, or prompts, on which image synthesis is based, contain very short and insufficient descriptions of the scenario.
[0004] Furthermore, various other approaches related to image synthesis are known from the prior art, such as the use of VLP systems (Vision-Language Pretraining).
[0005] As conventional solutions, the following documents are also known: Non-Patent Document 1 (Rombach et al., "High-Resolution Image Synthesis with Latent Diffusion Models" (arXiv:2112.10752)) and Non-Patent Document 2 (Li et al., "BLIP-2, Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models" (arXiv:2301.12597)). [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] (Rombach et al., "High-Resolution Image Synthesis with Latent Diffusion Models" (arXiv:2112.10752) [Non-patent document 2] “BLIP-2, Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models”(arXiv:2301.12597) Summary of the Invention
[0007] The subject matter of the present invention is a method having the features of claim 1, a machine learning model having the features of claim 9, a computer program having the features of claim 10, an apparatus having the features of claim 11, and a computer-readable storage medium having the features of claim 12. Further features and details of the invention are set out in the respective dependent claims, the description and the drawings, where it goes without saying that the features and details set out in relation to the method according to the invention also apply in relation to the machine learning model, the computer program, the apparatus and the computer-readable storage medium according to the invention, and therefore cross-reference is always possible to the disclosure of the present invention.
[0008] The subject of the present invention is in particular a method for generating a dataset for training, in particular fine-tuning and / or testing a machine learning system, preferably a machine learning model according to the present invention.
[0009] The machine learning system may include at least one machine learning model, preferably at least one machine learning model having a neural network, which may be tested and / or trained with a dataset. The trained and / or tested machine learning model may be used, for example, for image classification.
[0010] The machine learning system may further include a generative model configured to generate synthetic image data, particularly a generative text-image model such as Stable Diffusion.
[0011] The method can be used in connection with vehicle control. The method is particularly useful for generating datasets used for developing and / or improving vehicle control software. This can include machine learning, whereby the generated datasets are used for training and / or testing the machine learning.
[0012] The method according to the invention may include the provision of metadata, in particular digital metadata. The metadata may be specific to descriptions of different environmental scenarios, in particular environments, preferably vehicle environments. The descriptions may include, for example, location and / or weather and / or driving (travel) conditions and / or information thereon. The provision of metadata may be realized, for example, by retrieval from a data storage device or a data interface. The environmental scenario may include, for example, the arrangement of various objects and / or the arrangement of lanes and / or spatial obstacles and / or lane markings.
[0013] The metadata may be present in association with the image data assigned to these metadata. Thus, the method according to the present invention may further comprise providing specific image data for images depicting different environmental scenarios. This may be, for example, an image of an environment, preferably a vehicle environment, and a moving vehicle, for each of the different environmental scenarios. The metadata may, for example, be generated automatically when the image data is recorded and / or manually, for example by driver input. The metadata may, for example, indicate the location and / or time of recording of the image data. The metadata may further, for example, be supplemented automatically and / or manually with further information regarding the environmental scenario in which the image data was acquired / recorded. This information may, for example, be weather information at the recording location at the time of recording, obtained from a database.
[0014] Furthermore, the method according to the present invention may comprise generating a text prompt (prompt) based on the provided metadata, where information contained in the metadata (e.g. weather information) can be used in the text prompt, in particular automatically converted into a text prompt and / or adjusted to an existing text prompt.
[0015] The method according to the present invention then allows for the generation of a dataset based on the generated text prompts and, preferably, the provided image data. In this case, the generation of the dataset is preferably performed by a machine learning system or other generative machine learning system for synthesizing data, particularly images. In other words, automated text generation for prompt creation is possible, thereby enabling the generation of datasets with greater efficiency, quality, and / or diversity. In this case, the metadata can be used directly in the text generation to automatically generate detailed and diverse descriptions. The dataset is generated based on text prompts, for example, generated by a generative model, preferably a generative text-image model (such as Stable Diffusion). The provided image data can optionally serve as training data and / or test data or can be used for training and / or testing, e.g., for further control of the images generated based on the text prompts.
[0016] The dataset can be used, for example, to fine-tune and / or train a machine learning model that synthesizes image data. It is also possible that the dataset itself already contains synthetic image data. The dataset and / or synthetic image data can be used, for example, to train a machine learning model that functions in autonomous vehicle navigation and / or situation interpretation and / or perception during autonomous navigation. In this way, the trained model can be applied, for example, to an AI-controlled visual recognition system.
[0017] Optionally, the generated dataset may include multiple training data for training and / or testing a machine learning system. The training data may be configured as image data, preferably synthetic image data. Thus, the generated dataset or training data may provide and / or be designed to provide representations, in particular environmental representations, for different and / or newly generated environmental scenarios. This allows for the generation of a variety of synthetic images, which are subsequently used to train and fine-tune image models, such as Stable Diffusion.
[0018] Furthermore, it is conceivable that the training is intended to train a machine learning system for the classification of digital images, in particular image classification, with the generated dataset based on image pixels and / or pixels, preferably on the edges or pixel attributes of the images. These digital images may, for example, be digital images obtained by at least one sensor, preferably by at least one camera, in a vehicle, in particular preferably from the camera environment and / or the vehicle environment while the vehicle is moving. This recording can, for example, be realized by at least one camera in the vehicle. In this case, the classification may be aimed at recognizing objects and / or understanding traffic scenes in the environment represented by the digital images.
[0019] The classification can be used in various technical applications, one example being vehicle applications, where, based on the classification, in particular on at least one classification result, for example, at least one control action can be initiated and / or executed, preferably at least one control action for a vehicle or other technical system.
[0020] The classification result may include and / or be specific to at least one of the following results: object category, object identification, object and / or obstacle location (e.g., in the direction of travel or next to the direction of travel), obstacle presence, traffic scene description, hazard notification, number of objects, lane marking type and / or location and / or lane boundaries, location and / or status of traffic signal devices, lane location, etc.
[0021] Based on the classification result, at least one control action can be initiated and / or performed on the vehicle, including at least one of the following: braking, steering, accelerating, overtaking, emergency braking, activating warning devices, activating warning lights, activating driving direction indicators, light control, etc.
[0022] Classification allows, for example, the recognition of obstacles, whether they are directly in the direction of travel or not, and depending on their location (e.g., the expected vehicle trajectory), appropriate control actions, such as braking or avoidance, can be initiated.
[0023] For example, the brakes may be activated if the classification indicates the presence of an obstacle in the direction of travel and / or a high probability of a collision. It is also conceivable that the roadway and / or roadway boundaries may be recognized based on the classification, and the vehicle may move on the roadway at least partially automatically by control actions.
[0024] The vehicle may be configured as an automobile and / or a passenger car and / or an at least partially autonomous vehicle.
[0025] The method according to the present invention has the advantage that training and / or testing data can be generated with high diversity, especially when representing different environmental scenarios. This can improve the reliability of the training and / or testing and the resulting trained learning system for the classification task. Testing can be performed, for example, by splitting the generated dataset into test and training data and using the test data to check the training progress. In particular, a high diversity of the dataset can improve the generalization ability of the learning system, thus making the classification applicable to new situations, such as vehicle control.
[0026] "Classification" and "image classification" can also include "object detection" or "object-in-image detection," where classification specifically refers to the presence or absence of an object in a particular region of an image. Furthermore, the terms "classification" and "image classification" can also relate to "semantic segmentation," specifically pixel-by-pixel classification.
[0027] It is also optionally envisaged that the image data and / or metadata are obtained from sensor acquisition, preferably image acquisition, of the vehicle's and / or camera's surrounding environment and / or from acquisition by at least one sensor, in particular a camera, preferably on the vehicle, in which acquisition, preferably metadata is defined to describe the environmental scenario. This has the advantage that the method for generating text prompts can take into account a variety of environmental scenarios for image synthesis. Furthermore, since the description is based on accurate data about the environment, the method may be more reliable and accurate. More accurate text descriptions for generating synthetic images improve the accuracy of realistic environmental and object representation in the generated image data.
[0028] Furthermore, the machine learning system may include at least one generative model, preferably a generative model for generating synthetic images, and / or at least one machine learning model for use in at least partially autonomous driving. The synthetic image data generated in the image data can be used, for example, as a training or test data set for the machine learning model. This approach minimizes the effort and cost of data acquisition while increasing the quality and quantity of available data, significantly improving the training of the learning system. A further advantage is the flexibility of the training data to accommodate adaptation needs.
[0029] The step of providing image data can further include providing image data obtained from sensor acquisition and supplemented with metadata, preferably by the driver, during the journey through each environmental scenario. This allows the image data to be supplemented with metadata. This supplementation can optionally be based on data sources such as vehicle cameras or other sensors that record image data during the journey. Data sources not directly used for sensor acquisition can also be used to acquire metadata, such as a cloud server or a GPS system for acquiring weather data. The driver can also manually provide this or additional metadata. Access to these extended data sources allows for a more comprehensive representation of the environmental scenario, further improving the training and testing of the machine learning system. It also improves the accuracy and quality of the text prompts.
[0030] Optionally, the step of generating text prompts comprises converting information contained in the metadata into respective text prompts, thereby obtaining at least one of the following information from the metadata for the dataset to be generated, preferably for the dataset to be generated comprising a plurality of images: - environmental conditions, especially weather or time of day; - context details of each environmental scenario in the surroundings, preferably road type or traffic situation; location information, preferably obtained from GPS acquisition; This has the advantage that the generated text prompts are more detailed and provide more context for the image to be generated.
[0031] Within the scope of the present invention, the step of generating text prompts may further include a transformation of metadata, in which the metadata may be transformed into structured text prompts. In this case, the metadata may be randomly selected for transformation. The use of a probabilistic approach ensures high diversity of the text prompts, which has the advantage of also increasing the diversity of the dataset.
[0032] The generation of the dataset can preferably be performed by a generative machine learning model. Alternatively or additionally, to generate the dataset, the generated text prompts can be supplemented with a conditional spatial layout, thereby considering the application conditions of the machine learning system when generating the dataset. Thus, the combination of the generated text prompts and the conditional spatial layout ensures a realistic representation of the environment. This allows for a more realistic and accurate generation of the dataset. In this case, an edge detection method can preferably be used.
[0033] Furthermore, the subject of the present invention is a method for producing ... - generating a dataset according to the method of the present invention; training a machine learning model with at least the generated dataset; It is a machine learning model obtained by training with
[0034] Thus, the machine learning model according to the present invention has the same advantages as those described in detail with reference to the method according to the present invention.
[0035] The method according to the present invention can be used for manufacturing at least partially autonomous driving systems and / or driver assistance systems in vehicles. The vehicle can be configured, for example, as an automobile and / or a passenger car and / or an autonomously driving vehicle. The vehicle can have, for example, a vehicle device for providing the autonomous driving functions and / or the driver assistance system. The vehicle device can be configured for at least partially automatic control of the vehicle, in particular for acceleration and / or deceleration and / or steering. For this purpose, the control can be based on the output of a machine learning system, preferably a machine learning model according to the present invention.
[0036] The subject of the invention is also a computer program, in particular a computer program product, which comprises instructions that cause a computer to carry out the method according to the invention when the computer program is executed by the computer, and therefore has the same advantages as those explained in detail with reference to the method according to the invention.
[0037] The present invention also provides a data processing apparatus adapted to carry out the method according to the invention. This apparatus can be, for example, a computer that executes the computer program according to the invention. The computer can have at least one processor for executing the computer program. It can also have a non-volatile data storage device in which the computer program is stored and from which the processor can read and execute the computer program.
[0038] Furthermore, the subject of the invention is a computer-readable storage medium which carries the computer program according to the invention and / or which contains instructions which, when executed by a computer, cause the computer to carry out the method according to the invention. This storage medium can be configured as a data storage device, for example a hard disk and / or a non-volatile memory and / or a memory card. The storage medium can, for example, be integrated into the computer.
[0039] Furthermore, the methods of the present invention may be implemented as computer-implemented methods. Alternatively or additionally, at least one disclosed method step may be computer-implemented and / or performed automatically.
[0040] Further advantages, features and details of the invention will become apparent from the following detailed description of an embodiment of the invention, with reference to the accompanying drawings, in which the features mentioned in the claims and in the description can be essential to the invention either individually or in any combination. [Brief explanation of the drawings]
[0041] [Figure 1] 1 is a schematic diagram illustrating a method, an apparatus, a machine learning model, a storage medium, and a computer program according to an embodiment of the present invention. [Figure 2] 1 is a visual diagram illustrating a method according to an embodiment of the present invention. [Figure 3] 4 is a further visual diagram illustrating a method according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0042] FIG. 1 shows a schematic diagram of a method 100, an apparatus 10, a machine learning model 51, a storage medium 15, and a computer program 20 according to an embodiment of the present invention.
[0043] In this case, the illustrated method 100 is used to generate a dataset 60 for training and / or testing a machine learning system 50 .
[0044] In a first method step 101, images representing different environmental scenarios are provided with specific image data.
[0045] In a second method step 102, the descriptions of different environmental scenarios are provided with specific metadata 65. In this case, the metadata 65 may contain specific information about the environment represented by the recorded image data. The metadata 65 may further provide specific information about the conditions under which the recording took place, in particular about the real recording situation captured at the time of recording.
[0046] Then, according to a third method step 103, a text prompt 70 is generated based on the provided metadata 65, and information contained in the metadata 65 can be used in the text prompt 70. If the metadata 65 indicates, for example, that it was snowing or raining at the time of recording, this can be included in the corresponding text prompt 70.
[0047] Furthermore, according to a fourth method step 104, it is possible to generate a data set 60 based on the generated text prompt 70 and preferably on the provided image data.
[0048] 2 shows a visual illustration of an embodiment (including variations) of the present invention, within which the synergistic effects of models such as Stable Diffusion and ControlNet can be utilized to create synthetic images that resemble images captured by vehicle cameras.
[0049] The composite image 80 can be created with the following primary inputs: - a conditional spatial layout, mainly structured by edge detection methods such as Canny (see Figure 3 showing the edge map 75); - a descriptive text prompt to define generic image features, e.g., snowy roads; - optionally a further input 71 such as random noise; It can be created by:
[0050] Text prompts can sometimes be generated automatically through automated annotations such as BLIP (see related citations in the prior art). However, these annotations are usually too short and not detailed enough, resulting in poor image generation quality. Therefore, embodiments of the present invention can contribute to this field by creating improved and more detailed text descriptions. To this end, available metadata can be leveraged for better (more diverse and more accurate) image generation based on descriptive and detailed text prompts.
[0051] One problem with conventional solutions addressed by embodiments of the present invention is that limitations of current text prompting methods often result in the generation of synthetic images that lack variety and uniqueness. Conventional text-image models often generate images that are too general or lack sufficient variety to effectively train advanced AI recognition models. In addition, reliance on manual annotation to create text descriptions is not only inefficient but also unscalable for large datasets.
[0052] Embodiments of the present invention may at least partially address these problems by providing a method 100 for creating text prompts (prompts) that leverages available metadata in a recorded dataset, as exemplarily shown in Figure 1. This approach ensures that the generated text prompts are diverse and unique, thereby enabling the creation of synthetic images that are more representative of real-world scenarios and have sufficient diversity to train robust AI models.
[0053] Furthermore, by leveraging detailed metadata, the prompt creation process can be automated, increasing efficiency and reducing reliance on manual effort or automated methods that do not meet the criteria described above. The problems of the variability and uniqueness limitations of traditional methods and the inefficiency of manual annotation are at least partially addressed by a metadata-based approach to prompt development.
[0054] Embodiments of the present invention further provide a prompt engineering approach that utilizes only metadata associated with an image, i.e., the provided metadata can be the sole data basis for generating text prompts in a method according to embodiments of the present invention.
[0055] The method according to an embodiment of the present invention further includes extracting and converting metadata information (such as time of day, road type, weather conditions, etc.) into structured text prompts, which ensures that the text prompts are not only accurate but also cover a wide range of image features, thus increasing the variety and uniqueness of the generated images.
[0056] I. Metadata Collection According to the first possibility, a metadata-based prompt engineering approach can be provided using a database with comprehensive metadata, which is regularly created and updated with the utmost care for such applications. This is often done by a driver manually entering data collected during travel under various environmental conditions. This metadata includes a variety of information that may be associated with the image, such as environmental conditions, geographic location, time, specific events that occurred during data acquisition, and the context.
[0057] In this case, a method 100 according to an embodiment of the present invention can extract metadata in providing step 101 shown in Figure 1 and use it as the basis for a text prompt. The metadata can then be carefully analyzed to ensure relevance and accuracy in representing the image context. For example, if the metadata indicates that an image was taken on a rainy evening in a suburban area, method 100 can be used to create a text prompt (prompt) that accurately reflects these conditions. This process can include, for example, categorizing and prioritizing metadata elements to create a coherent and detailed description.
[0058] II. Probabilistic Randomization To increase diversity and avoid redundancy in text prompts, probabilistic models can be additionally used. These models follow a random principle to randomly select and combine various metadata elements, ensuring that each prompt is as unique as possible and reflects diverse scenarios. This randomness is essential for generating diverse image descriptions and is advantageous for effective training and fine-tuning of generative models.
[0059] By utilizing metadata collected by drivers, the approach described above not only provides high accuracy and relevance for text prompts, but also contributes to the creation of a comprehensive and diverse dataset that is valuable for improving the capabilities of text-to-image models and allows the generation of more accurate and diverse images that better reflect the complexity of real-world scenarios.
[0060] This metadata-based prompt engineering method offers several important advantages: The accuracy of the text description can be improved by directly using metadata, where the metadata can reflect factual and specific information about the image. - Metadata includes a wide variety of image features, from environmental conditions to context-related details, allowing for a significant increase in the variety of text prompts. - It can help automate and streamline the process of creating detailed and diverse text descriptions for image datasets.
[0061] Possible applications of embodiments of the present invention include improving the training and fine-tuning of text-image models such as Stable Diffusion. Furthermore, by providing a method for generating diverse and detailed text-image pairs, the quality and variety of generated images can be improved, and the method can also contribute to further development of data augmentation techniques and the development of more advanced generative models.
[0062] The above-described embodiments are merely illustrative of the present invention, and it goes without saying that the individual features of the embodiments can be freely combined without departing from the scope of the present invention, as long as they are technically useful.
Claims
1. 1. A method (100) for generating a dataset (60) for training and / or testing a machine learning system (50), the method (100) comprising: - providing (101) specific image data for images in which different environmental scenarios are represented; - providing (102) specific metadata (65) for the descriptions of the different environmental scenarios; - a step (103) of generating a text prompt (70) based on the provided metadata (65), wherein information contained in the metadata (65) is used for the text prompt (70); - generating (104) said data set (60) based on said generated text prompt (70) and preferably on said provided image data; A method comprising:
2. 2. The method (100) of claim 1, wherein the generated dataset (60) is a training dataset comprising a plurality of synthetic image data for training the machine learning system (50), thereby providing representations of the different and / or newly generated environmental scenarios for training.
3. 3. A method (100) according to claim 2, characterized in that the training is intended to train the machine learning system (50) with the generated dataset (60) to classify digital images, in particular digital images obtained from recordings of the surrounding environment of a moving vehicle and / or digital images obtained by a camera, based on the picture elements and / or pixels of the images, and preferably vehicle control is provided based on the classification.
4. 4. The method (100) according to claim 1, wherein the metadata (65) is obtained from a sensor acquisition, preferably an image acquisition, by at least one sensor, in particular at least one camera, preferably on a vehicle, in which acquisition the metadata (65) is defined to describe the environmental scenario.
5. 5. The method (100) according to any one of claims 1 to 4, characterized in that the machine learning system (50) comprises a generative model, preferably a generative model for generating synthetic images and / or a machine learning model for use in at least partially autonomous driving.
6. 6. The method (100) according to any one of claims 1 to 5, wherein the step (101) of providing image data comprises: - providing said image data obtained from sensor acquisition and supplemented by said metadata (65) during a drive in each of said environmental scenarios, preferably by a driver.
7. 7. The method (100) of any one of claims 1 to 6, wherein the step (103) of generating the text prompt (70) comprises: The information contained in the metadata (65) is converted into a respective text prompt, whereby at least one of the following information is extracted from the metadata (65) for the dataset to be generated, preferably a dataset comprising a plurality of images to be generated: - environmental conditions, especially weather or time of day; - context details of each said environmental scenario, preferably road type or traffic situation; location information, preferably obtained from GPS acquisition; and preferably further comprising representing it in an image.
8. 8. The method (100) of claim 1, wherein the step (103) of generating the text prompt (70) further comprises transforming the metadata (65), wherein the transforming converts the metadata (65) into a structured text prompt, and wherein the metadata (65) is randomly selected for transformation.
9. 9. The method (100) of any one of claims 1 to 8, characterized in that for the step of generating (104) the dataset, the generated text prompts (70) are supplemented with a conditional spatial layout, preferably by a generative machine learning model, thereby taking into account application conditions of the machine learning system (50) when generating the dataset.
10. A machine learning model (51) comprising the following steps: - generating a data set (60) by a method (100) according to any one of claims 1 to 9; - training said machine learning model (51) with said at least generated dataset (60); A machine learning model obtained by training with
11. A computer program (20) comprising instructions that, when executed by a computer (10), cause the computer (10) to carry out a method (100) according to any one of claims 1 to 9.
12. An apparatus (10) for processing data, configured to perform a method (100) according to any one of claims 1 to 9.
13. A computer-readable storage medium (15) comprising instructions that, when executed by a computer (10), cause the computer (10) to perform the steps of the method (100) of any one of claims 1 to 9.