Method and apparatus for constructing a dataset of images by feature configurability
By automatically generating datasets for visual perception models using feature-configurable generative artificial intelligence models, the problems of low efficiency and insufficient coverage of manual annotation are solved, thereby improving the accuracy and generalization ability of the models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ROBERT BOSCH GMBH
- Filing Date
- 2024-11-28
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, the construction of training and validation datasets for visual perception models relies on manual annotation, which is inefficient and costly, and it is difficult to cover various complex scenarios, especially edge cases, resulting in insufficient model generalization ability.
By using a feature-configurable approach, generative artificial intelligence models can be used to automatically generate image samples and construct datasets based on feature values, including corresponding ground truth labels, to ensure dataset diversity and relevance, thereby enhancing the training and validation effects of the model.
It improves the accuracy and generalization ability of visual perception models, reduces the inefficiency and high cost of manual annotation, and ensures that the dataset covers various scenarios, including edge cases.
Smart Images

Figure CN122116019A_ABST
Abstract
Description
Technical Field
[0001] This application relates to artificial intelligence, and more specifically, to methods, apparatus, computer-readable storage media, and computer program products for automatically generating images to constitute a dataset based on artificial intelligence models. Background Technology
[0002] In the context of advanced autonomous driving, visual perception models are typically incorporated into applications such as autonomous driving systems, automated parking systems, and driver assistance systems to perceive objects in the vehicle's surrounding environment. For example, visual perception models can identify objects in the vehicle's environment based on images captured by various sensors such as cameras and lidar, thereby assisting in making various driving decisions based on the identified objects during driving or parking.
[0003] To obtain a high-performance visual perception model, it needs to be trained on a training dataset and then validated using a validation dataset. The training and / or validation datasets used have a significant impact on the training efficiency, accuracy, and generalization ability of the visual perception model.
[0004] Visual perception models can be trained using supervised learning methods; therefore, the dataset should include not only the data itself but also corresponding ground truth labels. In current practice, data annotation is typically done manually, requiring significant manpower and failing to guarantee accuracy. Furthermore, occasional edge cases exist in real-world scenarios; training visual perception models based on actual collected images may result in weak generalization ability due to a lack of edge case samples. With the popularization of artificial intelligence technology, various generative AI models offer new image generation methods. Therefore, a method based on AI models to automatically generate images to form a dataset is desired. Summary of the Invention
[0005] The following brief introduction is provided to present some of the selected concepts in a simplified manner, which will be further described in the detailed description that follows. This brief introduction is not intended to highlight the key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.
[0006] In autonomous driving applications, accurately identifying objects in the vehicle's surrounding environment based on images captured by image sensors such as cameras and LiDAR is a crucial task. This task can typically be achieved using a visual perception model mounted on an autonomous driving system. This model can be trained using a training dataset containing a large number of road images, and then validated using a validation dataset containing road images different from the training dataset.
[0007] In practical applications, visual perception may face various complex scenarios, which can be collectively referred to as corner cases in this paper. For example, corner cases can include: visual obstruction due to weather conditions, vehicles or other obstacles appearing in unconventional locations on the road, reflections and ghosting caused by various glare, etc., all of which may cause the visual perception model to fail to accurately identify objects. Therefore, in addition to ordinary driving scenarios, training and validating the visual perception model's ability to handle corner cases is crucial. Furthermore, if the trained visual perception model's ability to handle a certain corner case is insufficient, a specific training dataset can be constructed for retraining.
[0008] To achieve this, a suitable dataset needs to be constructed for the training and / or validation phases. Typically, the dataset used for training and / or validation should include samples corresponding to typical cases and samples corresponding to edge cases. Furthermore, the visual perception model can be trained using supervised learning, therefore the dataset also needs to be correctly labeled. Currently, datasets are generally constructed based on images previously captured by image sensors and labeled manually; however, manual labeling is inefficient and costly. Moreover, actually captured images may not cover all possible road conditions, especially occasional edge cases, which could lead to weak generalization ability of the trained visual perception model.
[0009] Therefore, it is desirable to provide a method for constructing a dataset for a visual perception model using feature-configurable images. This method can automatically generate image samples based on a generative artificial intelligence model according to feature values configured for each feature, and construct a dataset for the visual perception model based on the generated image samples and their corresponding feature values. Furthermore, the method can extract candidate feature values for each feature and set rules for selecting each candidate feature value. Therefore, the method provided in this disclosure can automatically generate sample data with ground truth labels, and by setting selection rules, can construct a dataset that balances diversity and relevance. Using this dataset to train and / or validate a visual perception model can improve its accuracy and generalization ability.
[0010] On one hand, embodiments of this disclosure provide a method for constructing a dataset for a visual perception model using feature-configurable images. The method includes: configuring feature values for one or more features for an image to be generated, wherein the one or more features include features associated with a scene and / or features associated with a target; using the configured feature values of the one or more features as input to a generative artificial intelligence model to generate an image according to the respective feature values, wherein the generative artificial intelligence model is pre-trained; and using the configured feature values of the one or more features as labels for the generated images to construct a dataset for a visual perception model with the labeled images, wherein the dataset is used for training or validation.
[0011] On the other hand, embodiments of this disclosure provide a system comprising: at least one processor; and a memory coupled to the at least one processor having executable instructions stored thereon, which, when executed by the at least one processor, cause the at least one processor to perform the method according to any embodiment of this disclosure.
[0012] On the other hand, embodiments of this disclosure provide a computer-readable medium storing a computer program including instructions that, when executed by a processor, cause one or more units to perform the method described according to any embodiment of this disclosure.
[0013] On the other hand, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, implements the methods described according to any embodiment of this disclosure. Attached Figure Description
[0014] A further understanding of the nature and advantages of this disclosure can be achieved by referring to the accompanying drawings. In the drawings, similar components or features may have the same reference numerals.
[0015] Figure 1 An exemplary architecture diagram is shown, illustrating a dataset for a visual perception model composed of feature-configurable images according to an embodiment of this disclosure.
[0016] Figure 2 An exemplary architecture diagram for extracting candidate feature values according to embodiments of the present disclosure is shown.
[0017] Figure 3 An exemplary flowchart illustrating a dataset for a visual perception model is shown, comprising feature-configurable images according to an embodiment of the present disclosure.
[0018] Figure 4 A block diagram of a system according to an embodiment of the present disclosure is shown. Detailed Implementation
[0019] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed merely to enable those skilled in the art to better understand and implement the subject matter described herein, and are not intended to limit the scope, applicability, or examples set forth in the claims. The function and arrangement of the elements discussed may be changed without departing from the scope of this disclosure. Various processes or components may be omitted, substituted, or added as needed in the various examples. For example, the described methods may be performed in a different order than described, and steps may be added, omitted, or combined. Furthermore, features described in some examples may be combined in other examples.
[0020] As used herein, the term "comprising" and its variations are open terms meaning "including but not limited to". The term "based on" means "at least partially based on". The terms "one embodiment" and "an embodiment" mean "at least one embodiment". The term "another embodiment" means "at least one other embodiment". The terms "first", "second", etc., may refer to different or the same objects. Unless explicitly indicated by the context, the definition of a term remains consistent throughout the specification.
[0021] It is important to note that applying the methods disclosed herein may involve using user-related information, such as road images collected by various vehicle sensors. It is crucial to remember that the use of information involving user privacy requires user authorization, and its use must not exceed the scope of that authorization.
[0022] In autonomous driving scenarios, it is anticipated that objects in the vehicle's surrounding environment can be identified based on image samples captured by the vehicle's sensors to further support driving decisions of the autonomous driving system. In one embodiment, the sensors can be cameras, video cameras, LiDAR, or any suitable type of image sensor. In one embodiment, objects in the vehicle's surrounding environment can be other vehicles, non-motorized vehicles (such as bicycles, tricycles), pedestrians, vegetation, parking locks, walls, railings, curbs, lane lines, traffic signs, etc. In one embodiment, object recognition can be achieved through a visual perception model. For example, a visual perception model is first trained using a training dataset containing a large number of image samples based on supervised learning, and then validated using a validation dataset containing image samples different from the training dataset to determine if the trained visual perception model possesses ideal performance.
[0023] In practical applications, object recognition may encounter various complex scenarios, which can be collectively referred to as corner cases in this paper. For example, corner cases can include visibility obstruction due to various weather conditions (such as fog, rain, or snow); obstacles appearing in unconventional locations on the road, such as vehicles parked across the road; reflections and ghosting caused by various reflective objects (such as watered surfaces or mirrored car bodies); and object overlap caused by vehicles carrying other objects (such as bicycles carried on the rear of vehicles). These corner cases can all lead to misidentification of objects. Therefore, a high-quality dataset needs to be constructed for the training and / or validation phases. This dataset should include samples corresponding to various road conditions, including both common and corner cases.
[0024] This disclosure proposes a method for constructing a dataset for a visual perception model using feature-configurable images. This method allows configuring feature values for each feature, automatically generating image samples based on these feature values using a generative artificial intelligence model, and then constructing the dataset for the visual perception model based on the generated image samples and their corresponding feature values. This approach avoids the inefficient and costly manual annotation process, automatically generating samples with ground truth labels. Furthermore, candidate feature values can be extracted for each feature, and rules can be set for selecting among these candidate feature values to generate image samples on demand. This method allows control over the sample distribution of the constructed dataset, ensuring that the dataset includes samples of the desired type.
[0025] On one hand, to prevent the model under training from overfitting and to enhance its generalization ability, samples corresponding to various edge cases can be added to the training set. On the other hand, to enhance the model's ability to handle a certain edge case, samples corresponding to that edge case can be added to the training set. In this way, the accuracy and generalization ability of the trained visual perception model can be improved.
[0026] Figure 1 An exemplary architecture diagram is shown illustrating a dataset for a visual perception model composed of feature-configurable images, according to embodiments of this disclosure. It should be understood that the following examples are provided merely to aid in a better understanding of this disclosure and do not constitute any limitation on the scope of this disclosure.
[0027] like Figure 1As shown, feature values for one or more features 101-1 to 101-N can be configured for the image to be generated. In one embodiment, one or more features 101-1 to 101-N may include features associated with the image scene and / or features associated with targets in the image. For example, features associated with the image scene may include: scene time (e.g., daytime, nighttime), scene weather (e.g., sunny, cloudy, overcast, rainy, snowy, foggy), reflectivity (e.g., strong reflectivity, weak reflectivity), occlusion (e.g., strong occlusion, weak occlusion, no occlusion), etc. For example, features associated with targets may include: target type (e.g., motor vehicles, non-motor vehicles, pedestrians, plants, parking locks, walls, railings, curbs, lane lines, traffic signs, etc.), number of targets, target appearance (e.g., color), target pose (e.g., position, orientation), target size, etc.
[0028] In one embodiment, feature values of one or more features 101-1 to 101-N can be configured via prompts. For example, a user can configure the scene weather feature value to "sunny" and the target appearance feature value to "red" using prompts such as "scene weather is sunny" or "target appearance is red." Alternatively, the user can configure feature values of one or more features 101-1 to 101-N directly in an editable dialog box via software. For example, a user can directly edit feature values such as the number of targets, target pose, and target size using plugins such as ControlNet.
[0029] Alternatively or concurrently, feature values for one or more features 101-1 to 101-N can be automatically configured by a computer program, for example, by automatically selecting from a plurality of candidate feature values according to pre-set rules. In one embodiment, candidate feature values can be extracted for one or more features based on pre-acquired road images, and then the feature values to be configured for one or more features 101-1 to 101-N can be automatically selected from the extracted candidate feature values.
[0030] Figure 2 An exemplary architecture diagram for extracting candidate feature values according to embodiments of this disclosure is shown. It should be understood that the following examples are provided merely to aid in a better understanding of this disclosure and do not constitute any limitation on the scope of this disclosure.
[0031] like Figure 2 As shown, a pre-collected road image 202 can be obtained from the database 201. The feature extraction module 203 extracts feature values from the road image 202 based on one or more features 204-1 to 204-N. These features 204-1 to 204-N can correspond to... Figure 1One or more features 101-1 to 101-N are extracted. The extracted feature values can be used as candidate feature values, wherein the candidate feature values can be extracted and represented as text, numbers, feature vectors, or any form that a computer can process. It should be understood that the feature extraction module 203 can be implemented using any method applicable in the art.
[0032] In one embodiment, pre-collected road images can be first classified according to road conditions, such as into road image subsets corresponding to various normal conditions and road image subsets corresponding to various edge conditions. Then, feature values can be extracted based on multiple road image subsets. For example, feature values can be extracted for features 204-1 to 204-N for the road image subset showing "occlusion caused by extreme weather," obtaining multiple feature values for representative features from features 204-1 to 204-N. For example, "rainy day, snowy day, heavy fog" can be extracted as candidate feature values for scene weather features, and "strong occlusion" can be extracted as a candidate feature value for occlusion degree. Therefore, when generating image samples for road conditions such as "occlusion caused by extreme weather," the candidate feature values extracted for scene weather features and occlusion degree features can be automatically selected. Additionally, candidate feature values can be pre-set for other features, such as setting candidate feature values for target type ("motorized vehicles, non-motorized vehicles, pedestrians") and candidate feature values for target appearance ("white, gray," etc.).
[0033] based on Figure 2 In this embodiment, when configuring feature values for one or more features 101-1 to 101-N, rules for selecting from candidate feature values can be set. For example, rules can be set to iterate through combinations of candidate feature values. As another example, when a diverse and evenly distributed dataset is desired, selection can be set with equal probability among candidate feature values. For instance, if there are 5 candidate feature values for the scene weather feature, the probability of selecting each candidate feature value can be set to 20%, and so on. Furthermore, conditional probabilities for selecting candidate feature values sequentially can be set based on the correlation between candidate feature values of different features. For example, if the candidate feature value "sunny" has already been selected for the scene weather feature, the candidate feature value "strong reflection" can be selected for the reflection level feature and the candidate feature value "no occlusion" can be selected for the occlusion level feature with a higher probability, and so on.
[0034] Continue as Figure 1As shown, one or more corresponding feature embedding vectors 102-1 to 102-N can be obtained based on the feature values configured for one or more features 101-1 to 101-N. In one embodiment, the feature embedding vectors can be obtained based on the feature values using a CLIP model. It should be understood that any method applicable in the art can be used to obtain the feature embedding vectors.
[0035] like Figure 1 As shown, one or more feature embedding vectors 102-1 to 102-N can be input into a generative artificial intelligence model 103 to generate an image 104. In one embodiment, one or more feature embedding vectors 102-1 to 102-N can be concatenated and then input into the generative artificial intelligence model 103. In one embodiment, the generative artificial intelligence model 103 can be pre-trained, for example, based on multiple pre-acquired road images. In one embodiment, the generative artificial intelligence model 103 can be a diffusion model, such as a Stable-Diffusion model.
[0036] like Figure 1 As shown, a dataset 105 can be constructed by forming data-label pairs based on the generated image 104 and the feature values of one or more corresponding features 101-1 to 101-N. In one embodiment, feature values can be extracted from prompts entered by the user using a computer program and stored as truth labels for the corresponding features. Alternatively, feature values can be extracted from information entered by the user in an editable dialog box using a computer program and stored as truth labels for the corresponding features. Alternatively, when feature values of one or more features are automatically configured by a computer program, the configured feature values can be stored as truth labels for the corresponding features. Further, the generated image and the truth labels of the corresponding multiple features can be associated and stored in a structured manner, such as in various types of databases, to form data-label pairs.
[0037] Using the content of this disclosure in combination Figure 1 and Figure 2 The described method avoids the inefficient and costly manual annotation process, automatically generating samples with ground truth labels to construct the dataset. Furthermore, image samples can be generated on demand to control the sample distribution of the constructed dataset, thereby ensuring that the dataset includes samples of the desired type, especially image samples corresponding to edge cases.
[0038] Figure 3 An exemplary flowchart illustrating how feature-configurable images constitute a dataset for a visual perception model according to an embodiment of this disclosure is shown.
[0039] At step 310, feature values for one or more features can be configured for the image to be generated, wherein the one or more features include features associated with the scene and / or features associated with the target.
[0040] In one embodiment, the scene-associated features may include one or more of the following features: scene time, scene weather, reflectivity, and occlusion. For example, scene time may include morning, noon, evening, night, etc.; scene weather may include sunny, cloudy, overcast, rainy, snowy, foggy, etc.; reflectivity may include strong reflectivity, weak reflectivity, etc.; and occlusion may include strong occlusion, weak occlusion, no occlusion, etc.
[0041] In one embodiment, the features associated with the target may include one or more of the following features: target type, target quantity, target appearance, target pose, and target size. For example, target type may include motor vehicles, non-motor vehicles, pedestrians, plants, parking locks, walls, railings, curbs, lane lines, traffic signs, etc.; target appearance may include color, pattern, etc.; target pose may include position, orientation, etc.; and target size may include the length, width, and height of the bounding box, etc.
[0042] In one embodiment, configuring feature values for one or more features for the image to be generated may include configuring feature values for one or more features by receiving user input. For example, feature values for one or more features may be configured based on prompts input by the user, such as "scene weather is sunny" or "target appearance is red," to configure the feature value for the scene weather as "sunny" and the feature value for the target appearance as "red."
[0043] Alternatively, feature values for one or more features can be configured based on values edited by the user in a dialog box. For example, feature values for one or more features can be configured directly in an editable dialog box via software. For example, feature values such as the number of targets, target pose, and target size can be directly edited via plugins such as ControlNet.
[0044] In one embodiment, configuring feature values for the one or more features in the image to be generated can include automatically configuring the feature values for the one or more features via a computer program. For example, the computer program can iterate through all combinations of candidate feature values. For example, it can automatically select from multiple candidate feature values according to a pre-set probability.
[0045] In one embodiment, multiple candidate feature values can be extracted based on at least one of the one or more features from pre-acquired images. For example, pre-acquired road images can be first classified into multiple road image subsets according to road conditions, and then feature values can be extracted based on each road image subset.
[0046] In one embodiment, the feature values for one or more features may be configured for an image corresponding to an edge case.
[0047] At step 320, the feature values of one or more configured features can be used as input to a generative artificial intelligence model to generate an image according to the respective feature values, wherein the generative artificial intelligence model is pre-trained.
[0048] In one embodiment, one or more feature embedding vectors can be obtained based on the feature values of the one or more features, and the concatenation of the one or more feature embedding vectors can be input into the generative artificial intelligence model.
[0049] In one embodiment, the feature embedding vector can be obtained based on the eigenvalues using the CLIP model. It should be understood that any method applicable in the art can be used to obtain the feature embedding vector.
[0050] In one embodiment, the generative artificial intelligence model is a diffusion model, such as the Stable-Diffusion model.
[0051] At step 330, the feature values of one or more configured features can be used as labels for the generated images to construct a dataset for a visual perception model, wherein the dataset is used for training or validation.
[0052] Figure 4 A block diagram of a system according to an embodiment of the present disclosure is shown.
[0053] System 400 may include one or more processors 410 and memory 420. Memory 420 may store executable instructions. Processor 410 may execute executable instructions stored or encoded in memory 420, thereby achieving the above-described combination. Figures 1 to 3 The various operations and / or functions described. Although not in Figure 4As shown, but those skilled in the art will understand, system 400 may include various other components, such as various communication modules, bus modules, and possibly user interface modules. In one embodiment, system 400 may include an input module that can be configured to receive user input, such as prompts in conjunction with user input in the embodiments described above, values edited by the user in a dialog box, and a set probability of automatically selecting from multiple candidate feature values.
[0054] Embodiments of this disclosure also provide a computer-readable storage medium. The computer-readable storage medium may store executable instructions, which, when executed by a processor, can achieve the above-described combinations. Figures 1 to 3 The various operations and / or functions described. For example, computer-readable storage media may include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), static random access memory (SRAM), hard disk, flash memory, etc.
[0055] Embodiments of this disclosure also provide a computer program product. The computer program product may include a computer program. When executed by a processor, the computer program can achieve the above-described combinations. Figures 1 to 3 The various operations and / or functions described.
[0056] Specific embodiments of this disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0057] Not all steps and units in the above process and system structure diagrams are necessary; some steps or units can be omitted according to actual needs. The device structure described in the above embodiments can be a physical structure or a logical structure. That is, some units may be implemented by the same physical entity, some units may be implemented by multiple physical entities, or they may be jointly implemented by certain components in multiple independent devices.
[0058] The term "exemplary" as used throughout this specification means "serving as an example, instance, or illustration" and does not imply "preferred" or "advantageous" compared to other embodiments. Detailed descriptions are included for the purpose of providing an understanding of the described techniques. However, these techniques can be practiced without these detailed descriptions. In some instances, well-known structures and apparatuses are shown in block diagram form to avoid obscuring the concepts of the described embodiments.
[0059] The foregoing description of this disclosure is provided to enable any person skilled in the art to implement or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein can be applied to other variations without departing from the scope of this disclosure. Therefore, this disclosure is not limited to the examples and designs described herein, but is consistent with the widest scope of the principles and novel features disclosed herein.
Claims
1. A method for constructing a dataset for a visual perception model using feature-configurable images, comprising: Configure feature values for one or more features for the image to be generated, wherein the one or more features include features associated with the scene and / or features associated with the target; The feature values of one or more configured features are used as input to a generative artificial intelligence model to generate an image according to the respective feature values, wherein the generative artificial intelligence model is pre-trained; and The feature values of one or more configured features are used as labels for the generated images, and the images with the labels are used to construct a dataset for a visual perception model, wherein the dataset is used for training or validation.
2. The method according to claim 1, wherein, The scene-related features may include one or more of the following: scene time, scene weather, reflectivity, and occlusion.
3. The method according to claim 1, wherein, The features associated with the target may include one or more of the following: target type, target quantity, target appearance, target pose, and target size.
4. The method according to claim 1, wherein the feature values of one or more configured features are used as input to the generative artificial intelligence model, comprising: Based on the feature values of one or more of the aforementioned features, one or more corresponding feature embedding vectors are obtained; The concatenation of one or more feature embedding vectors is input into the generative artificial intelligence model.
5. The method according to claim 1, wherein, The generative artificial intelligence model is a diffusion model.
6. The method according to claim 1, wherein, Configure feature values for one or more features for the image to be generated, including: Configure feature values for one or more features by receiving user input; and / or The feature values for the one or more features are automatically configured by a computer program.
7. The method according to claim 6, wherein, The step of configuring feature values for one or more features by receiving user input includes: Configure feature values for one or more features based on prompts entered by the user; and / or The feature values for the one or more features are configured based on the values edited by the user in the dialog box.
8. The method according to claim 6, wherein, The automatic configuration of feature values for the one or more features via a computer program includes: For at least one of the aforementioned features, a feature value is automatically selected from multiple candidate feature values according to a pre-set probability.
9. The method according to claim 8, wherein, Multiple candidate feature values for at least one of the aforementioned features are extracted based on pre-acquired images.
10. The method according to claim 1, wherein, The feature values used for one or more features are configured for an image corresponding to an edge case.
11. A system comprising: At least one processor; A memory coupled to the at least one processor, thereon storing executable instructions that, when executed by the at least one processor, cause the at least one processor to implement the method according to any one of claims 1 to 10.
12. The system according to claim 11, further comprising: An input module is configured to receive user input, wherein the user input includes one or more of the following: prompt words, values edited in a dialog box, or probabilities of automatically selecting from multiple candidate feature values.
13. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 10.