Image Sample Generation Method and Visualization Analysis System for Autonomous Driving Scenarios

By performing mask extraction, in-depth analysis and weather recognition model construction on the autonomous driving image dataset, target image samples are generated, which solves the problem of insufficient diversity of the existing dataset and improves the adaptability and safety of the autonomous driving system.

CN120107921BActive Publication Date: 2025-07-22SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510585818.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-22
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing autonomous driving image datasets have insufficient diversity, especially the lack of extreme weather conditions and abnormal objects in scenarios, resulting in inaccurate predictions of models when dealing with rare or complex scenarios, and poses safety risks.

Method used

By obtaining the data set of autonomous driving pictures, mask extraction and data cleaning, in-depth analysis, constructing weather recognition models and situation-aware spatial representation models, generating target picture samples, and introducing extreme and rare situations through a controllable generation module, combining a visual analysis system for multi-dimensional display and evaluation.

Benefits of technology

It realizes more targeted, richer and more comprehensive image sample generation, improves the quality of image samples, and enhances the adaptability and safety of the autonomous driving system to complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107921B_ABST
    Figure CN120107921B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating picture samples and a visualization analysis system for an autonomous driving scenario. The method includes: obtaining an autonomous driving picture data set; performing mask extraction and data cleaning operations on the autonomous driving picture data set to obtain a target mask graph; performing in-depth analysis on the scenarios of the autonomous driving picture data set to obtain a depth graph; obtaining target simulated weather pictures, and constructing a target weather recognition model according to the target simulated weather pictures; constructing a target situation awareness spatial representation model according to the target mask graph and the depth graph; and obtaining target picture samples according to the target weather recognition model and the target situation awareness spatial representation model. The present invention can improve the quality of picture samples and can be widely applied to the technical field of image data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and in particular, to a method for generating picture samples and a visualization analysis system for an autonomous driving scenario. Background Art

[0002] Autonomous driving technology has injected new vitality into the automotive industry. The performance and reliability of current autonomous driving technology largely depend on high-quality training data sets. Most systems mainly rely on data collected from daily driving scenarios for training and evaluation. However, existing image data sets have significant deficiencies in terms of diversity. These data sets often do not include scenarios that are less common in the real world, such as extreme weather conditions or the appearance of abnormal objects. When dealing with rare or complex scenarios, the model performance is largely limited by the distribution of long-tail data (corner cases). When these situations occur, they may lead to inaccurate predictions of the models integrated into the autonomous driving system, thus bringing immeasurable risks. Summary of the Invention

[0003] In view of this, the main purpose of the embodiments of the present invention is to provide a method for generating picture samples and a visualization analysis system for an autonomous driving scenario, in order to solve at least one of the problems in the prior art and improve the quality of the picture samples.

[0004] To achieve the above object, on the one hand, an embodiment of the present invention provides a method for generating picture samples for an autonomous driving scenario, the method comprising the following steps:

[0005] Obtain an autonomous driving picture data set;

[0006] Perform mask extraction and data cleaning operations on the autonomous driving picture data set to obtain a target mask image;

[0007] Perform in-depth analysis on the scenarios of the autonomous driving picture data set to obtain a depth map;

[0008] Obtain a target simulated weather picture, and construct a target weather recognition model according to the target simulated weather picture;

[0009] Construct a target situation awareness spatial representation model according to the target mask image and the depth map;

[0010] Obtain a target picture sample according to the target weather recognition model and the target situation awareness spatial representation model.

[0011] In some embodiments, the method for generating picture samples for an autonomous driving scenario further comprises the following steps:

[0012] Display the first quantity of the first target object in the autonomous driving picture dataset and the second quantity of the second target object in the target picture sample, and generate a stacked bar chart;

[0013] Display the first distribution of the first target object in the autonomous driving picture dataset and the second distribution of the second target object in the target picture sample, and generate a heat map of the object position distribution;

[0014] Display the first multi-dimensional data of the autonomous driving picture dataset, and / or display the second multi-dimensional data of the target picture sample, and generate a multi-dimensional display diagram; wherein, the first multi-dimensional data includes the first quantity, the first depth value of the first target object, the first occlusion degree of the first target object, and the first weather condition; the second multi-dimensional data includes the second quantity, the second depth value of the second target object, the second occlusion degree of the second target object, and the second weather condition;

[0015] In response to a target dimension selection operation, display the third distribution of the autonomous driving picture dataset and the target picture sample in the target dimension, and generate a heat map of the target dimension distribution.

[0016] In some embodiments, the obtaining of the target simulated weather picture includes the following steps:

[0017] Construct a controllable weather generation model according to a cyclic generative adversarial network;

[0018] Generate a first simulated weather picture through the controllable weather generation model;

[0019] Preprocess the first simulated weather picture to obtain a second simulated weather picture;

[0020] Perform a weather type annotation operation on the second simulated weather picture to obtain the target simulated weather picture.

[0021] In some embodiments, the constructing of the target weather recognition model according to the target simulated weather picture includes the following steps:

[0022] Construct an initial weather recognition model;

[0023] Construct a first loss function;

[0024] Input the target simulated weather picture into the initial weather recognition model, and train the initial weather recognition model through a residual connection network and the first loss function to obtain the target weather recognition model.

[0025] In some embodiments, constructing a target situation awareness spatial representation model according to the target mask graph and the depth graph includes the following steps:

[0026] Fuse the semantic mask of the target mask graph and the depth graph to obtain a semantic depth fusion graph;

[0027] Construct an initial situation awareness spatial representation model based on a conditional variational autoencoder;

[0028] Construct a second loss function;

[0029] Input the semantic depth fusion graph into the initial situation awareness spatial representation model, and train the initial situation awareness spatial representation model through the second loss function to obtain the target situation awareness spatial representation model.

[0030] In some embodiments, obtaining a target picture sample according to the target weather recognition model and the target situation awareness spatial representation model includes the following steps:

[0031] Classify and identify the autonomous driving picture dataset according to the target weather recognition model to obtain the corresponding weather type;

[0032] Select a preset candidate position box through the autonomous driving picture dataset;

[0033] Obtain the target category of the first target object in the autonomous driving picture dataset, and obtain the depth information of the first target object;

[0034] Generate test condition information according to the preset candidate position box, the target category, and the depth information;

[0035] Input the test condition information and random Gaussian noise into the target situation awareness spatial representation model to generate a target candidate position box;

[0036] Generate the target picture sample according to the autonomous driving picture dataset, the weather type, and the target candidate position box.

[0037] In some embodiments, inputting the test condition information and random Gaussian noise into the target situation awareness spatial representation model to generate a target candidate position box includes the following steps:

[0038] Input the test condition information and the random Gaussian noise into the target situation awareness spatial representation model to generate an initial candidate position box;

[0039] Preset target constraint conditions; wherein, the target constraint conditions include an aspect ratio constraint condition, a ground intersection constraint condition, and an occlusion constraint condition;

[0040] According to the target constraint conditions, the initial candidate position box is optimized through linear approximation constraints to obtain the target candidate position box.

[0041] To achieve the above object, on the other hand, an embodiment of the present invention proposes a visual analysis system for generating picture samples for an autonomous driving scenario, and the visual analysis system includes:

[0042] A first module, configured to obtain an autonomous driving picture data set;

[0043] A second module, configured to perform mask extraction and data cleaning operations on the autonomous driving picture data set to obtain a target mask map;

[0044] A third module, configured to perform in-depth analysis on the scenario of the autonomous driving picture data set to obtain a depth map;

[0045] A fourth module, configured to obtain a target simulated weather picture, and construct a target weather recognition model according to the target simulated weather picture;

[0046] A fifth module, configured to construct a target situation awareness spatial representation model according to the target mask map and the depth map;

[0047] A sixth module, configured to obtain a target picture sample according to the target weather recognition model and the target situation awareness spatial representation model.

[0048] In some embodiments, the visual analysis system for generating picture samples for an autonomous driving scenario further includes:

[0049] A seventh module, configured to display a first quantity of a first target object in the autonomous driving picture data set and a second quantity of a second target object in the target picture sample, and generate a stacked bar chart;

[0050] An eighth module, configured to display a first distribution of a first target object in the autonomous driving picture data set and a second distribution of a second target object in the target picture sample, and generate an object position distribution heat map;

[0051] A ninth module, configured to display first multi-dimensional data of the autonomous driving picture data set and / or display second multi-dimensional data of the target picture sample, and generate a multi-dimensional display map; wherein, the first multi-dimensional data includes the first quantity, a first depth value of the first target object, a first occlusion degree of the first target object, and a first weather condition; the second multi-dimensional data includes the second quantity, a second depth value of the second target object, a second occlusion degree of the second target object, and a second weather condition;

[0052] The tenth module is configured to display the automatic driving picture dataset and the third distribution of the target picture sample in the target dimension in response to a target dimension selection operation, and generate a target dimension distribution heat map.

[0053] To achieve the above object, another aspect of the embodiments of the present invention provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned method for generating picture samples for an automatic driving scenario.

[0054] To achieve the above object, another aspect of the embodiments of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned method for generating picture samples for an automatic driving scenario.

[0055] To achieve the above object, another aspect of the embodiments of the present invention provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the computer device executes the above-mentioned method for generating picture samples for an automatic driving scenario.

[0056] The embodiments of the present invention at least include the following beneficial effects: The present invention provides a method for generating picture samples for an automatic driving scenario and a visualization analysis system. This solution obtains an automatic driving picture dataset; performs mask extraction and data cleaning operations on the automatic driving picture dataset to obtain a target mask map; deeply analyzes the scenarios of the automatic driving picture dataset to obtain a depth map; obtains a target simulated weather picture, and constructs a target weather recognition model according to the target simulated weather picture; constructs a target situation awareness spatial representation model according to the target mask map and the depth map; and obtains a target picture sample according to the target weather recognition model and the target situation awareness spatial representation model, which can realize the generation of more targeted, richer and more comprehensive picture samples and improve the quality of the picture samples. Description of the Drawings

[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0058] Figure 1 It is a flowchart of a method for generating picture samples for an autonomous driving scenario provided by an embodiment of the present invention;

[0059] Figure 2 It is a flowchart of another method for generating picture samples for an autonomous driving scenario provided by an embodiment of the present invention;

[0060] Figure 3 It is an overall flowchart of generating picture samples for an autonomous driving scenario provided by an embodiment of the present invention;

[0061] Figure 4 It is a flowchart of data feature extraction provided by an embodiment of the present invention;

[0062] Figure 5 It is a flowchart of controllable picture generation provided by an embodiment of the present invention;

[0063] Figure 6 It is a schematic diagram of the controllable generation effect of the global environment provided by an embodiment of the present invention;

[0064] Figure 7 It is a flowchart of local controllable generation provided by an embodiment of the present invention;

[0065] Figure 8 It is a schematic diagram of the generation effect of picture samples provided by an embodiment of the present invention;

[0066] Figures 9a - 9d It is a schematic diagram of visualizing multi-dimensional information of pictures provided by an embodiment of the present invention;

[0067] Figure 10 It is a schematic diagram of a visual analysis system for generating picture samples for an autonomous driving scenario provided by an embodiment of the present invention;

[0068] Figure 11 It is a schematic diagram of the controllable picture sample generation effect of a visual analysis system provided by an embodiment of the present invention;

[0069] Figure 12 It is a schematic diagram of the controllable picture sample generation effect of another visual analysis system provided by an embodiment of the present invention;

[0070] Figure 13 It is a schematic diagram of the comparative evaluation of the picture sample generation effect provided by an embodiment of the present invention;

[0071] Figure 14 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0072] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present invention. They are only examples of devices and methods consistent with some aspects of the embodiments of the present invention as detailed in the appended claims.

[0073] It should be noted that although the functional modules are divided in the system schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the system or the order in the flowchart. The terms "first / S100" and "second / S200" in the specification, claims and the above drawings may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "when" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0074] The terms "at least one", "a plurality", "each", "any one" and the like used in the present invention, at least one includes one, two or more, a plurality includes two or more, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.

[0075] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.

[0076] Before describing the embodiments of the present invention in detail, some nouns and terms involved in the embodiments of the present invention will be described first. The nouns and terms involved in the embodiments of the present invention are applicable to the following explanations.

[0077] Deep learning: Deep learning is a branch of machine learning. Based on the development of artificial neural networks, it simulates the thinking mode of the human brain through a multi-layer network structure. It generates more abstract high-level representations (such as objects or scenes) by extracting and combining low-level features (such as edges and colors) layer by layer. Deep learning models optimize parameters through the backpropagation algorithm, enabling the network to learn from large-scale datasets and automatically extract patterns in the data. Therefore, it is widely used in fields such as image recognition, speech recognition, natural language processing, medical image analysis, and autonomous driving.

[0078] Diffusion Model: A diffusion model is a type of probabilistic generative model that mainly generates new samples by simulating the process of noise gradually spreading and reverse denoising. This model starts from completely random noise and generates data consistent with the target distribution through denoising in multiple iterations, such as high-quality images or speech samples. The generation process of the diffusion model is similar to that of a Markov chain, so the computational cost is relatively large. However, the quality of the samples it generates is usually higher than that of traditional generative adversarial networks (GANs) and variational autoencoders (VAEs), especially showing remarkable performance in image generation and natural language tasks.

[0079] Controllable Data Generation: Controllable data generation refers to generating sample data that meets the requirements according to user-defined conditions or constraints. It allows users to control the style, content, or structure of the generated data by inputting parameters to meet the needs of specific scenarios.

[0080] Visual Analytics System: A visual analytics system is an interactive system that combines data visualization and analysis functions, aiming to help users quickly and intuitively understand the patterns and anomalies in complex data. By presenting data through visual means such as charts and dashboards, users can quickly capture the trends behind the data and deeply explore their internal relationships.

[0081] Human-Computer Interaction: Human-Computer Interaction (HCI) is an interdisciplinary field dedicated to studying the ways and underlying mechanisms of interaction between humans and computer systems. Its goal is to balance usability and functionality in design and improve the user experience by optimizing interface design and interaction patterns. Human-computer interaction involves multidisciplinary knowledge, such as psychology, cognitive science, computer science, and design science, and covers various scenarios in applications from desktop software interfaces to virtual reality, augmented reality, and smart home devices.

[0082] In today's digital age, the rise of autonomous driving technology has not only injected new vitality into the automotive industry but also led to a revolutionary change in future travel modes, aiming to provide people with a safer, more convenient, and efficient transportation experience. The performance and reliability of autonomous driving technology largely depend on high-quality training data sets.

[0083] Currently, most systems mainly rely on data collected from daily driving scenarios for training and evaluation. However, existing image datasets have significant deficiencies in terms of diversity, which affects the adaptability and robustness of autonomous driving systems. These datasets often do not include less common scenarios in the real world, such as extreme weather conditions or the appearance of abnormal objects. When dealing with rare or complex scenarios, this limitation can pose significant challenges to autonomous driving systems. In other words, the further improvement of model performance is largely restricted by the distribution of long-tail data (corner cases). When these situations occur, they may lead to inaccurate predictions of the models integrated into autonomous driving systems, thus bringing immeasurable risks.

[0084] Although some studies have specifically explored autonomous driving datasets under extreme weather conditions, these datasets are still difficult to maintain balance in certain dimensions. Therefore, to improve the safety and reliability of autonomous driving vehicles, it is crucial to ensure that the dataset covers all kinds of scenarios that drivers may encounter in the real world. Expanding the coverage of special scenarios usually needs to be achieved in two ways: one is to collect more data from the real world, and the other is to generate data samples of specific scenarios. For the former, collecting autonomous driving datasets usually relies on road tests. However, road tests are costly and difficult to cover diverse driving scenarios, such as different weather conditions and complex road conditions. For the latter, some studies use driving simulation software to generate virtual image samples to simulate real-world driving scenarios. Although these simulated data can supplement existing datasets to a certain extent, the fidelity of the simulation and the authenticity of the virtual scenarios are still controversial. This gap between simulation and reality reveals the limitations of relying solely on virtual environments for comprehensive dataset enhancement.

[0085] In view of this, as Figure 3 shown, through an end-to-end autonomous driving perception data generation system, it is possible to utilize the rich real-world scene information in the real autonomous driving picture dataset, perform data feature extraction (as Figure 4 shown) and other processing on the real autonomous driving picture dataset, and also introduce various extreme and rare situations through the global control and local control of the picture controllable generation module (as Figure 5 shown) to further improve the diversity and reliability of the dataset, thereby generating picture samples specifically, enabling the autonomous driving picture dataset to effectively enhance data in the direction of customization, and further enhancing the autonomous driving system's understanding and adaptability to the complexity and unpredictability of the real world. In addition, multi-dimensional visual analysis can also be carried out on the autonomous driving picture dataset and the generated picture samples, and by comparing the picture results generated by different controllable parameters through human-computer interaction and data-driven methods, the picture dataset can be accurately explored and evaluated.

[0086] AsFigure 1 As shown, an embodiment of the present invention provides a method for generating picture samples for an autonomous driving scenario, and this method may include but is not limited to steps S100 to S600:

[0087] Step S100, obtain an autonomous driving picture dataset;

[0088] Step S200, perform mask extraction and data cleaning operations on the autonomous driving picture dataset to obtain a target mask image;

[0089] Step S300, perform in-depth analysis on the scenarios of the autonomous driving picture dataset to obtain a depth map;

[0090] Step S400, obtain target simulated weather pictures, and construct a target weather recognition model according to the target simulated weather pictures;

[0091] Step S500, construct a target situation awareness spatial representation model according to the target mask image and the depth map;

[0092] Step S600, obtain target picture samples according to the target weather recognition model and the target situation awareness spatial representation model.

[0093] In step S100 of some embodiments, obtain an autonomous driving picture dataset , and this autonomous driving picture dataset includes pictures covering different weather conditions and various target object categories. Among them, different weather conditions may be sunny, rainy, foggy, strong light, etc., but are not limited thereto; target object categories may be vehicles, pedestrians, bicycles, motorcycles, traffic signs, etc., but are not limited thereto. Through the obtained autonomous driving picture dataset, richer real-scene information can be further obtained.

[0094] In step S200 of some embodiments, as Figure 4 shown, in data feature extraction, perform mask extraction on the obtained autonomous driving picture dataset to extract the target object masks in the pictures and obtain an initial mask image. Through these masks, each target object can be separately segmented, laying a foundation for subsequent depth estimation, prediction of the occlusion degree of target objects, etc. Further, perform data cleaning on the initial mask image, including removing pictures with unclear labels (which may include but are not limited to the shapes and object categories of the masks) and large noise, and deleting target objects that cannot be recognized or are too small in size, so as to obtain a target mask image. By performing mask extraction and data cleaning operations on the autonomous driving picture dataset, the target mask image can reflect the distribution and characteristics of target objects in the real environment.

[0095] In step S300 of some embodiments, asFigure 4 As shown, in data feature extraction, through a depth estimation model, the scene in the pictures of the autonomous driving picture dataset is analyzed for depth, generating a depth map for each pixel point in the picture, and thus the distance relationship and spatial position of the target objects in the scene can be inferred. In the subsequent steps, by combining the depth map with the mask information of the target objects (target mask map), the specific depth of each target object can be obtained, providing a basis for the subsequent object insertion algorithm.

[0096] In step S400 of some embodiments, as Figure 4 shown, in data feature extraction, by simulating complex weather conditions in the autonomous driving scene (such as rain, fog, sunlight of different intensities, etc., not limited thereto), a batch of simulated weather pictures with different weather conditions are generated, and these generated simulated weather pictures will be used to train a weather recognition model, so as to obtain a trained weather recognition model. This trained weather recognition model can accurately identify and classify the weather types of the pictures in the autonomous driving dataset, facilitating subsequent users to analyze the distribution of pictures with different weather conditions in the autonomous driving picture dataset.

[0097] In some embodiments, the step of obtaining the target simulated weather pictures may include, but is not limited to, steps S411 to S414:

[0098] Step S411, constructing a controllable weather generation model according to the CycleGAN;

[0099] Step S412, generating a first simulated weather picture through the controllable weather generation model;

[0100] Step S413, preprocessing the first simulated weather picture to obtain a second simulated weather picture;

[0101] Step S414, performing a weather type annotation operation on the second simulated weather picture to obtain the target simulated weather picture.

[0102] In steps S411 to S412 of some embodiments, as Figure 5 shown, in controllable picture generation, a controllable weather generation model is constructed according to the CycleGAN. By using the CycleGAN to achieve global control, that is, controlling different weather conditions such as lighting, rain, and fog, a first simulated weather picture is generated. Exemplarily, through the controllable weather generation model and using the CycleGAN, the intensity of lighting, fog, and rain is controlled to generate a batch of simulated weather pictures with different weather conditions as Figure 6 shown.

[0103] In some embodiments, a cycle generative adversarial network is used to achieve unsupervised image-to-image translation between a source domain (such as the current scene) and a target domain (the generated scene). The cycle generative adversarial network consists of two generators and two discriminators, which are used to implement the mapping from the source domain to the target domain and from the target domain to the source domain respectively.

[0104] In steps S413 to S414 of some embodiments, the batch-generated first simulated weather pictures are preprocessed, that is, operations such as unifying the size of the first simulated weather pictures and standardizing the colors of the first simulated weather pictures are performed to obtain second simulated weather pictures, and label annotation operations of weather types are performed on the second simulated weather pictures according to the weather conditions corresponding when generating the pictures to obtain target simulated weather pictures. By preprocessing the first simulated weather pictures to obtain target simulated weather pictures, the training efficiency and performance of the subsequent weather recognition model can be improved, the generalization ability and robustness of the weather recognition model can be enhanced, and at the same time, it is also convenient for data storage, transmission and analysis.

[0105] In some embodiments, the step of constructing a target weather recognition model according to the target simulated weather pictures may include, but is not limited to, steps S421 to S423:

[0106] Step S421, construct an initial weather recognition model;

[0107] Step S422, construct a first loss function;

[0108] Step S423, input the target simulated weather pictures into the initial weather recognition model, and train the initial weather recognition model through a residual connection network and the first loss function to obtain the target weather recognition model.

[0109] In steps S421 to S423 of some embodiments, the labeled target simulated weather pictures are used as the input of the initial weather recognition model, and the initial weather recognition model is trained using a residual connection network (ResNet) and a first loss function to obtain a target weather recognition model, so that the target weather recognition model can automatically recognize different weather conditions and the intensities of different weather conditions. Optionally, a first loss function for training the initial weather recognition model is constructed using cross-entropy loss, and the expression of this first loss function is:

[0110] ;

[0111] Where represents the first loss function; represents the total number of weather types, ; represents the actual label, that is, the Labels of weather types; Represents the predicted probability, that is, the probability that the predicted weather type is the th weather type.

[0112] In step S500 of some embodiments, as Figure 5 shown, in controllable image generation, based on the conditional variational autoencoder (CVAE) structure, an initial context-aware spatial representation model is constructed. By fusing the depth map and the target mask map as the input of the initial context-aware spatial representation model, combined with the second loss function, the initial context-aware spatial representation model is trained to obtain the target context-aware spatial representation model, enabling the context-aware spatial representation model to accurately learn the semantic and spatial information of different scenes, providing reasonable candidate regions for subsequent object insertion, and laying a good foundation for subsequent acquisition of the target candidate position box.

[0113] In some embodiments, step S500 may include but is not limited to steps S510 to S540:

[0114] Step S510, fuse the semantic mask of the target mask map and the depth map to obtain a semantic depth fusion map;

[0115] Step S520, construct an initial context-aware spatial representation model based on the conditional variational autoencoder;

[0116] Step S530, construct a second loss function;

[0117] Step S540, input the semantic depth fusion map into the initial context-aware spatial representation model, and through the second loss function, train the initial context-aware spatial representation model to obtain the target context-aware spatial representation model.

[0118] In step S510 of some embodiments, the target mask map semantic mask and the depth map are fused to obtain a semantic depth fusion map The expression of which is:

[0119] ;

[0120] Wherein, represents the semantic depth fusion map; represents the semantic mask of the target mask map; represents the depth map.

[0121] The semantic depth fusion map is obtained by fusing the target mask map and the depth map. In the subsequent steps, this semantic depth fusion map is used as a condition to input into the initial situation awareness spatial representation model for training, which enables the situation awareness spatial representation model to consider the depth information of the target object when generating the insertion position box (target candidate position box), so that the front and rear occlusion relationships of the inserted target object and the size of the target object match the real scene.

[0122] In step S520 of some embodiments, based on the structure of the conditional variational autoencoder (CVAE), an initial situation awareness spatial representation model is constructed, which includes two parts: an encoder and a decoder.

[0123] In some embodiments, the encoder can integrate the semantic segmentation map (i.e., the mask map) and the position information (i.e., the depth map), and represent the implicit features of different scenes through the latent variable to capture the position distribution suitable for inserting objects. Exemplarily, in the conditional variational autoencoder, the encoder maps the input image (i.e., the semantic depth fusion map ) to the latent space to generate the latent variable . In the conditional variational autoencoder, the decoder uses the latent variable and the conditional information (i.e., the preset candidate position box, the target category of the target object, and the depth information of the target object) to decode and generate the candidate position box for inserting the object. The candidate position box generated by the decoder can reflect the semantic information and spatial distribution of the surrounding environment. Optionally, by selecting an object in the pictures of the autonomous driving picture dataset that has the same category or similar position distribution as the target object to be generated, the candidate box of this object is used as the preset candidate position box of the conditional information. In some optional embodiments, the latent variable generated by the encoder can be replaced by the randomly generated random Gaussian noise .

[0124] In step S530 of some embodiments, based on the training objective of the conditional variational autoencoder, a second loss function for training the initial situation awareness spatial representation model is constructed. The training objective of the conditional variational autoencoder is to minimize the total loss, including the reconstruction loss and the KL divergence loss. Optionally, the expression of the second loss function is:

[0125] ;

[0126] where represents the second loss function; represents the reconstruction loss, which is used to ensure that the candidate box generated by the decoder can reproduce the features of the input image ; Represents the KL divergence loss, which is used to make the output of the encoder Consistent with the prior distribution of the latent space Keep consistent.

[0127] In step S540 of some embodiments, the semantic depth fusion map is input into the initial context-aware spatial representation model, and combined with the second loss function, the initial context-aware spatial representation model is trained to obtain the target context-aware spatial representation model, so that the context-aware spatial representation model can accurately learn the semantic and spatial information of different scenes and provide reasonable candidate regions for subsequent object insertion.

[0128] In some embodiments, step S600 may include but is not limited to steps S610 to S660:

[0129] Step S610, classify and identify the autonomous driving image dataset according to the target weather recognition model to obtain the corresponding weather type;

[0130] Step S620, select a preset candidate position box through the autonomous driving image dataset;

[0131] Step S630, obtain the target category of the first target object in the autonomous driving image dataset and obtain the depth information of the first target object;

[0132] Step S640, generate test condition information according to the preset candidate position box, the target category and the depth information;

[0133] Step S650, input the test condition information and random Gaussian noise into the target context-aware spatial representation model to generate a target candidate position box;

[0134] Step S660, generate the target image sample according to the autonomous driving image dataset, the weather type and the target candidate position box.

[0135] In step S610 of some embodiments, the target weather recognition model can automatically recognize the pictures in the autonomous driving image dataset to obtain the corresponding weather type of the pictures and the intensity of the corresponding weather type. The target weather recognition model can accurately recognize and classify and analyze the weather type, which is convenient for subsequent users to analyze the distribution of pictures under different weather conditions in the autonomous driving image dataset.

[0136] In step S620 of some embodiments, in the images of the autonomous driving image dataset, several candidate boxes can be selected as preset candidate position boxes. Exemplarily, by selecting one or more preset objects in the images of the autonomous driving image dataset, the candidate boxes of the preset objects are used as the preset candidate position boxes, where the preset objects are objects with the same category or similar position distribution as the target objects expected to be generated in the target image samples.

[0137] In steps S630 to S650 of some embodiments, obtain the target category of the target object in the autonomous driving image dataset (which can be a vehicle, a pedestrian, a bicycle, a motorcycle, a traffic sign, etc., but is not limited thereto), and obtain the depth information of the corresponding target object. Use the preset candidate position box, the target category of the target object, and the depth information of the target object as test condition information and then use this test condition information and randomly generated random Gaussian noise as inputs to the target situation awareness spatial representation model, and generate target candidate position boxes through the decoder of the trained conditional variational autoencoder (CVAE) . These target candidate position boxes represent areas where objects may be suitable for insertion. Optionally, the target candidate position box can be represented as: where represents the center coordinates of the target candidate position box, represents the width of the target candidate position box, represents the height of the target candidate position box.

[0138] In some embodiments, step S650 may further include but is not limited to steps S651 to S653:

[0139] Step S651, input the test condition information and the random Gaussian noise into the target situation awareness spatial representation model to generate initial candidate position boxes;

[0140] Step S652, preset target constraint conditions; where the target constraint conditions include aspect ratio constraint conditions, ground intersection constraint conditions, and occlusion constraint conditions;

[0141] Step S653, optimize the initial candidate position boxes according to the target constraint conditions through linear approximation constraints to obtain the target candidate position boxes.

[0142] In step S651 of some embodiments, input the test condition information and random Gaussian noise into the target situation awareness spatial representation model to generate initial candidate position boxes , which can be represented as: , where represents the center coordinates of the initial candidate position box, represents the width of the initial candidate position box, represents the height of the initial candidate position box.

[0143] In step S652 of some embodiments, the aspect ratio constraint condition, the ground intersection constraint condition, and the occlusion constraint condition are preset as target constraint conditions. Exemplarily, the aspect ratio constraint condition is:

[0144] ;

[0145] where is the lower limit of the aspect ratio; is the upper limit of the aspect ratio.

[0146] The ground intersection constraint condition is:

[0147] ;

[0148] where is the ground mask to ensure that the inserted target object intersects with the ground area.

[0149] The occlusion constraint condition is:

[0150] ;

[0151] where is the initial candidate position box; is the point inside the initial candidate position box ; is the point the depth at the position; is the initial candidate position box the depth of. The occlusion constraint condition ensures that the inserted object does not occlude other objects with a depth less than it.

[0152] In step S653 of some embodiments, according to the preset target constraint conditions, the initial candidate position box is optimized using the Constrained Optimization BY Linear Approximations (COBYLA) optimization algorithm to obtain the target candidate position box . Among them, COBYLA is a gradient-free constrained optimization algorithm suitable for complex non-linear constraint problems. Then there is the following COBYLA optimization problem:

[0153] ;

[0154] Among them, are the aspect ratio constraint, the occlusion constraint, and the ground intersection constraint respectively; are the respective weights.

[0155] In step S660 of some embodiments, according to the weather types and their corresponding intensities identified by classifying the autonomous driving picture dataset and the target weather recognition model, and the target candidate position boxes, target picture samples are batch-generated. Exemplarily, the effects of the generated target picture samples are as Figure 8 shown.

[0156] In some embodiments, as Figure 7 shown, by noise sampling, random Gaussian noise is obtained, and the random Gaussian noise and the mask containing depth information are input into the encoder of the context-aware spatial representation model, an initial candidate position box can be generated. Combining the COBYLA optimization algorithm and the preset aspect ratio constraint condition, ground intersection constraint condition, and occlusion constraint condition, the initial candidate position box is optimized to generate a target candidate position box. According to the weather type identified by classifying the target weather recognition model, combined with the generated target candidate position box, picture samples of the corresponding weather type can be generated, and target objects of the corresponding target category are inserted into the target candidate position boxes in the picture samples to generate target picture samples.

[0157] As Figure 2 shown, a method for generating picture samples for an autonomous driving scenario provided by an embodiment of the present invention may further include, but is not limited to, steps S700 to S1000:

[0158] Step S700, display the first quantity of the first target object in the autonomous driving picture dataset and the second quantity of the second target object in the target picture sample, and generate a stacked bar chart;

[0159] Step S800, display the first distribution of the first target object in the autonomous driving picture dataset and the second distribution of the second target object in the target picture sample, and generate an object position distribution heat map;

[0160] Step S900, display the first multi-dimensional data of the autonomous driving picture dataset, and / or display the second multi-dimensional data of the target picture sample, and generate a multi-dimensional display diagram; wherein, the first multi-dimensional data includes the first quantity, the first depth value of the first target object, the first occlusion degree of the first target object, and the first weather condition; the second multi-dimensional data includes the second quantity, the second depth value of the second target object, the second occlusion degree of the second target object, and the second weather condition;

[0161] Step S1000: In response to the target dimension selection operation, display the autonomous driving picture dataset and the third distribution of the target picture sample in the target dimension, and generate a target dimension distribution heat map.

[0162] In steps S700 to S1000 of some embodiments, as Figure 3 shown, in the visualization front-end and back-end system, by generating and displaying views containing multi-dimensional information (which may include stacked bar charts, object position distribution heat maps, multi-dimensional display maps, and target dimension distribution heat maps, but are not limited thereto), the multi-dimensional information of the autonomous driving picture dataset and the multi-dimensional information of the picture samples can be visualized and analyzed.

[0163] In step S700 of some embodiments, through a combined view of a bar chart and a stacked chart, display the first quantity of the first target object in the autonomous driving picture dataset and the second quantity of the second target object in the target picture sample, and generate a stacked bar chart. Exemplarily, as Figure 9a shown in the stacked bar chart, the green bars represent the quantity of objects contained in the original dataset (such as the autonomous driving picture dataset), and the yellow bars represent the quantity of objects contained in the generated picture dataset (such as the target picture sample). The more the quantity, the higher the bar. Here, a logarithm with base 10 forms a spacing on the y-axis coordinate, which can intuitively display objects with a large quantity gap, such as cars and buses.

[0164] In step S800 of some embodiments, by displaying the first distribution of the first target object in the autonomous driving picture dataset and the second distribution of the second target object in the target picture sample, generate an object position distribution heat map. Exemplarily, taking a car as the target object analysis object, there is a Figure 9b shown object position distribution heat map. Referring to the Figure 9b shown object position distribution heat map, it represents the distribution of all cars in the dataset. The darker the color, the more the quantity of cars in this area is superimposed. The generation is indicated by dividing triangles. The triangles in the upper right corner of each small square represent the superimposed distribution of objects after generation and before generation, and the triangles in the lower left corner of each small square represent the superimposed distribution of objects before generation.

[0165] In step S900 of some embodiments, by combining a radar chart and a violin plot, the first multi-dimensional data of the autonomous driving picture dataset is displayed, and / or the second multi-dimensional data of the target picture sample is displayed, to generate a multi-dimensional display graph. Among them, the first multi-dimensional data includes the first quantity, the first depth value of the first target object, the first occlusion degree of the first target object, and the first weather condition; the second multi-dimensional data includes the second quantity, the second depth value of the second target object, the second occlusion degree of the second target object, and the second weather condition. Exemplarily, taking the statistical analysis of multiple dimensions of an image dataset as an example, there is a multi-dimensional display graph as follows Figure 9c as shown. Refer to Figure 9c . The function of this view is to perform statistical analysis on multiple dimensions of the image dataset, including the number of target objects in the image, the depth of the target object, the occlusion degree, and weather conditions, such as raindrops, fog, and light. The middle fan-shaped part shows the KL divergence between the distribution of each dimension and the uniform distribution.

[0166] In step S1000 of some embodiments, in response to a target dimension selection operation, the third distribution of the autonomous driving picture dataset and the target picture sample on the target dimension is displayed, to generate a target dimension distribution heat map. Exemplarily, the user can correspondingly select two target dimensions as the x-axis and y-axis. In response to the user's target dimension selection operation, a target dimension distribution heat map as shown in Figure 9d is generated. Refer to Figure 9d . The middle heat map matrix demonstrates the distribution of the dataset on the two selected dimensions. Each grid of the heat map matrix contains an image sample. The histograms along the two axes show the distribution of the dataset on one dimension. The user can select data samples in the heat map matrix to further explore the distribution of the samples. After generating the image samples, each corresponding grid in the heat map matrix will be divided into two triangles. The lower left part of each corresponding grid represents the original image data in that area, while the upper right part of each corresponding grid shows the combined distribution of the original image data and the batch-generated images. Through this target dimension distribution heat map, it can help the user explore the distribution of the dataset on certain specified dimensions and provide more details.

[0167] The embodiments of the present invention further provide a picture sample generation visualization analysis system for an autonomous driving scenario, which can implement the above-mentioned picture sample generation method for an autonomous driving scenario. The visualization analysis system includes:

[0168] A first module for obtaining an autonomous driving picture dataset;

[0169] A second module for performing mask extraction and data cleaning operations on the autonomous driving picture dataset to obtain a target mask map;

[0170] A third module, configured to perform in-depth analysis on the scenarios of the autonomous driving picture dataset to obtain a depth map;

[0171] A fourth module, configured to obtain a target simulated weather picture and construct a target weather recognition model according to the target simulated weather picture;

[0172] A fifth module, configured to construct a target situation awareness spatial representation model according to the target mask map and the depth map;

[0173] A sixth module, configured to obtain a target picture sample according to the target weather recognition model and the target situation awareness spatial representation model.

[0174] In some embodiments, the picture sample generation visualization analysis system for the autonomous driving scenario further includes:

[0175] A seventh module, configured to display the first quantity of the first target object in the autonomous driving picture dataset and the second quantity of the second target object in the target picture sample, and generate a stacked bar chart;

[0176] An eighth module, configured to display the first distribution of the first target object in the autonomous driving picture dataset and the second distribution of the second target object in the target picture sample, and generate an object position distribution heat map;

[0177] A ninth module, configured to display the first multi-dimensional data of the autonomous driving picture dataset and / or display the second multi-dimensional data of the target picture sample, and generate a multi-dimensional display diagram; wherein, the first multi-dimensional data includes the first quantity, the first depth value of the first target object, the first occlusion degree of the first target object, and the first weather condition; the second multi-dimensional data includes the second quantity, the second depth value of the second target object, the second occlusion degree of the second target object, and the second weather condition;

[0178] A tenth module, configured to display the third distribution of the autonomous driving picture dataset and the target picture sample in the target dimension in response to a target dimension selection operation, and generate a target dimension distribution heat map.

[0179] It can be understood that the content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0180] Such as Figure 10As shown, users can use the visualization analysis system for generating image samples for the autonomous driving scenario to achieve semi-automated batch generation of target image samples. Exemplarily, first, the user selects an image data set and the corresponding scenario type in the data view 1101, and then selects the object to be analyzed subsequently through the bar chart 1102 of the number of objects. Next, the data of the normal image area is selected as the original generation data by framing in the image data set distribution view 1103, and the controllable elements are adjusted in the control panel 1104 to achieve batch generation of target image samples. In addition, the original generation data and the target image samples can be compared and displayed before and after generation to form a comparison and analysis view 1105 of the image generation effect.

[0181] In some embodiments, users can perform global control through the global controllable sub-module to adjust the intensity of global environmental factors such as raindrops, fog, and light, and generate image samples as Figure 12 shown. In addition, local control can also be performed through the local controllable sub-module. As Figure 11 shown, the target area is framed in the heat map to generate a specific object, and its depth, occlusion degree and other parameters are further adjusted to generate image samples as Figure 11 shown in. This method realizes the generation of a large number of images meeting expectations on demand, effectively reducing the time and effort for processing images one by one.

[0182] In some embodiments, for the quantitative evaluation of batch images, the embodiments of the present invention adopt four image generation evaluation metrics: FID (Fréchet Inception Distance), IS (Inception Score), Precision, and Recall. FID focuses on the overall similarity between the generated image and the real image, while IS evaluates the quality and diversity of a single image. A lower FID score indicates higher image quality and realism, while a higher IS score indicates good quality and diversity. These two metrics are often used together to evaluate the performance of the generation model. In addition, the embodiments of the present invention also adopt improved Precision and Recall metrics, where Precision measures the similarity between the generated image and the real data, and Recall measures the diversity of the generated image in capturing the real data distribution. In the visualization analysis system for generating image samples for the autonomous driving scenario according to the embodiments of the present invention, the above four image generation evaluation metrics are used to compare and evaluate the image generation effect, and a schematic diagram of the comparison and evaluation of the image generation effect as Figure 13 shown is obtained. Refer to Figure 13, the four metric sizes of the image datasets generated in different batches are represented by quadrilaterals of different colors. These four metrics are presented in the form of a radar chart and normalized to the same ratio. It should be noted that only the FID metric is better when it is lower, while other metrics such as IS are better when they are higher. Therefore, for FID, the reciprocal form, i.e., 1 / FID, is used. As Figure 13 shown, in order to display the real images and compare the images before and after generation, this batch of image samples is also presented as a list of images. Users can select different dimensions to sort side by side and compare the effects of the images before and after generation, providing intuitive guidance for subsequent image generation.

[0183] An embodiment of the present invention also provides an electronic device, which includes a processor and a memory. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned method for generating image samples for an autonomous driving scenario. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0184] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0185] Referring to Figure 14 , Figure 14 shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0186] A processor 1201, which can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention;

[0187] A memory 1202, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1202 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1202 and are called by the processor 1201 to execute a method for generating image samples for an autonomous driving scenario according to the embodiments of the present invention;

[0188] An input / output interface 1203 for implementing information input and output;

[0189] A communication interface 1204 for implementing communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0190] A bus 1205 for transmitting information between various components of the device (such as the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204);

[0191] Among them, the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204 are communicatively connected to each other inside the device through the bus 1205.

[0192] An embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned method for generating picture samples for the autonomous driving scenario.

[0193] It can be understood that the content in the above method embodiments is applicable to this storage medium embodiment. The functions specifically implemented by this storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0194] An embodiment of the present invention also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the computer device executes the foregoing method for generating picture samples for the autonomous driving scenario.

[0195] In summary, the method for generating picture samples for the autonomous driving scenario and the visualization analysis system in the embodiments of the present invention have the following advantages:

[0196] 1. In the method for generating picture samples for the autonomous driving scenario and the visualization analysis system in the embodiments of the present invention, by combining traditional algorithms and existing advanced generation models, controllable generation is performed from both the global and local aspects, and through the visualization front-end and back-end systems, the generated picture dataset is explored and analyzed.

[0197] 2. The embodiments of the present invention combine traditional algorithms and generative models, propose an efficient semi-automatic and controllable sample generation mechanism, and provide a complete framework process. Based on deep learning technology, the present invention maximally meets the personalized needs of users for generating picture data, efficiently generates realistic pictures that conform to real-world laws in batches, and assists in expanding the picture dataset for autonomous driving.

[0198] 3. The embodiments of the present invention design and implement an interactive visual analysis system for picture samples in the autonomous driving scenario. The present invention proposes novel visualization designs and interaction methods, which not only support the analysis and mining of multi-dimensional picture datasets, but also support global and local picture generation, and fine-grained exploration and evaluation from the whole to the individual, so as to realize the combination of controllable picture generation and quantitative evaluation.

[0199] 4. The embodiments of the present invention explore the diversity and balance of picture data based on visualization technology. Real-world picture data has limitations that cannot be directly evaluated and feasible analysis methods in different dimensions. Therefore, for the picture dataset of autonomous driving, the present invention proposes an intuitive, novel, efficient and reliable visualization design, and combines the design concept and interaction criteria of human-in-the-loop to present the diversity and balance information of picture data to users.

[0200] In some alternative embodiments, the functions / operations mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, two consecutive blocks shown continuously may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical processes presented herein. Alternative embodiments are expected, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.

[0201] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features described may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of such modules would be understood within the ordinary skills of an engineer. Thus, those of ordinary skill in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0202] If the described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0203] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with such instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0204] More specific examples (a non-exhaustive list) of computer-readable media include the following: electrical connections (electronic devices) having one or more wirings, portable computer diskettes (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0205] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.

[0206] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0207] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

[0208] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.

Claims

1. A method for generating picture samples for autonomous driving scenarios, characterized in that, Including the following steps: Obtain an autonomous driving image dataset; Perform mask extraction and data cleaning operations on the autonomous driving image dataset to obtain a target mask image; Perform in-depth analysis on the scenes of the autonomous driving image dataset to obtain a depth map; Obtain a target simulated weather image, and construct a target weather recognition model according to the target simulated weather image; Construct a target situation awareness spatial representation model according to the target mask image and the depth map; Obtain a target image sample according to the target weather recognition model and the target situation awareness spatial representation model; Among them, obtaining the target image sample according to the target weather recognition model and the target situation awareness spatial representation model includes the following steps: Classify and recognize the autonomous driving image dataset according to the target weather recognition model to obtain the corresponding weather type; Select a preset candidate position box through the autonomous driving image dataset; Obtain the target category of the first target object in the autonomous driving image dataset, and obtain the depth information of the first target object; Generate test condition information according to the preset candidate position box, the target category, and the depth information; Input the test condition information and random Gaussian noise into the target situation awareness spatial representation model to generate a target candidate position box; Generate the target image sample according to the autonomous driving image dataset, the weather type, and the target candidate position box.

2. The method for generating a picture sample for an autonomous driving scenario according to claim 1, wherein It further includes the following steps: Display the first quantity of the first target object in the autonomous driving image dataset and the second quantity of the second target object in the target image sample to generate a stacked bar chart; Display the first distribution of the first target object in the autonomous driving image dataset and the second distribution of the second target object in the target image sample to generate an object position distribution heat map; Display the first multi-dimensional data of the autonomous driving image dataset, and / or display the second multi-dimensional data of the target image sample to generate a multi-dimensional display diagram; wherein, the first multi-dimensional data includes the first quantity, the first depth value of the first target object, the first occlusion degree of the first target object, and the first weather condition; the second multi-dimensional data includes the second quantity, the second depth value of the second target object, the second occlusion degree of the second target object, and the second weather condition; In response to a target dimension selection operation, display the third distribution of the autonomous driving image dataset and the target image sample in the target dimension to generate a target dimension distribution heat map.

3. A method for generating a picture sample for an autonomous driving scenario according to claim 1, characterized in that, The obtaining of the target simulated weather image includes the following steps: Construct a controllable weather generation model according to a cyclic generative adversarial network; Generate a first simulated weather image through the controllable weather generation model; Perform preprocessing on the first simulated weather image to obtain a second simulated weather image; Perform a weather type annotation operation on the second simulated weather image to obtain the target simulated weather image.

4. A method for generating a picture sample for an autonomous driving scenario according to claim 1, characterized in that, The constructing of the target weather recognition model according to the target simulated weather image includes the following steps: Construct an initial weather recognition model; Construct the first loss function; Input the target simulated weather image into the initial weather recognition model, and train the initial weather recognition model through the residual connection network and the first loss function to obtain the target weather recognition model.

5. A method for generating picture samples for an autonomous driving scenario according to claim 1, characterized in that The constructing the target context-aware spatial representation model according to the target mask image and the depth image includes the following steps: Fuse the semantic mask of the target mask image and the depth image to obtain a semantic-depth fused image; Construct an initial context-aware spatial representation model based on the conditional variational autoencoder; Construct the second loss function; Input the semantic-depth fused image into the initial context-aware spatial representation model, and train the initial context-aware spatial representation model through the second loss function to obtain the target context-aware spatial representation model.

6. The method for generating a picture sample for an autonomous driving scenario according to claim 1, characterized in that The inputting the test condition information and the random Gaussian noise into the target context-aware spatial representation model to generate the target candidate position box includes the following steps: Input the test condition information and the random Gaussian noise into the target context-aware spatial representation model to generate an initial candidate position box; Preset the target constraint conditions; wherein, the target constraint conditions include the aspect ratio constraint condition, the ground intersection constraint condition, and the occlusion constraint condition; Optimize the initial candidate position box according to the target constraint conditions through linear approximation constraints to obtain the target candidate position box.

7. A visual analysis system for generating picture samples for autonomous driving scenarios, characterized in that, Include: The first module is used to obtain the autonomous driving image dataset; The second module is used to perform mask extraction and data cleaning operations on the autonomous driving image dataset to obtain the target mask image; The third module is used to perform depth analysis on the scenes of the autonomous driving image dataset to obtain the depth image; The fourth module is used to obtain the target simulated weather image and construct the target weather recognition model according to the target simulated weather image; The fifth module is used to construct the target context-aware spatial representation model according to the target mask image and the depth image; The sixth module is used to obtain the target image sample according to the target weather recognition model and the target context-aware spatial representation model; Wherein, the sixth module is specifically used for: Classify and recognize the autonomous driving image dataset according to the target weather recognition model to obtain the corresponding weather type; Select the preset candidate position boxes through the autonomous driving image dataset; Obtain the target category of the first target object in the autonomous driving image dataset, and obtain the depth information of the first target object; Generate test condition information according to the preset candidate position box, the target category, and the depth information; Input the test condition information and the random Gaussian noise into the target context-aware spatial representation model to generate the target candidate position box; Generate the target image sample according to the autonomous driving image dataset, the weather type, and the target candidate position box.

8. A visualization analysis system for generating picture samples for an autonomous driving scenario according to claim 7, characterized in that Further include: The seventh module is used to display the first quantity of the first target object in the autonomous driving image dataset and the second quantity of the second target object in the target image sample, and generate a stacked bar chart; An eighth module, configured to display a first distribution of a first target object in the autonomous driving picture dataset and a second distribution of a second target object in the target picture sample, and generate an object position distribution heat map; A ninth module, configured to display first multi-dimensional data of the autonomous driving picture dataset and / or display second multi-dimensional data of the target picture sample, and generate a multi-dimensional display graph; wherein, the first multi-dimensional data includes the first quantity, a first depth value of the first target object, a first occlusion degree of the first target object, and a first weather condition; the second multi-dimensional data includes the second quantity, a second depth value of the second target object, a second occlusion degree of the second target object, and a second weather condition; A tenth module, configured to, in response to a target dimension selection operation, display a third distribution of the autonomous driving picture dataset and the target picture sample in the target dimension, and generate a target dimension distribution heat map.

9. An electronic device, characterized in that, Comprising a processor and a memory; The memory is used for storing programs; The processor executes the programs to implement the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • All-weather automatic driving data generation method and system based on diffusion model

    CN117079248A

  • Rainy day target detection method based on paired image data generation and knowledge distillation

    CN119649328A