Automatic driving scene-oriented picture sample generation method and visual analysis system

By generating picture samples of autonomous driving scenarios, the problem of insufficient diversity of existing data sets is solved, and the quality of the data set and the adaptability of the autonomous driving system are improved.

CN120107921AActive Publication Date: 2025-06-06SUN YAT SEN UNIV

Patent Information

Application Number
CN202510585818.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing autonomous driving image datasets have significant shortcomings in terms of diversity and cannot effectively cover less common scenarios in the real world, such as extreme weather conditions or the emergence of abnormal objects, resulting in inaccurate model predictions and immeasurable risks.

Method used

It provides a picture sample generation method for autonomous driving scenarios. By obtaining the data set of autonomous driving pictures, mask extraction and data cleaning, in-depth analysis of the scene, constructing a target weather recognition model and a situation-aware spatial representation model, and generating target picture samples.

Benefits of technology

It has achieved more targeted, richer and more comprehensive image sample generation, improved the quality of image samples, and enhanced the understanding and adaptability of the self-driving system to the complexity and unpredictability of the real world.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107921A_ABST
    Figure CN120107921A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic driving scene-oriented picture sample generation method and a visual analysis system. The method comprises the following steps of: obtaining an automatic driving picture data set; performing mask extraction and data cleaning operation on the automatic driving picture data set to obtain a target mask picture; performing depth analysis on a scene of the automatic driving picture data set to obtain a depth map; obtaining a target simulated weather picture, and constructing a target weather recognition model according to the target simulated weather picture; according to the target mask graph and the depth map, constructing a target context awareness space representation model; and obtaining a target picture sample according to the target weather recognition model and the target context awareness space representation model. The method can improve the quality of the picture sample, and can be widely applied to the technical field of image data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image data processing technology, and in particular to a method for generating image samples and a visual analysis system for autonomous driving scenarios. Background Art

[0002] Autonomous driving technology has injected new vitality into the automotive industry. The performance and reliability of current autonomous driving technology largely rely on high-quality training datasets. Most systems rely mainly on data collected from daily driving scenarios for training and evaluation. However, existing image datasets are significantly deficient in diversity. These datasets often do not include less common scenes in the real world, such as extreme weather conditions or the appearance of abnormal objects. When dealing with rare or complex scenarios, model performance is largely limited by the distribution of long-tail data (corner cases). When these situations occur, they may cause inaccurate predictions of models integrated into autonomous driving systems, bringing immeasurable risks. Summary of the invention

[0003] In view of this, the main purpose of the embodiments of the present invention is to provide a method for generating image samples and a visual analysis system for autonomous driving scenarios, in order to solve at least one of the problems in the prior art and improve the quality of image samples.

[0004] To achieve the above object, an embodiment of the present invention provides a method for generating image samples for an autonomous driving scenario, the method comprising the following steps: Get the autonomous driving image dataset; Performing mask extraction and data cleaning operations on the autonomous driving image dataset to obtain a target mask map; Performing depth analysis on the scene of the autonomous driving image dataset to obtain a depth map; Acquire a target simulated weather picture, and construct a target weather recognition model according to the target simulated weather picture; Constructing a target context-aware spatial representation model according to the target mask map and the depth map; A target picture sample is obtained according to the target weather recognition model and the target situation perception space representation model.

[0005] In some embodiments, the method for generating image samples for autonomous driving scenarios further includes the following steps: Displaying a first number of first target objects in the autonomous driving image dataset and a second number of second target objects in the target image sample to generate a stacked bar chart; Displaying a first distribution of a first target object in the autonomous driving image dataset and a second distribution of a second target object in the target image sample to generate an object position distribution heat map; Display the first multi-dimensional data of the autonomous driving image data set, and / or display the second multi-dimensional data of the target image sample, to generate a multi-dimensional display diagram; wherein the first multi-dimensional data includes the first quantity, the first depth value of the first target object, the first occlusion degree of the first target object, and the first weather condition; the second multi-dimensional data includes the second quantity, the second depth value of the second target object, the second occlusion degree of the second target object, and the second weather condition; In response to the target dimension selection operation, the autonomous driving image dataset and the third distribution of the target image samples on the target dimension are displayed to generate a target dimension distribution heat map.

[0006] In some embodiments, obtaining the target simulated weather picture includes the following steps: Construct a controllable weather generation model based on the cyclic generative adversarial network; Generate a first simulated weather picture by using the controllable weather generation model; Preprocessing the first simulated weather picture to obtain a second simulated weather picture; A weather type labeling operation is performed on the second simulated weather picture to obtain the target simulated weather picture.

[0007] In some embodiments, constructing a target weather recognition model according to the target simulated weather picture comprises the following steps: Construct an initial weather recognition model; Construct the first loss function; The target simulated weather picture is input into the initial weather recognition model, and the initial weather recognition model is trained through a residual connection network and the first loss function to obtain the target weather recognition model.

[0008] In some embodiments, constructing a target context-aware spatial representation model according to the target mask map and the depth map comprises the following steps: Fusing the semantic mask of the target mask map with the depth map to obtain a semantic depth fusion map; Based on the conditional variational autoencoder, an initial context-aware spatial representation model is constructed; Construct the second loss function; The semantic depth fusion map is input into the initial context-aware spatial representation model, and the initial context-aware spatial representation model is trained through the second loss function to obtain the target context-aware spatial representation model.

[0009] In some embodiments, obtaining a target image sample according to the target weather recognition model and the target context-aware spatial representation model comprises the following steps: According to the target weather recognition model, the autonomous driving image dataset is classified and recognized to obtain a corresponding weather type; Selecting a preset candidate location box through the autonomous driving image dataset; Obtaining a target category of a first target object in the autonomous driving image dataset, and obtaining depth information of the first target object; Generate test condition information according to the preset candidate position frame, the target category and the depth information; Inputting the test condition information and random Gaussian noise into the target context perception spatial representation model to generate a target candidate position frame; The target image sample is generated according to the autonomous driving image dataset, the weather type, and the target candidate location box.

[0010] In some embodiments, the step of inputting the test condition information and random Gaussian noise into the target context-aware spatial representation model to generate a target candidate position frame includes the following steps: Inputting the test condition information and the random Gaussian noise into the target context perception space representation model to generate an initial candidate position frame; Presetting target constraints; wherein the target constraints include aspect ratio constraints, ground intersection constraints, and occlusion constraints; According to the target constraint condition, the initial candidate position frame is optimized through linear approximation constraints to obtain the target candidate position frame.

[0011] To achieve the above purpose, another aspect of an embodiment of the present invention provides a system for generating visual analysis of image samples for autonomous driving scenarios, the visual analysis system comprising: The first module is used to obtain the autonomous driving image dataset; The second module is used to perform mask extraction and data cleaning operations on the autonomous driving image dataset to obtain a target mask map; The third module is used to perform depth analysis on the scene of the autonomous driving image dataset to obtain a depth map; The fourth module is used to obtain a target simulated weather picture and construct a target weather recognition model according to the target simulated weather picture; A fifth module is used to construct a target context-aware spatial representation model according to the target mask map and the depth map; The sixth module is used to obtain target image samples according to the target weather recognition model and the target situation perception space representation model.

[0012] In some embodiments, the system for generating a visual analysis of image samples for an autonomous driving scenario further includes: A seventh module is used to display a first number of first target objects in the autonomous driving image data set and a second number of second target objects in the target image sample to generate a stacked bar chart; An eighth module is used to display a first distribution of a first target object in the autonomous driving image data set and a second distribution of a second target object in the target image sample, and generate an object position distribution heat map; A ninth module is used to display the first multi-dimensional data of the autonomous driving image data set, and / or display the second multi-dimensional data of the target image sample, to generate a multi-dimensional display diagram; wherein the first multi-dimensional data includes the first quantity, the first depth value of the first target object, the first occlusion degree of the first target object, and the first weather condition; the second multi-dimensional data includes the second quantity, the second depth value of the second target object, the second occlusion degree of the second target object, and the second weather condition; The tenth module is used to display the autonomous driving image dataset and the third distribution of the target image samples on the target dimension in response to the target dimension selection operation, and generate a target dimension distribution heat map.

[0013] To achieve the above-mentioned purpose, another aspect of an embodiment of the present invention provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned method for generating image samples for autonomous driving scenarios.

[0014] To achieve the above-mentioned purpose, another aspect of an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned method for generating image samples for autonomous driving scenarios.

[0015] To achieve the above-mentioned purpose, another aspect of an embodiment of the present invention provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the aforementioned method for generating image samples for an autonomous driving scenario.

[0016] The embodiments of the present invention include at least the following beneficial effects: The present invention provides a method for generating image samples and a visualization analysis system for autonomous driving scenarios. The scheme obtains an autonomous driving image dataset; performs mask extraction and data cleaning operations on the autonomous driving image dataset to obtain a target mask map; performs in-depth analysis on the scene of the autonomous driving image dataset to obtain a depth map; obtains a target simulated weather image, and constructs a target weather recognition model based on the target simulated weather image; constructs a target situational awareness spatial representation model based on the target mask map and the depth map; obtains target image samples based on the target weather recognition model and the target situational awareness spatial representation model, thereby achieving the generation of more targeted, richer, and more comprehensive image samples and improving the quality of image samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 is a flowchart of a method for generating image samples for an autonomous driving scenario provided by an embodiment of the present invention; Figure 2 is a flowchart of another method for generating picture samples for autonomous driving scenarios provided by an embodiment of the present invention; Figure 3 This is an overall flow chart of image sample generation for autonomous driving scenarios provided by an embodiment of the present invention; Figure 4 is a flow chart of data feature extraction provided by an embodiment of the present invention; Figure 5 This is a flow chart of controllable generation of pictures provided by an embodiment of the present invention; Figure 6 It is a schematic diagram of the controllable generation effect of the global environment provided by an embodiment of the present invention; Figure 7 It is a flowchart of local controllable generation provided by an embodiment of the present invention; Figure 8 is a schematic diagram of the image sample generation effect provided by an embodiment of the present invention; Figure 9a-9d is a schematic diagram of visualizing multi-dimensional information of an image provided by an embodiment of the present invention; Fig.10 is a schematic diagram of a visual analysis system for generating image samples for autonomous driving scenarios provided by an embodiment of the present invention; Fig.11 is a schematic diagram of a controllable image sample generation effect of a visual analysis system provided by an embodiment of the present invention; Fig.12 is a schematic diagram of another controllable image sample generation effect of a visual analysis system provided by an embodiment of the present invention; Fig.13 is a schematic diagram of comparative evaluation of image sample generation effects provided by an embodiment of the present invention; Fig.14 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present invention, and they are only examples of devices and methods consistent with some aspects of the embodiments of the present invention as detailed in the attached claims.

[0020] It should be noted that, although the functional modules are divided in the system schematic diagram and the logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than the module division in the system or the flow chart. The terms "first / S100" and "second / S200" in the specification and claims and the above-mentioned drawings may be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiment of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determination".

[0021] The terms "at least one", "multiple", "each", "any", etc. used in the present invention, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.

[0023] Before describing the embodiments of the present invention in detail, some nouns and terms involved in the embodiments of the present invention are first described. The nouns and terms involved in the embodiments of the present invention are subject to the following explanations.

[0024] Deep learning: Deep learning is a branch of machine learning. It is based on the development of artificial neural networks and simulates the way the human brain thinks through a multi-layer network structure. It generates more abstract high-level representations (such as objects or scenes) by extracting and combining low-level features (such as edges and colors) layer by layer. The deep learning model optimizes parameters through the back-propagation algorithm, enabling the network to learn from large-scale data sets and automatically extract patterns in the data. Therefore, it is widely used in image recognition, speech recognition, natural language processing, medical image analysis, and autonomous driving.

[0025] Diffusion model: Diffusion model is a type of probabilistic generation model that generates new samples by simulating the gradual diffusion of noise and the reverse denoising process. The model starts with completely random noise and generates data consistent with the target distribution through denoising in multiple iterations, such as high-quality images or speech samples. The generation process of the diffusion model is similar to that of the Markov chain, so the amount of calculation is large, but the quality of the samples it generates is usually higher than that of traditional generative adversarial networks (GANs) and variational autoencoders (VAEs), especially in image generation and natural language tasks.

[0026] Controllable data generation: Controllable data generation refers to generating sample data that meets the requirements based on user-defined conditions or constraints. It allows users to control the style, content or structure of generated data through input parameters to meet the needs of specific scenarios.

[0027] Visual analysis system: The visual analysis system is an interactive system that combines data visualization and analysis functions, designed to help users quickly and intuitively understand the patterns and anomalies in complex data. By presenting data through visualization methods such as charts and dashboards, users can quickly capture the trends behind the data and deeply explore its internal relationships.

[0028] Human-Computer Interaction: Human-Computer Interaction (HCI) is an interdisciplinary field dedicated to studying the interaction between humans and computer systems and the mechanisms behind them. Its goal is to balance usability and functionality in design, and to improve user experience by optimizing interface design and interaction modes. Human-Computer Interaction involves multidisciplinary knowledge, such as psychology, cognitive science, computer science and design, and covers a variety of scenarios in applications, from desktop software interfaces to virtual reality, augmented reality and smart home devices.

[0029] In today's digital age, the rise of autonomous driving technology has not only injected new vitality into the automotive industry, but also led to a disruptive change in future travel modes, committed to providing people with a safer, more convenient and efficient transportation experience. The performance and reliability of autonomous driving technology largely rely on high-quality training data sets.

[0030] Currently, most systems rely mainly on data collected from everyday driving scenarios for training and evaluation. However, existing image datasets are significantly deficient in diversity, which affects the adaptability and robustness of autonomous driving systems. These datasets often do not include less common scenes in the real world, such as extreme weather conditions or the appearance of abnormal objects. This limitation may pose significant challenges to autonomous driving systems when dealing with rare or complex scenes. In other words, further improvement of model performance is largely limited by the distribution of long-tail data (corner cases). When these situations occur, they may cause inaccurate predictions of models integrated into autonomous driving systems, thereby bringing immeasurable risks.

[0031] Although some studies have specifically explored autonomous driving datasets under extreme weather conditions, these datasets are still difficult to balance in some dimensions. Therefore, in order to improve the safety and reliability of autonomous vehicles, it is crucial to ensure that the dataset covers various scenarios that drivers may encounter in the real world. Expanding the coverage of special scenarios usually requires two ways: one is to collect more data from the real world, and the other is to generate data samples for specific scenarios. For the former, collecting autonomous driving datasets usually relies on road testing. However, road testing is expensive and difficult to cover diverse driving scenarios, such as different weather and complex road conditions. For the latter, some studies use driving simulation software to generate virtual image samples to simulate real-world driving scenarios. Although these simulated data can supplement existing datasets to a certain extent, the fidelity of the simulation and the authenticity of the virtual scenes are still controversial. This gap between simulation and reality reveals the limitations of relying solely on virtual environments for comprehensive dataset enhancement.

[0032] In view of this, if Figure 3As shown in the figure, through an end-to-end autonomous driving perception data generation system, it is possible to use the rich real-life scene information in the real autonomous driving image dataset to extract data features from the real autonomous driving image dataset (such as Figure 4 ) and other processing, and can also generate modules through image controllable generation (such as Figure 5 The global control and local control (as shown in the figure) introduce various extreme and rare situations to further improve the diversity and reliability of the data set, thereby generating image samples in a targeted manner, so that the autonomous driving image data set can be effectively enhanced in the customized direction, thereby enhancing the autonomous driving system's understanding and adaptability to the complexity and unpredictability of the real world. In addition, multi-dimensional visualization analysis can be performed on the autonomous driving image data set and the generated image samples, and the image results generated by different controllable parameters can be compared through human-computer interaction and data-driven methods, so as to accurately explore and evaluate the image data set.

[0033] like Figure 1 As shown, an embodiment of the present invention provides a method for generating image samples for an autonomous driving scenario, and the method may include but is not limited to steps S100 to S600: Step S100, obtaining an autonomous driving image dataset; Step S200, performing mask extraction and data cleaning operations on the autonomous driving image dataset to obtain a target mask map; Step S300, performing depth analysis on the scene of the autonomous driving image dataset to obtain a depth map; Step S400, obtaining a target simulated weather picture, and constructing a target weather recognition model according to the target simulated weather picture; Step S500, constructing a target context-aware spatial representation model according to the target mask map and the depth map; Step S600, obtaining a target picture sample according to the target weather recognition model and the target situation perception space representation model.

[0034] In step S100 of some embodiments, an autonomous driving image dataset is obtained. The autonomous driving image dataset includes images of different weather conditions and multiple target object categories, where the different weather conditions can be sunny, rainy, foggy, strong light, etc., but not limited to these; the target object categories can be vehicles, pedestrians, bicycles, motorcycles, or traffic signs, etc., but not limited to these. The acquired autonomous driving image dataset can further obtain rich real scene information.

[0035] In step S200 of some embodiments, Figure 4 As shown, in data feature extraction, the acquired autonomous driving image dataset Perform mask extraction to extract the target object mask in the image and obtain the initial mask map. These masks can be used to separate each target object, laying a good foundation for subsequent depth estimation and prediction of the degree of occlusion of the target object. Furthermore, the initial mask map is cleaned, including removing images with unclear labels (including but not limited to the shape of the mask and the category of the object) and high noise, and deleting unrecognizable or undersized target objects, so as to obtain the target mask map. By performing mask extraction and data cleaning operations on the autonomous driving image dataset, the target mask map can reflect the distribution and characteristics of the target objects in the real environment.

[0036] In step S300 of some embodiments, Figure 4 As shown in the figure, in data feature extraction, the scene in the picture of the autonomous driving picture dataset is deeply analyzed through the depth estimation model to generate a depth map of each pixel in the picture, and then the distance relationship and spatial position of the target object in the scene can be inferred. In the subsequent steps, the depth map is combined with the mask information of the target object (target mask map) to obtain the specific depth of each target object, providing a basis for the subsequent object insertion algorithm.

[0037] In step S400 of some embodiments, Figure 4 As shown, in the data feature extraction, by simulating the complex weather conditions in the autonomous driving scene (such as rain, fog, sunshine of different intensities, etc., but not limited to these), simulated weather pictures with different weather conditions are generated in batches. These generated simulated weather pictures will be used to train the weather recognition model, thereby obtaining a trained weather recognition model. The trained weather recognition model can accurately identify and classify the weather types of the pictures in the autonomous driving data set, so as to facilitate subsequent users to analyze the distribution of pictures with different weather conditions in the autonomous driving picture data set.

[0038] In some embodiments, the step of obtaining the target simulated weather picture may include but is not limited to steps S411 to S414: Step S411, constructing a controllable weather generation model according to a cyclic generative adversarial network; Step S412, generating a first simulated weather picture through the controllable weather generation model; Step S413, preprocessing the first simulated weather picture to obtain a second simulated weather picture; Step S414, performing a weather type labeling operation on the second simulated weather picture to obtain the target simulated weather picture.

[0039] In steps S411 to S412 of some embodiments, Figure 5As shown in the figure, in the controllable generation of images, a controllable weather generation model is constructed based on a cyclic generative adversarial network (CycleGAN). Global control is achieved by using a cyclic generative adversarial network, that is, different weather conditions such as light, rain and fog are controlled to generate a first simulated weather image. For example, the controllable weather generation model uses a cyclic generative adversarial network to control the intensity of light, fog and rain to generate the following Figure 6 The batch shown has simulated weather pictures with different weather conditions.

[0040] In some embodiments, the cyclic generative adversarial network is used to achieve unsupervised image-to-image translation between a source domain (such as the current scene) and a target domain (the generated scene). The cyclic generative adversarial network consists of two generators and two discriminators, which are used to achieve the mapping from the source domain to the target domain and from the target domain to the source domain, respectively. In steps S413 to S414 of some embodiments, the first simulated weather pictures generated in batches are subjected to data preprocessing, that is, the size of the first simulated weather pictures is unified, the color of the first simulated weather pictures is standardized, and other preprocessing operations are performed to obtain a second simulated weather picture, and the second simulated weather picture is labeled with a weather type label according to the weather conditions corresponding to the generation of the picture to obtain a target simulated weather picture. By performing data preprocessing operations on the first simulated weather pictures to obtain the target simulated weather picture, the training efficiency and performance of the subsequent weather recognition model can be improved, the generalization ability and robustness of the weather recognition model can be enhanced, and the storage, transmission and analysis of data can also be facilitated.

[0041] In some embodiments, the step of constructing a target weather recognition model according to the target simulated weather picture may include but is not limited to steps S421 to S423: Step S421, constructing an initial weather recognition model; Step S422, constructing a first loss function; Step S423, inputting the target simulated weather picture into the initial weather recognition model, training the initial weather recognition model through a residual connection network and the first loss function, and obtaining the target weather recognition model.

[0042] In steps S421 to S423 of some embodiments, the labeled target simulated weather image is used as the input of the initial weather recognition model, and the initial weather recognition model is trained using a residual connection network (ResNet) and a first loss function to obtain a target weather recognition model, so that the target weather recognition model can automatically recognize different weather conditions and the intensity of different weather conditions. Optionally, a first loss function for training the initial weather recognition model is constructed using a cross entropy loss, and the expression of the first loss function is: ; in, represents the first loss function; Represents the total number of weather types, ; Represents the actual label, i.e. Labels for weather types; Represents the prediction probability, that is, the predicted weather type is The probability of a weather type.

[0043] In step S500 of some embodiments, Figure 5 As shown, in the controllable generation of pictures, based on the conditional variational autoencoder (CVAE) structure, an initial context-aware spatial representation model is constructed. The fusion of the depth map and the target mask map is used as the input of the initial context-aware spatial representation model. Combined with the second loss function, the initial context-aware spatial representation model is trained to obtain the target context-aware spatial representation model. The context-aware spatial representation model can accurately learn the semantics and spatial information of different scenes, provide reasonable candidate areas for subsequent object insertion, and lay a good foundation for the subsequent acquisition of target candidate position boxes.

[0044] In some embodiments, step S500 may include but is not limited to steps S510 to S540: Step S510, fusing the semantic mask of the target mask map and the depth map to obtain a semantic depth fusion map; Step S520, constructing an initial context-aware spatial representation model based on a conditional variational autoencoder; Step S530, constructing a second loss function; Step S540: input the semantic depth fusion map into the initial context-aware spatial representation model, and train the initial context-aware spatial representation model through the second loss function to obtain the target context-aware spatial representation model.

[0045] In step S510 of some embodiments, the target mask image Semantic mask of And the depth map After fusion, we can get a semantic deep fusion map The expression is: ; in, Represents the semantic deep fusion map; A semantic mask representing the target mask map; Represents a depth map.

[0046] A semantic depth fusion map is obtained by fusing the target mask map and the depth map. In the subsequent steps, the semantic depth fusion map is used as a conditional input into the initial context-aware spatial representation model for training. This enables the context-aware spatial representation model to consider the depth information of the target object when generating an insertion position frame (target candidate position frame), so that the front and back occlusion relationship of the inserted target object and the size of the target object match the real scene.

[0047] In step S520 of some embodiments, an initial context-aware spatial representation model is constructed based on the structure of a conditional variational autoencoder (CVAE), including two parts: an encoder and a decoder.

[0048] In some embodiments, the encoder can integrate the semantic segmentation map (i.e., mask map) and the position information (i.e., depth map) through the latent variable To represent the implicit features of different scenes, thereby capturing the distribution of locations suitable for inserting objects. For example, in a conditional variational autoencoder, the encoder takes the input image (i.e. semantic deep fusion graph ) is mapped to the latent space to generate latent variables In a conditional variational autoencoder, the decoder uses latent variables and condition information (i.e., preset candidate location box, target category of target object, and depth information of target object) to decode and generate candidate location box for inserting object. The candidate location box generated by decoder can reflect semantic information and spatial distribution of surrounding environment. Optionally, an object with the same category or similar location distribution as the target object to be generated is selected from the picture of autonomous driving picture data set, and the candidate box of the object is used as the preset candidate location box of condition information. In some optional embodiments, the latent variable generated by encoder Random Gaussian noise can be generated by replace.

[0049] In step S530 of some embodiments, a second loss function for training the initial context-aware spatial representation model is constructed by the training objective of the conditional variational autoencoder, and the training objective of the conditional variational autoencoder is to minimize the total loss, including the reconstruction loss and the KL divergence loss. Optionally, the expression of the second loss function is: ; in, represents the second loss function; Represents the reconstruction loss, which is used to ensure that the candidate boxes generated by the decoder can reproduce the input image Features; represents the KL divergence loss, which is used to transform the encoder output and the prior distribution of the latent space Be consistent.

[0050] In step S540 of some embodiments, the semantic depth fusion map is input into the initial context-aware spatial representation model, and the initial context-aware spatial representation model is trained in combination with the second loss function to obtain a target context-aware spatial representation model, so that the context-aware spatial representation model can accurately learn the semantics and spatial information of different scenes and provide reasonable candidate areas for subsequent object insertion.

[0051] In some embodiments, step S600 may include but is not limited to steps S610 to S660: Step S610, classifying and identifying the autonomous driving image dataset according to the target weather recognition model to obtain a corresponding weather type; Step S620, selecting a preset candidate position box through the autonomous driving image dataset; Step S630, obtaining a target category of a first target object in the autonomous driving image dataset, and obtaining depth information of the first target object; Step S640, generating test condition information according to the preset candidate position frame, the target category and the depth information; Step S650, inputting the test condition information and random Gaussian noise into the target context perception space representation model to generate a target candidate position frame; Step S660: Generate the target image sample according to the autonomous driving image dataset, the weather type, and the target candidate location box.

[0052] In step S610 of some embodiments, the target weather recognition model can be used to automatically recognize the images in the autonomous driving image dataset to obtain the weather type corresponding to the image and the intensity of the corresponding weather type. The target weather recognition model can accurately identify and classify the weather type, which is convenient for subsequent users to analyze the distribution of images of different weather conditions in the autonomous driving image dataset.

[0053] In step S620 of some embodiments, several candidate boxes may be selected as preset candidate position boxes in the picture of the autonomous driving picture dataset. Exemplarily, one or more preset objects are selected from the picture of the autonomous driving picture dataset, and the candidate boxes of the preset objects are used as preset candidate position boxes, wherein the preset objects are objects of the same category or similar position distribution as the target objects expected to be generated in the target picture sample.

[0054] In some embodiments, in steps S630 to S650, the target category of the target object in the autonomous driving image dataset is obtained (which may be a vehicle, pedestrian, bicycle, motorcycle or traffic sign, etc., but not limited to this), and the depth information of the corresponding target object is obtained. The preset candidate position box, the target category of the target object and the depth information of the target object are used as the test condition information. , and then the test condition information and randomly generated random Gaussian noise Input into the target context-aware spatial representation model, and generate the target candidate location box through the decoder of the trained conditional variational autoencoder (CVAE) These target candidate position boxes represent areas that may be suitable for inserting objects. Optionally, the target candidate position boxes can be expressed as: ,in, Represents the center coordinates of the target candidate location box, Indicates the width of the target candidate location box, Indicates the height of the target candidate location box.

[0055] In some embodiments, step S650 may also include but is not limited to steps S651 to S653: Step S651, inputting the test condition information and the random Gaussian noise into the target context perception space representation model to generate an initial candidate position frame; Step S652, presetting target constraints; wherein the target constraints include aspect ratio constraints, ground intersection constraints, and occlusion constraints; Step S653: According to the target constraint condition, the initial candidate position frame is optimized through linear approximation constraints to obtain the target candidate position frame.

[0056] In step S651 of some embodiments, the test condition information and random Gaussian noise Input into the target context-aware spatial representation model to generate the initial candidate location box , which can be expressed as: ,in, Represents the center coordinates of the initial candidate location box, Indicates the width of the initial candidate location box, Indicates the height of the initial candidate location box.

[0057] In step S652 of some embodiments, aspect ratio constraints, ground intersection constraints, and occlusion constraints are preset as target constraints. For example, the aspect ratio constraints are: ; in, is the lower limit of the aspect ratio; is the upper limit of the aspect ratio.

[0058] The ground intersection constraint condition is: ; in, is the ground mask, ensuring that the inserted target object intersects with the ground area.

[0059] The occlusion constraint condition is: ; in, is the initial candidate position box; is the point inside the initial candidate location box ; For point Depth at location; is the initial candidate location box The occlusion constraint ensures that the inserted object will not occlude other objects with a depth less than it.

[0060] In step S653 of some embodiments, according to the preset target constraint conditions, a linear approximation constraint (Constrained Optimization BY Linear Approximations, COBYLA) optimization algorithm is used to optimize the initial candidate position frame. Optimize to get the target candidate position box Among them, COBYLA is a gradient-independent constrained optimization algorithm suitable for complex nonlinear constraint problems. The following COBYLA optimization problem is: ; in, They are aspect ratio constraint, occlusion constraint, and ground intersection constraint; For each weight.

[0061] In step S660 of some embodiments, target image samples are generated in batches according to the autonomous driving image dataset, the weather type and its corresponding intensity obtained by classification and identification of the target weather recognition model, and the target candidate position frame. For example, the generated target image samples have the following effects: Figure 8 shown.

[0062] In some embodiments, Figure 7As shown, random Gaussian noise is obtained by noise sampling, and the random Gaussian noise and the mask containing depth information are input into the encoder of the context-aware spatial representation model to generate an initial candidate location frame. The initial candidate location frame is optimized by combining the COBYLA optimization algorithm and the preset aspect ratio constraints, ground intersection constraints, and occlusion constraints to generate a target candidate location frame. According to the weather type classified and identified by the target weather recognition model, combined with the generated target candidate location frame, a picture sample of the corresponding weather type can be generated, and the target object of the corresponding target category is inserted into the target candidate location frame in the picture sample to generate a target picture sample.

[0063] like Figure 2 As shown, the method for generating image samples for an autonomous driving scenario provided by an embodiment of the present invention may also include but is not limited to steps S700 to S1000: Step S700, displaying a first number of first target objects in the autonomous driving image data set and a second number of second target objects in the target image sample, and generating a stacked bar chart; Step S800, displaying a first distribution of a first target object in the autonomous driving image dataset and a second distribution of a second target object in the target image sample, and generating an object position distribution heat map; Step S900, displaying the first multi-dimensional data of the autonomous driving image data set, and / or displaying the second multi-dimensional data of the target image sample, to generate a multi-dimensional display diagram; wherein the first multi-dimensional data includes the first quantity, the first depth value of the first target object, the first occlusion degree of the first target object, and the first weather condition; the second multi-dimensional data includes the second quantity, the second depth value of the second target object, the second occlusion degree of the second target object, and the second weather condition; Step S1000, in response to the target dimension selection operation, displays the autonomous driving image dataset and the third distribution of the target image samples on the target dimension, and generates a target dimension distribution heat map.

[0064] In steps S700 to S1000 of some embodiments, Figure 3 As shown, in the visualization front-end and back-end systems, by generating and displaying views containing multi-dimensional information (which may include stacked bar charts, object location distribution heat maps, multi-dimensional display maps, and target dimension distribution heat maps, but are not limited to these), the multi-dimensional information of the autonomous driving image data set and the multi-dimensional information of the image samples can be visualized and analyzed.

[0065] In step S700 of some embodiments, a stacked bar chart is generated by displaying a first number of first target objects in the autonomous driving image dataset and a second number of second target objects in the target image sample through a combined view of a bar chart and a stacked chart. Figure 9a In the stacked bar chart shown in the figure, the green bars represent the number of objects in the original dataset (such as the autonomous driving image dataset), and the yellow bars represent the number of objects in the generated image dataset (such as the target image sample). The higher the number, the higher the bar height. Here, the logarithm with the base 10 forms a spacing of the y-axis coordinates, which can intuitively show objects with a large difference in number, such as cars and buses.

[0066] In step S800 of some embodiments, a heat map of object position distribution is generated by displaying a first distribution of a first target object in the autonomous driving image dataset and a second distribution of a second target object in the target image sample. Figure 9b The heat map of object position distribution is shown. Figure 9b The heat map of object position distribution shown in the figure shows the distribution of all cars in the data set. The darker the color, the more cars are contained in this area. The triangles are divided to indicate that new objects are generated in this range. The triangles in the upper right corner of each small square represent the superposition of objects before and after generation, and the triangles in the lower left corner of each small square represent the superposition of objects before generation.

[0067] In step S900 of some embodiments, the first multi-dimensional data of the autonomous driving image dataset is displayed by combining a radar chart with a violin chart, and / or the second multi-dimensional data of the target image sample is displayed to generate a multi-dimensional display chart. The first multi-dimensional data includes the first quantity, the first depth value of the first target object, the first occlusion degree of the first target object, and the first weather condition; the second multi-dimensional data includes the second quantity, the second depth value of the second target object, the second occlusion degree of the second target object, and the second weather condition. Exemplarily, taking the statistical analysis of multiple dimensions of the image dataset as an example, as shown in FIG. Fig.9c The multi-dimensional display diagram shown. Fig.9c , which performs statistical analysis on multiple dimensions of the image dataset, including the number of objects in the image, the depth of the objects, the degree of occlusion, and weather conditions such as raindrops, fog, and lighting. The fan-shaped part in the middle shows the KL divergence between the distribution of each dimension and the uniform distribution.

[0068] In step S1000 of some embodiments, in response to the target dimension selection operation, the autonomous driving image dataset and the third distribution of the target image samples on the target dimension are displayed to generate a target dimension distribution heat map. For example, the user can select two target dimensions as the x-axis and the y-axis accordingly, and in response to the user's target dimension selection operation, a heat map of the target dimension distribution is generated. Figure 9d The target dimension distribution heat map is shown. Figure 9d , the heat map matrix in the middle demonstrates the distribution of the dataset in two selected dimensions, and each grid of the heat map matrix contains image samples. The histograms along the two axes show the distribution of the dataset in one dimension. Users can select data samples in the heat map matrix to further explore the distribution of the samples. After the image samples are generated, each corresponding grid in the heat map matrix will be divided into two triangles, and the lower left part of each corresponding grid represents the original image data in the area, while the upper right part of each corresponding grid shows the combined distribution of the original image data and the batch generated images. This target dimension distribution heat map can help users explore the distribution of the dataset in certain specified dimensions and provide more details.

[0069] The embodiment of the present invention further provides a system for generating visual analysis of picture samples for autonomous driving scenarios, which can implement the above-mentioned method for generating picture samples for autonomous driving scenarios. The visual analysis system includes: The first module is used to obtain the autonomous driving image dataset; The second module is used to perform mask extraction and data cleaning operations on the autonomous driving image dataset to obtain a target mask map; The third module is used to perform depth analysis on the scene of the autonomous driving image dataset to obtain a depth map; The fourth module is used to obtain a target simulated weather picture and construct a target weather recognition model according to the target simulated weather picture; A fifth module is used to construct a target context-aware spatial representation model according to the target mask map and the depth map; The sixth module is used to obtain target image samples according to the target weather recognition model and the target situation perception space representation model.

[0070] In some embodiments, the system for generating a visual analysis of image samples for an autonomous driving scenario further includes: A seventh module is used to display a first number of first target objects in the autonomous driving image data set and a second number of second target objects in the target image sample to generate a stacked bar chart; An eighth module is used to display a first distribution of a first target object in the autonomous driving image data set and a second distribution of a second target object in the target image sample, and generate an object position distribution heat map; A ninth module is used to display the first multi-dimensional data of the autonomous driving image data set, and / or display the second multi-dimensional data of the target image sample, to generate a multi-dimensional display diagram; wherein the first multi-dimensional data includes the first quantity, the first depth value of the first target object, the first occlusion degree of the first target object, and the first weather condition; the second multi-dimensional data includes the second quantity, the second depth value of the second target object, the second occlusion degree of the second target object, and the second weather condition; The tenth module is used to display the autonomous driving image dataset and the third distribution of the target image samples on the target dimension in response to the target dimension selection operation, and generate a target dimension distribution heat map.

[0071] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0072] like Fig.10 As shown, users can generate a visual analysis system for image samples for autonomous driving scenarios to achieve semi-automatic batch generation of target image samples. Exemplarily, the user first selects an image dataset and a corresponding scene type in the data view 1101, and then selects the object to be analyzed later through the bar chart 1102 of the number of objects. Next, in the image dataset distribution view 1103, select the data in the normal image area as the original generated data, and adjust the controllable elements in the control panel 1104 to achieve batch generation of target image samples. In addition, the original generated data and the target image samples before and after generation can be compared and displayed to form a picture generation effect comparison and analysis view 1105.

[0073] In some embodiments, the user can use the global controllable submodule to adjust the intensity of global environmental factors such as raindrops, fog, and light to generate Fig.12 In addition, local control is also performed through local controllable submodules, such as Fig.11 As shown in , the target area is selected in the heat map to generate a specific object, and its depth, occlusion degree and other parameters are further adjusted to generate the following Fig.11 This method can generate a large number of expected images on demand, effectively reducing the time and effort of processing images one by one.

[0074] In some embodiments, for the quantitative evaluation of batch images, the embodiments of the present invention adopt four image generation evaluation indicators: FID (Fréchet Inception Distance), IS (Inception Score), Precision, and Recall. FID focuses on the overall similarity between the generated image and the real image, while IS evaluates the quality and diversity of a single image. A lower FID score indicates higher image quality and realism, while a higher IS score indicates good quality and diversity. These two indicators are often used together to evaluate the performance of the generation model. In addition, the embodiment of the present invention also adopts improved Precision and Recall indicators, where Precision measures the similarity between the generated image and the real data, and Recall measures the diversity of the generated image in capturing the distribution of the real data. In the image sample generation visualization analysis system for autonomous driving scenarios in the embodiment of the present invention, the above four image generation evaluation indicators are used to compare and evaluate the image generation effects, and the following are obtained. Fig.13 The schematic diagram of the comparative evaluation of the image sample generation effect is shown. Fig.13 , using quadrilaterals of different colors to represent the sizes of the four indicators of the image datasets generated in different batches, presenting these four indicators in the form of a radar chart and normalizing them to the same scale. It is worth noting that only the FID indicator is better when it is lower, while other indicators such as IS are better when they are higher. Therefore, for FID, the inverse form is used, that is, 1 / FID. Fig.13 As shown in the figure, in order to show the real images and compare the images before and after generation, this batch of image samples is also presented as an image list. Users can choose different dimensions to sort and compare the effects of the images before and after generation, providing intuitive guidance for subsequent image generation.

[0075] The embodiment of the present invention further provides an electronic device, which includes a processor and a memory, wherein the memory stores a computer program, and when the processor executes the computer program, the above-mentioned method for generating image samples for an autonomous driving scenario is implemented. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.

[0076] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0077] refer to Fig.14 , Fig.14 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes: The processor 1201 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention; The memory 1202 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1202 can store an operating system and other applications. When the technical solution provided in the embodiments of this specification is implemented by software or firmware, the relevant program code is stored in the memory 1202, and the processor 1201 calls and executes a method for generating image samples for an autonomous driving scenario in an embodiment of the present invention; Input / output interface 1203, used to implement information input and output; The communication interface 1204 is used to realize the communication interaction between the device and other devices. The communication can be realized through a wired manner (such as USB, network cable, etc.) or a wireless manner (such as mobile network, WIFI, Bluetooth, etc.); A bus 1205 that transmits information between various components of the device (e.g., the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204); The processor 1201 , the memory 1202 , the input / output interface 1203 and the communication interface 1204 are connected to each other in communication within the device via the bus 1205 .

[0078] An embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned method for generating image samples for autonomous driving scenarios.

[0079] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0080] The embodiment of the present invention further provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the aforementioned method for generating image samples for an autonomous driving scenario.

[0081] In summary, the image sample generation method and visualization analysis system for autonomous driving scenarios according to the embodiments of the present invention have the following advantages: 1. In the image sample generation method and visualization analysis system for autonomous driving scenarios in the embodiments of the present invention, controllable generation is performed from both global and local aspects by combining traditional algorithms and existing advanced generation models, and the generated image data set is explored and analyzed through a visualization front-end and back-end system.

[0082] 2. The embodiment of the present invention combines traditional algorithms and generation models to propose an efficient semi-automatic and controllable sample generation mechanism and provides a complete framework process. The present invention generates image data that meets the personalized needs of users to the greatest extent based on deep learning technology, efficiently generates realistic images that conform to the laws of reality in batches, and assists in expanding the autonomous driving image data set.

[0083] 3. The embodiment of the present invention designs and implements an interactive visual analysis system for image samples of autonomous driving scenarios. The present invention proposes a novel visualization design and interaction method, which not only supports multi-dimensional image data set analysis and mining, but also supports global and local image generation, fine-grained exploration and evaluation from the whole to the individual, so as to realize the combination of controllable image generation and quantitative evaluation.

[0084] 4. The embodiment of the present invention explores the diversity and balance of image data based on visualization technology. Real image data has limitations that cannot be directly evaluated and feasible analysis methods in different dimensions. To this end, the present invention proposes an intuitive, novel, efficient and reliable visualization design for the autonomous driving image dataset, combining the design concept and interaction criteria of people in the loop to present the diversity and balance information of image data to users.

[0085] In some selectable embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Selectable embodiments are expected, wherein the order of various operations is changed and the sub-operation of a part for which is described as a larger operation is performed independently.

[0086] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise specified, one or more of the functions and / or features described may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present invention. More specifically, in view of the properties, functions, and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional skills of the engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0087] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.

[0088] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0089] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0090] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0091] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0092] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.

[0093] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present invention.

Claims

1. A method for generating image samples for autonomous driving scenarios, characterized in that: The following steps are involved: Get the autonomous driving image dataset; Performing mask extraction and data cleaning operations on the autonomous driving image dataset to obtain a target mask map; Performing depth analysis on the scene of the autonomous driving image dataset to obtain a depth map; Acquire a target simulated weather picture, and construct a target weather recognition model according to the target simulated weather picture; Constructing a target context-aware spatial representation model according to the target mask map and the depth map; A target picture sample is obtained according to the target weather recognition model and the target situation perception space representation model.

2. The method for generating image samples for autonomous driving scenarios according to claim 1, characterized in that: The following steps are also included: Displaying a first number of first target objects in the autonomous driving image dataset and a second number of second target objects in the target image sample to generate a stacked bar chart; Displaying a first distribution of a first target object in the autonomous driving image dataset and a second distribution of a second target object in the target image sample to generate an object position distribution heat map; Display the first multi-dimensional data of the autonomous driving image data set, and / or display the second multi-dimensional data of the target image sample, to generate a multi-dimensional display diagram; wherein the first multi-dimensional data includes the first quantity, the first depth value of the first target object, the first occlusion degree of the first target object, and the first weather condition; the second multi-dimensional data includes the second quantity, the second depth value of the second target object, the second occlusion degree of the second target object, and the second weather condition; In response to the target dimension selection operation, the autonomous driving image dataset and the third distribution of the target image samples on the target dimension are displayed to generate a target dimension distribution heat map.

3. The method for generating image samples for autonomous driving scenarios according to claim 1, characterized in that: The step of obtaining the target simulated weather picture comprises the following steps: Construct a controllable weather generation model based on the cyclic generative adversarial network; Generate a first simulated weather picture by using the controllable weather generation model; Preprocessing the first simulated weather picture to obtain a second simulated weather picture; A weather type labeling operation is performed on the second simulated weather picture to obtain the target simulated weather picture.

4. The method for generating image samples for autonomous driving scenarios according to claim 1, characterized in that: The step of constructing a target weather recognition model according to the target simulated weather picture comprises the following steps: Construct an initial weather recognition model; Construct the first loss function; The target simulated weather picture is input into the initial weather recognition model, and the initial weather recognition model is trained through a residual connection network and the first loss function to obtain the target weather recognition model.

5. The method for generating image samples for autonomous driving scenarios according to claim 1, characterized in that: The step of constructing a target context-aware spatial representation model according to the target mask map and the depth map comprises the following steps: Fusing the semantic mask of the target mask map with the depth map to obtain a semantic depth fusion map; Based on the conditional variational autoencoder, an initial context-aware spatial representation model is constructed; Construct the second loss function; The semantic depth fusion map is input into the initial context-aware spatial representation model, and the initial context-aware spatial representation model is trained through the second loss function to obtain the target context-aware spatial representation model.

6. The method for generating image samples for autonomous driving scenarios according to claim 1, characterized in that: The step of obtaining a target image sample according to the target weather recognition model and the target situation perception space representation model comprises the following steps: According to the target weather recognition model, the autonomous driving image dataset is classified and recognized to obtain a corresponding weather type; Selecting a preset candidate location box through the autonomous driving image dataset; Obtaining a target category of a first target object in the autonomous driving image dataset, and obtaining depth information of the first target object; Generate test condition information according to the preset candidate position frame, the target category and the depth information; Inputting the test condition information and random Gaussian noise into the target context perception spatial representation model to generate a target candidate position frame; The target image sample is generated according to the autonomous driving image dataset, the weather type, and the target candidate location box.

7. The method for generating image samples for autonomous driving scenarios according to claim 6, characterized in that: The step of inputting the test condition information and random Gaussian noise into the target context perception spatial representation model to generate a target candidate position frame comprises the following steps: Inputting the test condition information and the random Gaussian noise into the target context perception space representation model to generate an initial candidate position frame; Presetting target constraints; wherein the target constraints include aspect ratio constraints, ground intersection constraints, and occlusion constraints; According to the target constraint condition, the initial candidate position frame is optimized through linear approximation constraints to obtain the target candidate position frame.

8. A visual analysis system for generating image samples for autonomous driving scenarios, characterized in that: include: The first module is used to obtain the autonomous driving image dataset; The second module is used to perform mask extraction and data cleaning operations on the autonomous driving image dataset to obtain a target mask map; The third module is used to perform depth analysis on the scene of the autonomous driving image dataset to obtain a depth map; The fourth module is used to obtain a target simulated weather picture and construct a target weather recognition model according to the target simulated weather picture; A fifth module is used to construct a target context-aware spatial representation model according to the target mask map and the depth map; The sixth module is used to obtain target image samples according to the target weather recognition model and the target situation perception space representation model.

9. The visual analysis system for generating image samples for autonomous driving scenarios according to claim 8, characterized in that: Also includes: A seventh module is used to display a first number of first target objects in the autonomous driving image data set and a second number of second target objects in the target image sample to generate a stacked bar chart; An eighth module is used to display a first distribution of a first target object in the autonomous driving image data set and a second distribution of a second target object in the target image sample, and generate an object position distribution heat map; A ninth module is used to display the first multi-dimensional data of the autonomous driving image data set, and / or display the second multi-dimensional data of the target image sample, to generate a multi-dimensional display diagram; wherein the first multi-dimensional data includes the first quantity, the first depth value of the first target object, the first occlusion degree of the first target object, and the first weather condition; the second multi-dimensional data includes the second quantity, the second depth value of the second target object, the second occlusion degree of the second target object, and the second weather condition; The tenth module is used to display the autonomous driving image dataset and the third distribution of the target image samples on the target dimension in response to the target dimension selection operation, and generate a target dimension distribution heat map.

10. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • All-weather automatic driving data generation method and system based on diffusion model

    CN117079248A

  • Rainy day target detection method based on paired image data generation and knowledge distillation

    CN119649328A

Cited By

  • Model training method and device and electronic equipment

    CN120953955A