Visual annotation data generation method and electronic equipment

By building vision sensors in the simulator for autonomous driving simulation, the simulation image data is migrated to the real environment using the simulation migration model to generate visual annotation data, which solves the problems of high data annotation cost, low data utilization rate and model overfitting, and achieves efficient and robust neural network training.

CN120148035APending Publication Date: 2025-06-13ECARX (HUBEI) TECHCO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510199718.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, there are problems such as high cost of data labeling, low data utilization, overfitting of model training and difficulty in generalization.

Method used

By building the vehicle's vision sensor in the simulator, performing autonomous driving simulation, obtaining simulation image data and truth value tags, inputting them into the training-completed simulation migration model, migrating the simulation image data to the real environment, and generating visual label data.

Benefits of technology

Automatic labeling of autonomous driving visual data is realized, the cost of manual labeling is reduced, the data labeling efficiency is improved, and the problem of large differences and overfitting between simulated data and real data is solved, so that simulated data can be used to train robust neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148035A_ABST
    Figure CN120148035A_ABST
Patent Text Reader

Abstract

The invention provides a visual annotation data generation method and electronic equipment, and the method comprises the steps: constructing a visual sensor of a vehicle in a simulator, carrying out the simulation of the automatic driving of the vehicle through the simulator, and obtaining the simulation image data collected by the visual sensor in a simulation environment and a corresponding true value label, the method comprises the steps of training a simulation migration model, inputting simulation image data into the trained simulation migration model, migrating the simulation image data from a simulation environment to a real environment to obtain corresponding real image data, determining visual annotation data according to the real image data and a corresponding truth value label, and achieving automatic annotation of automatic driving visual data. According to the method, manual annotation is not needed, the data annotation efficiency is improved, the data annotation cost is reduced, the problems that the difference between the simulation data and the real data is large, the details of the simulation data are not rich enough and the like can be solved, and therefore the simulation data can be used for training the neural network with high robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data annotation, and particularly to a method for generating visual annotation data and an electronic device. Background Art

[0002] With the development of the autonomous driving field and the improvement of computing power, new perception algorithm models based on deep learning have emerged continuously. These models can automatically perform in-depth representation and extraction of environmental features, have stronger expressive power and adaptability, and can extract high-level semantic information from raw sensor data, thus solving the accuracy and robustness problems of traditional models in complex scenarios and dynamic environments, making high-level intelligent driving possible.

[0003] Building a neural network model with strong robustness requires the support of a large amount of annotated data. Traditional annotation methods require corresponding data acquisition devices and rely on manual annotation, which requires a large amount of human and time costs, and can only be targeted at a single training task, resulting in low data utilization. The simulation-based annotation method will cause the simulation data to lack details in the real world, making the simulation data highly repetitive, and the models trained on it often show overfitting to the simulation data, that is, the models perform well on the simulation data but are difficult to generalize on the real data. Summary of the Invention

[0004] In view of the above-mentioned defects or deficiencies in the prior art, this application aims to provide a method for generating visual annotation data and an electronic device to solve the problems of high data annotation cost, low data utilization, overfitting in model training, and difficulty in generalization in the prior art.

[0005] An embodiment of this application provides a method for generating visual annotation data, and the method includes:

[0006] Construct a visual sensor of a vehicle in a simulator, and simulate the autonomous driving of the vehicle through the simulator to obtain simulation image data collected by the visual sensor in the simulation environment and corresponding ground truth labels;

[0007] Input the simulation image data into a trained simulation transfer model to transfer the simulation image data from the simulation environment to the real environment, and obtain corresponding real image data;

[0008] Determine visual annotation data according to the real image data and the corresponding ground truth labels.

[0009] Optionally, the training process of the simulation transfer model includes:

[0010] Collect a plurality of original simulation images based on the simulator, and select sample simulation images from the plurality of original simulation images;

[0011] Get multiple sample real images and get a pre-trained style transfer model;

[0012] Based on each sample simulated image and each sample real image, the style transfer model is trained to obtain the simulated transfer model.

[0013] Optionally, collecting a plurality of original simulation images based on the simulator includes:

[0014] Obtaining an environment configuration file, wherein the environment configuration file is used to configure a simulation environment in the simulator;

[0015] Running the simulator based on the environment configuration file to obtain a plurality of original simulation images;

[0016] The environment configuration file is updated, and the step of running the simulator based on the environment configuration file is returned to be executed.

[0017] Optionally, updating the environment configuration file includes:

[0018] At least one of the weather type, road condition parameter, light parameter, obstacle quantity, obstacle type and obstacle position of the environment configuration file is updated.

[0019] Optionally, before selecting a sample simulation image from a plurality of original simulation images, the method further includes:

[0020] Acquire real scene distribution information, wherein the real scene distribution information includes the proportion of images in each real scene in the real data set;

[0021] Based on the real scene distribution information and the simulation scene corresponding to each of the original simulation images, some of the original simulation images are eliminated.

[0022] Optionally, before obtaining the real scene distribution information, the following is also included:

[0023] For each of the original simulation images, determining an occlusion rate in the original simulation image;

[0024] Part of the original simulation image is eliminated based on the occlusion rate.

[0025] Optionally, selecting a sample simulation image from a plurality of original simulation images includes:

[0026] Get the preset number of sample images;

[0027] With the goal that the selected sample simulation images can cover all real scenarios and that the selected sample simulation images can meet the real scenario distribution information, select sample simulation images that meet the preset number of sample images from multiple original simulation images.

[0028] Optionally, constructing the vehicle's vision sensor in the simulator includes:

[0029] Obtain sensor parameters, where the sensor parameters include the number of sensors, the type of each sensor, and the deployment location of each sensor;

[0030] Based on the sensor parameters, construct the vehicle's vision sensor in the simulator.

[0031] Optionally, the method further includes:

[0032] Construct the vehicle's radar sensor in the simulator;

[0033] During the process of simulating the vehicle's autonomous driving through the simulator, obtain the simulated point cloud data collected by the radar sensor in the simulation environment;

[0034] Determine the visual annotation data according to the real image data and the corresponding ground truth labels, including:

[0035] Determine the real image data, the corresponding simulated point cloud data, and the ground truth labels as the visual annotation data.

[0036] An embodiment of the present application further provides an electronic device, where the electronic device includes:

[0037] A processor and a memory;

[0038] The processor is used to execute the steps of the method for generating visual annotation data provided in any embodiment of the present application by calling the program or instructions stored in the memory.

[0039] An embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores a program or instructions, and the program or instructions cause a computer to execute the steps of the method for generating visual annotation data provided in any embodiment of the present application.

[0040] In summary, the present application proposes a method for generating visual annotation data. In this method, a visual sensor of a vehicle is constructed in a simulator, and the simulator is used to simulate the autonomous driving of the vehicle, obtaining the simulated image data collected by the visual sensor in the simulated environment and the corresponding ground truth labels. Then, the simulated image data is input into the trained simulation transfer model to transfer the simulated image data from the simulated environment to the real environment, obtaining the corresponding real image data. Based on the real image data and the corresponding ground truth labels, the visual annotation data is determined, realizing the automatic annotation of autonomous driving visual data without manual annotation, which can solve problems such as high manual annotation cost and low data utilization rate, improve the efficiency of data annotation, and reduce the cost of data annotation. Moreover, this method first generates simulated image data and then transfers the simulated image data to the real environment, making the simulated data indistinguishable from the real collected data, which can solve problems such as large differences between simulated data and real data and insufficient richness of details in simulated data, so that the simulated data can be used to train a neural network with strong robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 is a flowchart of a method for generating visual annotation data provided by an embodiment of the present application;

[0043] Figure 2 is a schematic diagram of the formation process of visual annotation data provided by an embodiment of the present application;

[0044] Figure 3 is a schematic diagram of the structure of a device for generating visual annotation data provided by an embodiment of the present application;

[0045] Figure 4 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The following will further elaborate on the present application in conjunction with the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention and are not intended to limit the invention. Additionally, it should be noted that for the sake of convenience of description, only the parts related to the invention are shown in the drawings.

[0047] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0048] As mentioned in the background art, in view of the problems in the prior art, the present application proposes a method for generating visual annotation data. Figure 1 is a flowchart of a method for generating visual annotation data provided by an embodiment of the present application. Refer to Figure 1 The method for generating visual annotation data specifically includes:

[0049] S110. Build a visual sensor of the vehicle in the simulator, and simulate the vehicle's autonomous driving through the simulator to obtain the simulated image data collected by the visual sensor in the simulation environment and the corresponding ground truth labels.

[0050] Among them, the simulator can be an autonomous driving simulation tool, such as CARLA. The simulator can be used to simulate the vehicle's autonomous driving environment, including weather, traffic scenarios, and map layouts, etc. Algorithms for autonomous driving vehicles can be tested and developed in the simulator. The simulator can support the simulation of multiple sensors, such as cameras, lidar, etc., so that the developed algorithms can obtain environmental information just like on a real vehicle.

[0051] Before building the visual sensor of the vehicle in the simulator, the corresponding open source code repository of the simulator can be obtained first, and the corresponding OpenDrivingLab map model can be pulled. Further, according to the requirements of the visual annotation data, such as the camera parameters, radar parameters, etc. required for generating the simulated image data, a visual sensor can be built in the simulator.

[0052] In a specific implementation manner, building the visual sensor of the vehicle in the simulator includes:

[0053] Obtain sensor parameters, where the sensor parameters include the number of sensors, the type of each sensor, and the deployment location of each sensor; based on the sensor parameters, build the visual sensor of the vehicle in the simulator.

[0054] Specifically, the sensor parameters can be obtained first, where the sensor parameters can include the relevant parameters of the visual sensor and the radar sensor. Exemplarily, the sensor parameters can be input by the user in the display interface of the simulator, or the sensor parameters can be determined based on the visual perception scheme selected by the user.

[0055] For example, the user can select a visual perception solution with 6 pinhole cameras plus 1 lidar, or the user can select a visual perception solution with 4 fisheye cameras. Based on the selected solution, the corresponding sensor parameters can be queried, including the number of visual sensors, the type of each visual sensor, the deployment location, etc., and then the deployment of each sensor can be completed in the simulator.

[0056] Through the above embodiments, visual sensors can be constructed in the simulator according to the requirements for visual annotation data, so that the simulated image data collected in the simulator can meet the requirements.

[0057] In addition to deploying visual sensors in the simulator, the ground truth labels output by the simulator can also be specified according to the requirements for visual annotation data, such as object position, object category, object size, etc. For example, the ground truth labels can include 2D ground truth annotations of vehicles, 2D ground truth annotations of pedestrians, 2D ground truth annotations of obstacles, lane line annotations, parking space frame annotations, etc. If the sensors constructed in the simulator include radar sensors, the ground truth labels can also include 3D ground truth annotations of vehicles, 3D ground truth annotations of pedestrians, 3D ground truth annotations of obstacles, etc.

[0058] Furthermore, the simulator can be run to simulate vehicle autonomous driving, and simulated image data and corresponding ground truth labels can be collected in real time. It should be noted that the number of collected simulated image data can be multiple, and according to the requirements for visual annotation data, after a batch of simulated image data is collected, the simulation environment in the simulator (such as weather, road surface or light) can be adjusted, and then the simulated image data and corresponding ground truth labels in the new simulation environment can be collected continuously. Exemplarily, the simulation environment in the simulator can be adjusted by modifying the environment configuration file.

[0059] S120: Input the simulated image data into the trained simulation transfer model to transfer the simulated image data from the simulation environment to the real environment, and obtain the corresponding real image data.

[0060] Specifically, after obtaining the simulated image data based on the simulator, it can be input into the trained simulation transfer model. The simulation transfer model can be used to transfer the input image data from the simulation environment to the real environment, and supplement the details of the real environment for the input image data, so that the input image data is closer to the data collected in reality.

[0061] In the embodiments of the present application, the simulation migration model can be obtained by updating and training a trained style migration model based on a sample data set. The sample data set can include multiple sample real images and multiple sample simulation images. The number of sample real images and the number of sample simulation images can be the same or different. Fine-tuning and training the style migration model based on the sample data set can enable the model to learn the deep features of the simulation environment and the real environment, and complete the mapping from the simulation environment to the real environment.

[0062] In a specific embodiment, the training process of the simulation migration model includes the following steps:

[0063] Step 11: Collect multiple original simulation images based on the simulator, and select sample simulation images from the multiple original simulation images;

[0064] Step 12: Obtain multiple sample real images, and obtain a pre-trained style migration model;

[0065] Step 13: Train the style migration model based on each sample simulation image and each sample real image to obtain the simulation migration model.

[0066] Among them, in Step 11, the simulator can be run to collect simulation data to obtain multiple original simulation images. Before collecting the original simulation images, the map model can also be pulled and sensors can be built in the simulator.

[0067] To ensure the reliability of the trained simulation migration model, during the process of collecting multiple original simulation images from the simulator, images under different simulation environments can be collected to cover as many simulation environments as possible, ensuring that the model trained subsequently can be applied to the simulation image data under various simulation environments.

[0068] Regarding the above Step 11, in one example, collecting multiple original simulation images based on the simulator includes:

[0069] Obtain the environment configuration file; run the simulator based on the environment configuration file to obtain multiple original simulation images; update the environment configuration file, and return to execute the step of running the simulator based on the environment configuration file.

[0070] Among them, the environment configuration file can be used to configure the simulation environment in the simulator, such as CarlaSettings.ini. Specifically, the environment configuration file can be pulled first, and the simulator can be run based on the pulled environment configuration file to configure the simulation environment in the simulator through this environment configuration file, and collect each original simulation image under this simulation environment in real time.

[0071] After the acquisition of the simulation environment is completed, the environment configuration file can be updated. For example, some parameters in the environment configuration file can be adjusted to update the simulation environment in the simulator, so as to continue to collect the original simulation images in the updated simulation environment. Repeat this process until the original simulation images in multiple simulation environments are collected.

[0072] Among them, the parameters in the environment configuration file include but are not limited to the parameters related to weather, road surface, light, obstacles, etc.

[0073] Optionally, updating the environment configuration file includes: updating at least one of the weather type, road surface condition parameters, light parameters, number of obstacles, type of obstacles, and position of obstacles in the environment configuration file.

[0074] Among them, the road surface condition parameters can be used to describe the road traffic types in the simulation environment, such as urban roads, elevated roads, crowded sections, intersections, etc., and can also be used to describe the road surface types in the simulation environment, such as asphalt roads, cement roads, gravel roads, etc. The light parameters can be used to describe the intensity of light in the simulation environment.

[0075] Among them, the number of obstacles can describe the number of objects other than vehicles in the simulation environment, including but not limited to pedestrians, other vehicles, roadblocks, etc. The type of obstacles can describe the types of objects other than vehicles in the simulation environment, such as pedestrians, other vehicles, roadblocks, etc. The position of obstacles can describe the position of obstacles relative to the vehicle in the simulation environment, or the position in the road network.

[0076] Specifically, the user can update the environment configuration file by changing at least one of the weather type, road surface condition parameters, light parameters, number of obstacles, type of obstacles, and position of obstacles, so as to facilitate the subsequent collection of the original simulation images according to the updated environment configuration file.

[0077] Through the above optional implementation manners, the change of the simulation environment can be realized by changing the weather type, road surface condition, light, number of obstacles, type of obstacles or position of obstacles in the simulation environment, and the original simulation images in different weather, road surface, light and other scenarios can be collected, ensuring the comprehensiveness of the original simulation images and enabling the original simulation images to cover all real scenarios as much as possible.

[0078] Specifically, in step 11 above, after all the original simulation images are collected, considering that the number of the collected original simulation images is large, in order to ensure the efficiency of model training, some original simulation images can be selected from all the original simulation images as the sample simulation images.

[0079] In the embodiments of the present application, during the process of selecting sample simulation images from the original simulation images, in order to make the selected sample simulation images conform to the real scene distribution as much as possible, the original simulation images can also be screened in combination with the real scene distribution.

[0080] For step 11 above, in one example, before selecting sample simulation images from multiple original simulation images, it further includes:

[0081] Step 110: Obtain real scene distribution information, where the real scene distribution information includes the proportion of images in each real scene in the real dataset;

[0082] Step 111: Based on the real scene distribution information and the simulation scenes corresponding to each original simulation image, eliminate some original simulation images.

[0083] Specifically, in step 110, the real dataset can be obtained first, and then the proportion of images in each real scene included in the real data can be determined.

[0084] Furthermore, in step 111, according to the real scene distribution information, sample balancing can be performed on all original simulation images, and some of the original simulation images can be eliminated, so that the image proportion of the simulation scenes of the remaining original simulation images after elimination conforms to the real scene distribution information.

[0085] Exemplarily, if the number of images in the rainy day scene in the real dataset is small, the original simulation images in the rainy day scene can be reduced.

[0086] Through the above steps 110-step 111, the balanced processing of the original simulation images can be realized, and the remaining original simulation images can meet the real scene distribution information, which is convenient for the subsequent selected sample simulation images to also conform to the real scene distribution information, thereby ensuring the robustness of the trained model and enabling the model to ensure the migration effect in various scenes.

[0087] In the embodiments of the present application, considering that there may be objects with too severe occlusion in some original simulation images, and such images have little effect on the model training, therefore, before eliminating some original simulation images based on the real scene distribution information, the original simulation images with too severe occlusion can also be eliminated from all original simulation images first.

[0088] Optionally, before obtaining the real scene distribution information, it further includes:

[0089] For each original simulation image, determine the occlusion rate in the original simulation image; eliminate some original simulation images based on the occlusion rate.

[0090] Among them, for each original simulation image, the original simulation image can be input into the object detection model first to obtain the object detection bounding box of the original simulation image, and the occluded area in the object detection bounding box can be determined. Then, based on the size of the occluded area and the size of the object detection bounding box, the occlusion rate in the original simulation image can be calculated. For example, the occlusion rate can be calculated according to the number of pixels in the occluded area and the number of pixels in the object detection bounding box.

[0091] Furthermore, the occlusion rate of each original simulation image can be compared with a preset occlusion rate threshold, and the original simulation images with an occlusion rate greater than the preset occlusion rate threshold can be removed. Based on this method, the original simulation images with too severe occlusion can be removed to avoid being selected for model training later, ensuring the accuracy of model training.

[0092] After filtering out the images with a severe occlusion rate from the original simulation images and equalizing the original simulation images, sample simulation images can be selected from all the original simulation images. During the process of selecting sample simulation images, the random sampling method can be used. To ensure the reliability of the subsequent simulation migration model, attention can also be paid to the diversity of the sample simulation images during the process of selecting sample simulation images, and more scenarios should be covered as much as possible.

[0093] Regarding step 11 above, in one example, selecting sample simulation images from multiple original simulation images includes:

[0094] Obtain the preset number of sample images;

[0095] With the goal that the selected sample simulation images can cover all real scenarios and the selected sample simulation images can meet the real scenario distribution information, select the sample simulation images that meet the preset number of sample images from multiple original simulation images.

[0096] Among them, the preset number of sample images can be the number of simulation images pre-set for training the simulation migration model. For example, 1000 images.

[0097] Specifically, according to the preset number of sample images, sample simulation images that meet the preset number of sample images can be selected from the original simulation images, and during the selection process, it is ensured that the selected sample simulation images can meet the real scenario distribution information, and all the selected sample simulation images can cover all real scenarios.

[0098] Exemplarily, scenarios such as day, night, rainy day, foggy day, etc. need to be covered, and scenarios such as urban roads, elevated roads, crowded sections, intersections, etc. also need to be covered, and, scenarios such as other vehicles driving in the oncoming direction and other vehicles driving in the adjacent lane also need to be covered.

[0099] Through the above examples, the selected sample simulation images can meet the preset number of sample images, and the selected sample simulation images can cover all real scenarios and meet the real scenario distribution. Furthermore, the simulation transfer model trained based on the sample simulation images is more robust, ensuring the image transfer effect in various scenarios.

[0100] After selecting the sample simulation images, further, in step 12, multiple sample real images can be obtained. The sample real images are images collected in real scenarios. For example, sample real images can be selected from open-source datasets for autonomous driving (such as nuscenes, waymo, etc.) or existing real datasets.

[0101] It should be noted that in the process of selecting the sample real images, the sample real images can also be obtained with the goal that the selected sample real images can cover all real scenarios and the selected sample real images can meet the real scenario distribution information.

[0102] In addition to obtaining the sample real images and the sample simulation images, a pre-trained style transfer model also needs to be obtained. Among them, the style transfer model can be a pre-trained model for image style transfer, used for style transfer between two scenarios, such as, transferring from day to night, from sunny to rainy, etc.

[0103] Exemplarily, the style transfer model can adopt img2img-turbo. The style transfer model can be trained based on a large model for image generation (such as Stable Diffusion).

[0104] Further, in step 13, the style transfer model can be fine-tuned (which can be understood as optimized training) through each sample simulation image and each sample real image, so that the style transfer model can learn the deep features in the simulation environment and the real environment, update the network parameters in the style transfer model, and obtain the simulation transfer model.

[0105] Through the above steps 11 - 13, sample simulation images can be obtained through the simulator. Furthermore, by combining the sample real images and the sample simulation images and using the style transfer model for fine-tuning, a simulation transfer model that can achieve the transfer from the simulation environment to the real environment can be obtained. There is no need to reconstruct the model framework and train the model from scratch. By using the model framework of the style transfer model and combining the purpose of transferring from the simulation environment to the real environment, fine-tuning is performed on it. While ensuring the model output accuracy of the simulation transfer model, the model training efficiency of the simulation transfer model is greatly improved.

[0106] In an embodiment of the present application, after the simulation transfer model is trained, for the use of the simulation transfer model, simulation image data can be input into the trained simulation transfer model to obtain real image data corresponding to the simulation image data output by the model.

[0107] S130. Determine visual annotation data according to the real image data and the corresponding ground truth label.

[0108] Specifically, after obtaining the real image data through the simulation transfer model, visual annotation data can be formed based on the real image data and the corresponding ground truth label (i.e., the ground truth label corresponding to the simulation image data output by the simulator).

[0109] In an embodiment of the present application, considering the annotation of autonomous driving data, there may be a need to output 3D point cloud data together for the training of other models. Therefore, 3D point cloud data can also be collected and output by the simulator.

[0110] In a specific implementation manner, the method provided in the embodiment of the present application further includes: constructing a radar sensor of the vehicle in the simulator; during the process of simulating the autonomous driving of the vehicle by the simulator, obtaining the simulated point cloud data collected by the radar sensor in the simulation environment;

[0111] Determining visual annotation data according to the real image data and the corresponding ground truth label includes: determining the real image data, the corresponding simulated point cloud data, and the ground truth label as visual annotation data.

[0112] Specifically, a radar sensor can be constructed while constructing a visual sensor in the simulator, and then during the process of simulating the autonomous driving of the vehicle by the simulator, the simulated point cloud data corresponding to each simulation image data is collected by the radar sensor.

[0113] Further, after converting the simulation image data into real image data, the real image data, the corresponding simulated point cloud data, and the ground truth label can be output together as visual annotation data for the training of other models.

[0114] Through the above implementation manner, the simulated point cloud data can be output together while outputting the real image data and the ground truth label, so as to facilitate the training of 2D and 3D application-related models.

[0115] Figure 2 is a schematic diagram of the formation process of visual annotation data provided by an embodiment of the present application, as Figure 2As shown, the simulator can first generate simulated image data and ground truth labels. The simulated image data can be a surround-view camera image or a fisheye camera image. The ground truth labels can include 2D object annotations, 3D object annotations, lane line annotations, and parking space frame annotations. The simulated image data can be migrated from the simulated environment to the real environment through a simulated migration model to obtain real image data. The real image data can be a surround-view camera image or a fisheye camera image. Finally, the ground truth labels and the real image data can be used as annotation data for training an autonomous driving perception model.

[0116] The method for generating visual annotation data provided in an embodiment of the present application constructs a visual sensor of a vehicle in a simulator and simulates the autonomous driving of the vehicle through the simulator to obtain simulated image data collected by the visual sensor in the simulated environment and corresponding ground truth labels. Then, the simulated image data is input into a trained simulated migration model to migrate the simulated image data from the simulated environment to the real environment to obtain corresponding real image data. According to the real image data and the corresponding ground truth labels, visual annotation data is determined, realizing automatic annotation of autonomous driving visual data without manual annotation, which can solve problems such as high manual annotation cost and low data utilization rate, improve the efficiency of data annotation, and reduce the cost of data annotation. Moreover, this method first generates simulated image data and then migrates the simulated image data to the real environment, making the simulated data indistinguishable from the real collected data, which can solve problems such as large differences between simulated data and real data and insufficient details in simulated data, so that the simulated data can be used to train a neural network with strong robustness.

[0117] Figure 3 It is a schematic structural diagram of a device for generating visual annotation data provided in an embodiment of the present application. The device for generating visual annotation data includes a simulation module 310, a migration module 320, and an output module 330, where:

[0118] The simulation module 310 is configured to construct a visual sensor of a vehicle in a simulator and simulate the autonomous driving of the vehicle through the simulator to obtain simulated image data collected by the visual sensor in the simulated environment and corresponding ground truth labels;

[0119] The migration module 320 is configured to input the simulated image data into a trained simulated migration model to migrate the simulated image data from the simulated environment to the real environment to obtain corresponding real image data;

[0120] The output module 330 is configured to determine visual annotation data according to the real image data and the corresponding ground truth labels.

[0121] On the basis of the above-mentioned embodiments, optionally, the device further includes a model training module, wherein the model training module is used to collect a plurality of original simulated images based on the simulator, and select a sample simulated image from the plurality of original simulated images; obtain a plurality of sample real images, and obtain a pre-trained style transfer model; and train the style transfer model based on each sample simulated image and each sample real image to obtain the simulation transfer model.

[0122] Based on the above embodiments, optionally, the model training module is also used to obtain an environment configuration file, wherein the environment configuration file is used to configure the simulation environment in the simulator; run the simulator based on the environment configuration file to obtain multiple original simulation images; update the environment configuration file, and return to execute the step of running the simulator based on the environment configuration file.

[0123] Based on the above implementation modes, optionally, the model training module is also used to update at least one of the weather type, road condition parameters, light parameters, obstacle quantity, obstacle type and obstacle position of the environmental configuration file.

[0124] On the basis of the above-mentioned implementation modes, optionally, the model training module is also used to obtain real scene distribution information, wherein the real scene distribution information includes the proportion of images in each real scene in the real data set; based on the real scene distribution information and the simulation scene corresponding to each of the original simulation images, some of the original simulation images are eliminated.

[0125] On the basis of the above-mentioned implementation modes, optionally, the model training module is further used to determine the occlusion rate in the original simulation image for each of the original simulation images; and to eliminate part of the original simulation images based on the occlusion rate.

[0126] On the basis of the above-mentioned implementation modes, optionally, the model training module is also used to obtain a preset number of sample images; with the goal that the selected sample simulation images can cover all real scenes and the selected sample simulation images can satisfy the real scene distribution information, sample simulation images that meet the preset number of sample images are selected from multiple original simulation images.

[0127] On the basis of the above-mentioned embodiments, optionally, the simulation module 310 is also used to obtain sensor parameters, which include the number of sensors, the type of each sensor, and the deployment location of each sensor; based on the sensor parameters, the vehicle's visual sensor is constructed in the simulator.

[0128] Based on the above embodiments, optionally, the simulation module 310 is further configured to construct a radar sensor of the vehicle in the simulator; during the simulation of the vehicle's autonomous driving through the simulator, obtain the simulated point cloud data collected by the radar sensor in the simulation environment; the output module 330 is further configured to determine the real image data, the corresponding simulated point cloud data, and the ground truth label as visual annotation data.

[0129] The visual annotation data generation device provided in the embodiments of the present application can execute the steps in the visual annotation data generation method provided in the method embodiments of the present application, and the implementation steps and beneficial effects are not described herein again.

[0130] Figure 4 It is a schematic structural diagram of an electronic device provided in the embodiments of the present application. As Figure 4 shown, the electronic device 400 includes one or more processors 401 and a memory 402.

[0131] The processor 401 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 400 to perform desired functions.

[0132] The memory 402 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 401 may run the program instructions to implement the visual annotation data generation method of any embodiment of the present application described above and / or other desired functions. Various contents such as initial extrinsic parameters and thresholds may also be stored in the computer-readable storage medium.

[0133] In one example, the electronic device 400 may further include: an input device 403 and an output device 404, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). The input device 403 may include, for example, a keyboard, a mouse, etc. The output device 404 may output various information to the outside, including warning prompt information, braking force, etc. The output device 404 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0134] Of course, for simplicity, Figure 4Only some of the components related to this application in the electronic device 400 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 400 may further include any other appropriate components.

[0135] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions that, when run by a processor, cause the processor to execute the steps of the method for generating visual annotation data provided in any embodiment of the present application.

[0136] The computer program product can be written in any combination of one or more programming languages for the program code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0137] In addition, an embodiment of the present application may also be a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions, when run by a processor, cause the processor to execute the steps of the method for generating visual annotation data provided in any embodiment of the present application.

[0138] The computer-readable storage medium may adopt any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but not be limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0139] It should be noted that the terms used in this application are only for describing specific embodiments and do not limit the scope of this application. As shown in the specification and claims of this application, unless the context clearly indicates otherwise, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method or device including the said element.

[0140] It should also be noted that the orientation or positional relationship indicated by terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to this application. Unless otherwise clearly specified and defined, terms such as "installed", "connected", "joined" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0141] Specific examples are used in this article to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only for helping to understand the method and its core idea of this application. The above is only the preferred implementation manner of this application. It should be noted that due to the limited nature of written expression and objectively infinite specific structures, for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements, refinements or changes can be made, or the above technical features can be combined in an appropriate manner; these improvements, refinements, changes or combinations, or directly applying the concept and technical solution of the invention to other occasions without improvement, should all be regarded as the protection scope of this application.

Claims

1. A method for generating visual annotation data, characterized in that: include: Constructing a visual sensor of the vehicle in a simulator, and simulating the automatic driving of the vehicle through the simulator, to obtain simulated image data collected by the visual sensor in the simulation environment and corresponding true value labels; Inputting the simulated image data into a trained simulation migration model to migrate the simulated image data from the simulated environment to a real environment to obtain corresponding real image data; Visual annotation data is determined according to the real image data and the corresponding true value label.

2. The method according to claim 1, characterized in that The training process of the simulation migration model includes: Collecting a plurality of original simulation images based on the simulator, and selecting a sample simulation image from the plurality of original simulation images; Get multiple sample real images and get a pre-trained style transfer model; Based on each sample simulated image and each sample real image, the style transfer model is trained to obtain the simulated transfer model.

3. The method according to claim 2, characterized in that A plurality of original simulation images are collected based on the simulator, including: Obtaining an environment configuration file, wherein the environment configuration file is used to configure a simulation environment in the simulator; Running the simulator based on the environment configuration file to obtain a plurality of original simulation images; The environment configuration file is updated, and the step of running the simulator based on the environment configuration file is returned to be executed.

4. The method according to claim 3, characterized in that Update the environment configuration file, including: At least one of the weather type, road condition parameter, light parameter, obstacle quantity, obstacle type and obstacle position of the environment configuration file is updated.

5. The method according to claim 2, characterized in that: Before selecting a sample simulated image from a plurality of original simulated images, the method further includes: Acquire real scene distribution information, wherein the real scene distribution information includes the proportion of images in each real scene in the real data set; Based on the real scene distribution information and the simulation scene corresponding to each of the original simulation images, some of the original simulation images are eliminated.

6. The method according to claim 5, characterized in that Before obtaining the real scene distribution information, it also includes: For each of the original simulation images, determining an occlusion rate in the original simulation image; Part of the original simulation image is eliminated based on the occlusion rate.

7. The method according to claim 5, characterized in that The step of selecting a sample simulation image from a plurality of original simulation images comprises: Get the preset number of sample images; With the goal that the selected sample simulation images can cover all real scenes and that the selected sample simulation images can satisfy the real scene distribution information, sample simulation images that meet the preset number of sample images are selected from multiple original simulation images.

8. The method according to claim 1, characterized in that The visual sensor of the vehicle is constructed in the simulator, including: Acquire sensor parameters, where the sensor parameters include the number of sensors, the type of each sensor, and the deployment location of each sensor; Based on the sensor parameters, a vision sensor of the vehicle is constructed in the simulator.

9. The method according to claim 1, characterized in that: The method further comprises: Build the vehicle's radar sensor in the simulator; In the process of simulating the automatic driving of the vehicle by the simulator, the simulated point cloud data collected by the radar sensor in the simulated environment is obtained; Determining visual annotation data according to the real image data and the corresponding true value label includes: The real image data, and the corresponding simulated point cloud data and true value labels are determined as visual annotation data.

10. An electronic device, characterized in that: The electronic device comprises: Processor and memory; The processor is used to execute the steps of the method for generating visual annotation data according to any one of claims 1 to 9 by calling the program or instruction stored in the memory.