Data annotation method, training method of automatic driving model and related equipment
By constructing and fusing the occupancy grid of sensor data, the parallax problem caused by the installation position difference between high-performance sensors and mass-produced vehicle sensors is solved, and the accuracy and efficiency of automatic labeling are improved, which is suitable for data annotation and training of autonomous driving models.
Patent Information
- Application Number
- CN202510168270.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the difference in the installation position between the high-performance sensor and the mass-produced vehicle's own sensor leads to parallax problems, affecting the accuracy of the automatic labeling data.
By obtaining the initial sensor data of the vehicle, an initial occupancy grid is built and processed to generate the target occupancy grid, fuse the observation status grid of different sensors, generate target labeling data, and solve the parallax problem.
It improves the accuracy and efficiency of automatic labeling, and can quickly handle data collection and labeling of rare driving scenarios.
Smart Images

Figure CN120256946A_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed in the present application relate to the field of autonomous driving technology, and more specifically, to a data annotation method, a training method for an autonomous driving model, and related devices. Background Art
[0002] In recent years, with the continuous development of new energy vehicles and the continuous expansion of intelligent driving scenarios, the training data required for perception models (such as, environmental perception, object detection, etc.) relied on by autonomous driving has increased exponentially. How to perform automatic annotation stably and efficiently in various driving scenarios to obtain a large amount of ground truth data is an urgent problem to be solved currently. Summary of the Invention
[0003] According to the embodiments of the present application, the present application provides a data annotation method, a training method for an autonomous driving model, and related devices to solve the above problems.
[0004] The first aspect of the present application discloses a data annotation method, including: obtaining initial sensor data of a vehicle, where the vehicle includes a first sensor and at least two second sensors, and the initial sensor data includes first sensor data corresponding to the first sensor and at least two second sensor data corresponding to the at least two second sensors; constructing an initial occupancy grid based on the initial sensor data; processing the initial occupancy grid to generate a target occupancy grid corresponding to the at least two second sensor data; and processing the target occupancy grid to generate target annotation data.
[0005] In some embodiments, the processing the initial occupancy grid to generate a target occupancy grid corresponding to the at least two second sensor data includes: generating at least two observation state grids based on the initial occupancy grid, where the at least two observation state grids correspond to the at least two second sensors; and fusing the at least two observation state grids to obtain the target occupancy grid.
[0006] In some embodiments, the generating at least two observation state grids based on the initial occupancy grid includes: determining position information of the at least two second sensors in the initial occupancy grid based on calibration information of the at least two second sensors; and calculating the initial occupancy grid based on the position information to generate the at least two observation state grids.
[0007] In some embodiments, calculating the initial occupancy grid based on the position information to generate the at least two observed state grids includes: deleting the occluded voxels in the observed state grid corresponding to the second sensor based on the field of view angle of the second sensor and the voxel occlusion relationship in the initial occupancy grid, to obtain the observed state grid.
[0008] In some embodiments, obtaining the initial sensor data of the vehicle includes: collecting first collection data by using the first sensor; collecting at least two second collection data by using the at least two second sensors; and synchronizing the first collection data with the at least two second collection data to obtain the initial sensor data.
[0009] In some embodiments, constructing the initial occupancy grid based on the initial sensor data includes: inputting the initial sensor data into a multi-modal occupancy grid model to generate the initial occupancy grid, where the initial occupancy grid includes voxel information of surrounding obstacles of the vehicle.
[0010] In some embodiments, processing the target occupancy grid to generate target annotation data includes: inputting the target occupancy grid into a target detection model to generate a target detection result; and post-processing the target detection result to obtain the target annotation data, where the post-processing includes at least one of trajectory tracking, motion compensation, and parameter tuning.
[0011] A second aspect of the present application discloses a method for training an autonomous driving model, including: obtaining annotation data; and training the autonomous driving model by using the annotation data; where the annotation data is obtained based on the data annotation method in the first aspect.
[0012] A third aspect of the present application discloses an electronic device, including a memory and a processor coupled to each other, where the processor is configured to execute program instructions stored in the memory to implement the data annotation method in the first aspect, or to implement the method for training the autonomous driving model in the second aspect.
[0013] A fourth aspect of the present application discloses a non-volatile computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, the data annotation method in the first aspect is implemented, or the method for training the autonomous driving model in the second aspect is implemented.
[0014] The beneficial effects of the present application are as follows: obtaining the initial sensor data of the vehicle, where the initial sensor data includes the first sensor data corresponding to the first sensor and at least two second sensor data corresponding to at least two second sensors, constructing an initial occupancy grid based on the initial sensor data, and further, fusing the initial occupancy grid to generate a target occupancy grid corresponding to at least two second sensor data, and obtaining target annotation data through processing the target occupancy grid, thereby improving the accuracy and efficiency of automatic annotation. Description of the Drawings
[0015] The present application will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0016] Figure 1 is a flowchart of the data annotation method according to an embodiment of the present application;
[0017] Figure 2 is a schematic diagram of the initial occupancy grid according to an embodiment of the present application;
[0018] Figure 3 is a schematic diagram of the generation of the observation state grid according to an embodiment of the present application;
[0019] Figure 4 is a schematic diagram of the generation of the observation state grid according to another embodiment of the present application;
[0020] Figure 5 is a schematic diagram of the target occupancy grid according to an embodiment of the present application;
[0021] Figure 6 is a flowchart of the training method of the automatic driving model according to an embodiment of the present application;
[0022] Figure 7 is a schematic diagram of the structure of the electronic device according to an embodiment of the present application;
[0023] Figure 8 is a schematic diagram of the structure of the non-volatile computer-readable storage medium according to an embodiment of the present application. Detailed Embodiments
[0024] Referring to "embodiments" in the present application means that specific features, structures, or characteristics described in connection with the embodiments can be included in at least two embodiments of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0025] In this application, the term "and / or" is merely a description of the relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, in this text, the character " / " generally indicates that the associated objects before and after are in an "or" relationship. Furthermore, the term "multiple" in this text means two or more than two. Additionally, the term "at least one" in this text means any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set composed of A, B, and C. Additionally, the terms "first", "second", and "third" in this application are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features.
[0026] Currently, other high-performance sensors are usually added to the original configuration of mass-produced vehicles to enhance the environmental perception ability of data acquisition vehicles and provide ground truth data for the mass production model solution of autonomous driving. In this process, the installation positions of high-performance sensors are different from those of the sensors of the mass-produced vehicles themselves, resulting in a parallax problem between different sensors, which affects the accuracy of automatically labeled data.
[0027] To this end, this application proposes a data annotation method, a training method for an autonomous driving model, and related devices to solve the above problems.
[0028] To enable those skilled in the art to better understand the technical solutions of this application, the technical solutions of this application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0029] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the data annotation method according to an embodiment of this application. The execution subject of this method can be an electronic device with computing functions, such as a microcomputer, a server, and mobile devices such as a laptop computer and a tablet computer.
[0030] It should be noted that if there are substantially the same results, the method of this application is not limited to Figure 1 the process sequence shown.
[0031] In some possible implementation manners, this method can be implemented by a processor invoking computer-readable instructions stored in a memory. As Figure 1 shown, this method can include the following steps:
[0032] S11: Obtain the initial sensor data of the vehicle. The vehicle includes a first sensor and at least two second sensors. The initial sensor data includes the first sensor data corresponding to the first sensor and the at least two second sensor data corresponding to the at least two second sensors.
[0033] The first sensor can be at least one or more of a high-performance surround-view lidar, a side-view blind-spot-supplementing high-precision radar, etc. The second sensors can be sensors such as ultrasonic radars and cameras on mass-produced vehicles. The first sensor and at least two second sensors are installed on the vehicle. For example, the first sensor can be installed on a mass-produced vehicle to serve as a data acquisition vehicle.
[0034] Obtain the initial sensor data of the vehicle. The initial sensor data includes the first sensor data corresponding to the first sensor and the at least two second sensor data corresponding to the at least two second sensors. For example, the first sensor data can be high-precision data obtained by the first sensor, and the second sensor data can be conventional-precision data obtained by the second sensors.
[0035] S12: Construct an initial occupancy grid based on the initial sensor data.
[0036] Construct an initial occupancy grid based on the initial sensor data. For example, use the initial sensor data to construct an occupancy grid of the surrounding environment where the vehicle is located. The initial occupancy grid can be used to represent obstacle information, etc., of the surrounding environment of the vehicle jointly determined by the first sensor data and the at least two second sensor data.
[0037] S13: Process the initial occupancy grid to generate target occupancy grids corresponding to the at least two second sensor data.
[0038] Process the initial occupancy grid to generate target occupancy grids corresponding to the at least two second sensor data. For example, determine at least two observation state grids corresponding to the at least two second sensor data from the initial occupancy grid, and fuse the at least two observation state grids to obtain a complete target occupancy grid corresponding to the at least two second sensor data.
[0039] S14: Process the target occupancy grid to generate target annotation data.
[0040] Perform post-processing and optimization on the target occupancy grid to generate target annotation data. The target annotation data can be used directly as ground truth data for training relevant models of mass-produced vehicles.
[0041] In this embodiment, the initial sensor data of the vehicle is obtained. The initial sensor data includes the first sensor data corresponding to the first sensor and at least two second sensor data corresponding to at least two second sensors. An initial occupancy grid is constructed based on the initial sensor data. Further, the initial occupancy grid is fused to generate a target occupancy grid corresponding to at least two second sensor data. The target annotation data is obtained through processing the target occupancy grid, thereby improving the accuracy and efficiency of automatic annotation.
[0042] In some embodiments, obtaining the initial sensor data of the vehicle includes: collecting with the first sensor to obtain first collection data; collecting with at least two second sensors to obtain at least two second collection data; synchronously matching the first collection data and the at least two second collection data to obtain the initial sensor data.
[0043] A first sensor and at least two second sensors are installed on the vehicle. Collecting with the first sensor to obtain first collection data. For example, the first collection data is the sequential point cloud data S1 collected by a high-performance lidar. Collecting with at least two second sensors to obtain at least two second collection data. For example, the second collection data can be the sequential data S2 collected by the sensors of a production vehicle. Synchronously matching the first collection data and the at least two second collection data to obtain the initial sensor data. For example, according to the principle of the closest timestamp matching, the S1 data and the S2 data are synchronously matched at the frame level, and the data frame after time synchronization is used as the initial sensor data.
[0044] In some embodiments, constructing the initial occupancy grid based on the initial sensor data includes: inputting the initial sensor data into a multimodal occupancy grid model to generate the initial occupancy grid, where the initial occupancy grid includes the voxel information of the surrounding obstacles of the vehicle.
[0045] Input the initial sensor data into the Multimodal Occupancy Grid Model, and use the multimodal occupancy grid model to construct and update the 3D space occupancy information around the vehicle, that is, generate the initial occupancy grid. Among them, the multimodal occupancy grid model includes Co-Occ (Co-Occupancy Model), Fb-occ (Feature-based Occupancy Model), Open-occupancy (Open-world Occupancy Model), etc. The initial occupancy grid includes the voxel information of the surrounding obstacles of the vehicle, such as the occupancy states of static 3D targets and dynamic 3D targets around the vehicle.
[0046] In some examples, a vehicle is equipped with sensor P1, sensor P2, and sensor P3. Among them, sensor P1 can be the first sensor, and sensor P2 and sensor P3 can be the second sensors. Initial sensor data of the vehicle is obtained. For example, the time-series data collected by sensor P1, sensor P2, and sensor P3 are synchronized to obtain the initial sensor data. The obtained initial sensor data is input into a multi-modal occupancy grid model, and then an initial occupancy grid M(a) can be obtained, as Figure 2 shown Figure 2 is a schematic diagram of the initial occupancy grid of an embodiment of the present application. In Figure 2 it, the entire occupancy grid (for example, a large cube) is composed of multiple voxels (for example, small cubes). The observation state of each voxel in the occupancy grid can be represented by color. Colored ones indicate occupancy (voxel visible), and non-colored ones indicate free (voxel invisible).
[0047] In some embodiments, the initial occupancy grid is processed to generate target occupancy grids corresponding to at least two second sensor data, including: generating at least two observation state grids based on the initial occupancy grid, where the at least two observation state grids correspond to the at least two second sensors; fusing the at least two observation state grids to obtain the target occupancy grid.
[0048] Generating at least two observation state grids based on the initial occupancy grid, where the at least two observation state grids correspond to the at least two second sensors. For example, based on the above initial occupancy grid M(a), two observation state grids can be generated, and the two observation state grids correspond to sensor P2 and sensor P3 respectively. Further, the at least two observation state grids are fused to obtain the target occupancy grid. For example, the two observation state grids corresponding to sensor P2 and sensor P3 are merged to obtain a merged occupancy grid, that is, the target occupancy grid.
[0049] In some embodiments, generating at least two observation state grids based on the initial occupancy grid includes: determining the position information of the at least two second sensors in the initial occupancy grid based on the calibration information of the at least two second sensors; calculating the initial occupancy grid based on the position information to generate at least two observation state grids.
[0050] In some examples, at least two observation state grids are generated based on the initial occupancy grid, and the at least two observation state grids correspond to the at least two second sensors. Among them, the position information of the at least two second sensors in the initial occupancy grid is determined based on the calibration information of the at least two second sensors, as Figure 2As shown, the positions of sensors P2 and P3, as well as the position of sensor P1, can be determined in the initial occupancy grid M(a). The calibration information of the sensors can be calculated based on the calibration information from the inherent coordinate systems of the respective sensors to the vehicle rear axle coordinate system. The position information of the sensors includes the positions and orientations of the respective sensors in the occupancy grid, etc. Based on the position information, the initial occupancy grid is calculated to generate at least two observed state grids. For example, according to the positions and orientations of at least two second sensors in the initial occupancy grid, etc., the actual observed states of the at least two second sensors are calculated, and then the corresponding observed state grids are obtained.
[0051] In some embodiments, calculating the initial occupancy grid based on the position information to generate at least two observed state grids includes: deleting the occluded voxels in the observed state grid corresponding to the second sensor based on the field of view angle of the second sensor and the voxel occlusion relationship in the initial occupancy grid to obtain the observed state grid.
[0052] Deleting the occluded voxels in the observed state grid corresponding to the second sensor based on the field of view angle of the second sensor and the voxel occlusion relationship in the initial occupancy grid to obtain the observed state grid. For example, any one of the at least two second sensors is determined, and based on the field of view angle of the second sensor and the voxel occlusion relationship in the initial occupancy grid, the occluded voxels in the observed state grid corresponding to the second sensor are deleted, and then the observed state grid corresponding to the second sensor is obtained.
[0053] Taking the perspective of sensor P2 as an example, as Figure 3 shown, Figure 3 is a schematic diagram of the generation of the observed state grid in an embodiment of the present application. Based on the initial occupancy grid M(a), the occupancy grid M(b) to be revised is determined. There are voxels V1, V2, V3, V4, and V5, etc. in the occupancy grid M(b) to be revised, and the vertex coordinates of voxel V1 are (1, 0, 0). Based on the field of view angle of the second sensor and the voxel occlusion relationship in the initial occupancy grid, the occluded voxels in the observed state grid corresponding to the second sensor are deleted. For example, by calculating based on the principle that light travels in a straight line, starting from the perspective of sensor P2, it can be determined that voxels V2, V3, and V4 are occluded by voxel V1, that is, voxels V2, V3, and V4 are occluded voxels. Therefore, voxels V2, V3, and V4 are deleted from the occupancy grid M(b) to be revised. For example, the states of voxels V2, V3, and V4 are modified to invisible, and then the revised occupancy grid M(c), that is, the observed state grid corresponding to sensor P2, is obtained.
[0054] Taking the perspective of sensor P3 as an example, as Figure 4 shown, Figure 4It is a schematic diagram of generating an observation status grid according to another embodiment of the present application. Based on the initial occupancy grid M(a), the occupancy grid M(d) to be revised is determined. There are voxels V2, V3, V4, V5, V6, etc. in the occupancy grid M(d) to be revised. One vertex coordinate of voxel V6 is (1, 4, 0). Based on the field of view angle of the second sensor and the voxel occlusion relationship in the initial occupancy grid, the occluded voxels in the observation status grid corresponding to the second sensor are deleted. For example, by calculating according to the principle that light travels in a straight line, starting from the perspective of sensor P3, it can be determined that voxels V3, V4, and V5 are occluded by voxel V6, that is, voxels V3, V4, and V5 are occluded voxels. Therefore, voxels V3, V4, and V5 are deleted from the occupancy grid M(b) to be revised. For example, the status of voxels V3, V4, and V5 is modified to invisible, and then the revised occupancy grid M(e) is obtained, that is, the observation status grid corresponding to sensor P3.
[0055] Further, in some examples, at least two observation status grids are fused to obtain a target occupancy grid. For example, the revised occupancy grid M(c) and the revised occupancy grid M(e) are merged to obtain the target occupancy grid M(f) corresponding to sensors P2 and P3, as Figure 5 shown, Figure 5 It is a schematic diagram of the target occupancy grid according to an embodiment of the present application. By post-processing and optimizing the target occupancy grid M(f), the target annotation data corresponding to sensors P2 and P3 can be generated.
[0056] In the embodiment of the present application, based on the calibration information of at least two second sensors, the position information of at least two second sensors in the initial occupancy grid is determined. Based on the position information, the initial occupancy grid is calculated to generate at least two observation status grids. Further, at least two observation status grids are fused to obtain a target occupancy grid, and the target annotation data is obtained by processing the target occupancy grid. Among them, the parallax problem between different sensors is effectively solved, and rapid data collection and annotation for long-tail problems such as rare driving scenarios can be realized.
[0057] In some embodiments, processing the target occupancy grid to generate target annotation data includes: inputting the target occupancy grid into a target detection model to generate a target detection result; post-processing the target detection result to obtain the target annotation data, and the post-processing includes at least one of trajectory tracking, motion compensation, and parameter tuning.
[0058] For example, based on the above-mentioned initial occupancy grid M(a), two observation state grids can be generated, namely the revised occupancy grid M(c) and the revised occupancy grid M(e), and the two grids are merged to obtain the target occupancy grid M(f). Further, the target occupancy grid M(f) is input into the target detection model to generate the target detection result. For example, the target occupancy grid M(f) is input into the target detection network to obtain the 3D target detection result in the BEV (Bird's-Eye-View) view. Among them, the target detection network includes FSDv2 (Full Self-Driving version 2), CenterPoint, etc., and the target detection result includes the 3D Box of the target object under the vehicle's rear axle coordinates. The target annotation data is obtained by post-processing the target detection result using relevant algorithms, and the relevant algorithms include at least one of the trajectory tracking algorithm, the motion compensation algorithm, the GTSAM (Georgia Tech Smoothing And Mapping) tuning algorithm, etc.
[0059] Please refer to Figure 6 , Figure 6 which is a schematic flowchart of the training method of the autonomous driving model according to the embodiments of the present application. This method can be applied to an electronic device with functions such as computing. It should be noted that if there are substantially the same results, the method of the present application is not limited to Figure 6 the process sequence shown.
[0060] In some possible implementation manners, this method can be implemented by the processor calling the computer-readable instructions stored in the memory. As Figure 6 shown, this method can include the following steps:
[0061] S61: Obtain annotation data.
[0062] Obtain the annotation data of the data collection vehicle. Among them, a first sensor and at least two second sensors are installed on the data collection vehicle. The annotation data is obtained based on the above-mentioned data annotation method, that is, obtain the initial sensor data of the vehicle, construct the initial occupancy grid based on the initial sensor data, process the initial occupancy grid to generate the target occupancy grid corresponding to at least two second sensor data, and process the target occupancy grid to generate the annotation data.
[0063] S62: Use the annotation data to train the autonomous driving model.
[0064] Training an autonomous driving model using labeled data. For example, the obtained labeled data is directly used as ground truth data for training the autonomous driving model of a production vehicle. Among them, at least two of the above-mentioned second sensors are installed on the production vehicle, and the autonomous driving model includes, but is not limited to, dynamic and static BEV models, dynamic and static occupancy network models, end-to-end models, etc.
[0065] Those skilled in the art can understand that in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0066] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of the electronic device according to an embodiment of the present application. The electronic device 70 includes a memory 71 and a processor 72 that are coupled to each other. The processor 72 is configured to execute program instructions stored in the memory 71 to implement the steps of the data annotation method embodiment described above, or to implement the steps of the autonomous driving model training method embodiment described above. In a specific implementation scenario, the electronic device 70 may include, but is not limited to, a microcomputer, a server, which is not limited herein.
[0067] Specifically, the processor 72 is configured to control itself and the memory 71 to implement the steps of the data annotation method embodiment described above, or to implement the steps of the autonomous driving model training method embodiment described above. The processor 72 may also be referred to as a CPU (Central Processing Unit), and the processor 72 may be an integrated circuit chip with signal processing capabilities. The processor 72 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 72 may be implemented jointly by integrated circuit chips.
[0068] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of the non-volatile computer-readable storage medium according to an embodiment of the present application. The non-volatile computer-readable storage medium 80 is used to store a computer program 801. When the computer program 801 is executed by a processor, for example, by the above-mentioned Figure 7When the processor 72 in the embodiment is executed, it is used to implement the steps of the above-mentioned embodiment of the data annotation method, or to implement the steps of the above-mentioned embodiment of the training method of the autonomous driving model.
[0069] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. For the same or similar parts, reference can be made to each other. For the sake of brevity, they will not be elaborated herein.
[0070] In several embodiments provided in the present application, it should be understood that the disclosed methods and related devices can be implemented in other ways. For example, the above-described embodiments of related devices are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication disconnection between each other can be through some interfaces. The indirect coupling or communication disconnection of devices or units can be in electrical, mechanical or other forms.
[0071] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0072] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes.
[0073] Those skilled in the art will readily appreciate that many modifications and variations can be made to the apparatus and methods while maintaining the teachings of the present application. Therefore, the above disclosure should be considered to be limited only by the scope of the appended claims.
Claims
1. A data annotation method, characterized in that, Including: Obtain initial sensor data of a vehicle, where the vehicle includes a first sensor and at least two second sensors, and the initial sensor data includes first sensor data corresponding to the first sensor and at least two second sensor data corresponding to the at least two second sensors; Construct an initial occupancy grid based on the initial sensor data; Process the initial occupancy grid to generate a target occupancy grid corresponding to the at least two second sensor data; Process the target occupancy grid to generate target annotation data.
2. The method according to claim 1, characterized in that, The processing the initial occupancy grid to generate a target occupancy grid corresponding to the at least two second sensor data includes: Generate at least two observation state grids based on the initial occupancy grid, where the at least two observation state grids correspond to the at least two second sensors; Fuse the at least two observation state grids to obtain the target occupancy grid.
3. The method according to claim 2, wherein The generating at least two observation state grids based on the initial occupancy grid includes: Determine position information of the at least two second sensors in the initial occupancy grid based on calibration information of the at least two second sensors; Calculate the initial occupancy grid based on the position information to generate the at least two observation state grids.
4. The method according to claim 3, wherein The calculating the initial occupancy grid based on the position information to generate the at least two observation state grids includes: Delete occluded voxels in the observation state grid corresponding to the second sensor based on the field of view angle of the second sensor and the voxel occlusion relationship in the initial occupancy grid to obtain the observation state grid.
5. The method according to claim 1, wherein The obtaining initial sensor data of a vehicle includes: Collect using the first sensor to obtain first collection data; Collect using the at least two second sensors to obtain at least two second collection data; Synchronize the first collection data and the at least two second collection data to obtain the initial sensor data.
6. The method according to claim 1, characterized in that, The constructing an initial occupancy grid based on the initial sensor data includes: Input the initial sensor data into a multi-modal occupancy grid model to generate the initial occupancy grid, where the initial occupancy grid includes voxel information of surrounding obstacles of the vehicle.
7. The method according to claim 1, characterized in that, The processing the target occupancy grid to generate target annotation data includes: Input the target occupancy grid into a target detection model to generate a target detection result; Perform post-processing on the target detection result to obtain the target annotation data, and the post-processing includes at least one of trajectory tracking, motion compensation, and parameter tuning.
8. A training method for an autonomous driving model, characterized in that, Including: Obtain annotation data; Train an autonomous driving model using the annotation data; Wherein, the annotation data is obtained based on the data annotation method described in any one of claims 1-7.
9. An electronic device, characterized in that, Including a mutually coupled memory and a processor, and the processor is configured to execute program instructions stored in the memory to implement the data annotation method described in any one of claims 1 to 7, or to implement the training method of the autonomous driving model described in claim 8.
10. A non-volatile computer-readable storage medium storing program instructions thereon, characterized in that, When the program instructions are executed by a processor, they implement the data annotation method according to any one of claims 1 to 7, or implement the training method of the autonomous driving model according to claim 8.