Training data labeling method, device and equipment for visual 3D target detection model, and computer program product

By using multiple labeled vehicles to collect data in the intersection area and performing data fusion, pre-labeled information of visual 3D targets is generated, which solves the problems of high cost and frequent disassembly and assembly of LiDAR equipment, and realizes efficient training data labeling and model building.

CN120877019APending Publication Date: 2025-10-31ZHIDAO NETWORK TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510975060.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In the training of roadside visual 3D target detection models, existing technologies rely on high-cost LiDAR equipment that requires frequent disassembly and reassembly, resulting in low data acquisition efficiency. This fails to meet the needs of large-scale applications and limits the development of vehicle-road-cloud collaborative systems.

Method used

Multiple labeled vehicles are parked in the intersection area, and their LiDAR and autonomous driving perception systems collect data. Through data fusion and matching, pre-labeled information of visual 3D targets is generated to build a training dataset, avoiding the frequent deployment and disassembly of LiDAR.

Benefits of technology

It improves the efficiency and accuracy of training data labeling, reduces equipment deployment and maintenance costs, enhances the flexibility and efficiency of data collection, and meets the needs of large-scale data collection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877019A_ABST
    Figure CN120877019A_ABST
Patent Text Reader

Abstract

The invention discloses a training data labeling method, a training data labeling device and training data labeling equipment for a visual 3D target detection model, and a computer program product. Acquiring image data acquired by a plurality of road end cameras deployed in the current intersection area and vehicle end sensing data of a plurality of labeled vehicles parked at specified positions of the intersection area; comprehensive processing is carried out on the vehicle end sensing data, and pre-labeling information of a visual 3D target is generated; matching the image data of the plurality of road end cameras with pre-annotation information to obtain the pre-annotation information corresponding to the image data of each road end camera; and constructing training data according to the image data of the plurality of road end cameras and the corresponding pre-annotation information. According to the method, the pre-annotation information is generated by deploying the annotation vehicle which is convenient to move, a reliable annotation basis is provided for the road end image data, the training data annotation efficiency and accuracy are improved, and dependence on a road end laser radar is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of visual model training technology, and in particular to a method, apparatus and equipment, and computer program product for training data annotation of a visual 3D object detection model. Background Technology

[0002] With the continuous development of intelligent transportation systems, vehicle-road-cloud collaborative technology has become a key direction for improving the overall efficiency of transportation systems. A vehicle-road-cloud collaborative system is a complex large system composed of vehicles, other traffic participants, roadside infrastructure, cloud control platforms, related support platforms, and communication networks. Its core objective is to achieve information interaction and collaboration between vehicles, between vehicles and roadside infrastructure, and between vehicles and the cloud, thereby significantly improving the perception, decision-making, and execution capabilities of the entire transportation system.

[0003] In vehicle-road-cloud cooperative systems, roadside perception, as a crucial component, undertakes the task of real-time monitoring and information acquisition of the road environment, providing fundamental data support for the system's decision-making. Among these technologies, visual perception, with its cost advantage, has become a key option for roadside perception. Effective processing and analysis of roadside visual information can provide critical traffic information to various stakeholders in the transportation system, contributing to improved traffic safety and efficiency.

[0004] From the perspective of road area division, it can be mainly divided into two types: road segment areas and intersection areas. Compared with road segment areas, intersection areas have more complex road conditions, with numerous traffic participants, including vehicles, pedestrians, and cyclists, and the direction and state of traffic flow are also more variable. Therefore, deploying roadside visual perception systems in intersection areas has higher value, enabling a more comprehensive understanding of the traffic conditions at intersections and providing strong support for applications such as traffic management and autonomous driving.

[0005] Roadside visual 3D object detection is an essential capability for roadside visual perception. It uses images acquired by roadside cameras as input and, through advanced algorithms and models, detects detailed information about targets (such as vehicles, pedestrians, and cyclists) within the area, including their 3D position (3D coordinates of the center point on the ground plane), 3D size, heading angle, and category. This information is crucial for the decision-making of autonomous vehicles and the optimized scheduling of the entire traffic system. It provides vehicles with more accurate environmental perception, helping them make reasonable driving decisions, and also provides traffic management departments with real-time traffic data for scientific traffic planning and management.

[0006] However, achieving high-quality roadside visual 3D object detection hinges on possessing a high-quality visual 3D object detection model. The quality and diversity of the training dataset are core elements for training such a model. Existing technologies require data collection at various intersections to enable the visual 3D model to adapt to as many types of intersection scenarios as possible. During data collection, LiDAR equipment is typically installed to obtain the ground truth 3D values ​​of the targets.

[0007] However, LiDAR equipment suffers from significant economic costs. Due to its high price, it cannot be permanently installed at every intersection in practical applications. The current common practice is to install LiDAR equipment at one intersection, collect data from that intersection, and then remove and relocate the equipment to another intersection to continue collecting data. This approach not only requires substantial manpower and time for equipment installation, disassembly, and transportation, but also results in low data collection efficiency, failing to meet the demands of large-scale, high-efficiency data acquisition. This, in turn, limits the training and application of high-quality visual 3D target detection models, becoming a bottleneck restricting the further development of roadside visual perception technology in vehicle-road-cloud cooperative systems. Summary of the Invention

[0008] This application provides a method, apparatus, and computer program product for annotating training data of a visual 3D object detection model, so as to improve the annotation efficiency of training data for roadside 3D visual models.

[0009] The embodiments of this application adopt the following technical solutions:

[0010] In a first aspect, embodiments of this application provide a method for labeling training data for a visual 3D object detection model, the method comprising:

[0011] Under the condition that the preset training data labeling conditions are met in the current intersection area, image data collected by multiple roadside cameras deployed in the current intersection area, and vehicle-side perception data of multiple labeled vehicles parked at designated locations in the current intersection area are obtained.

[0012] The vehicle-side perception data of multiple labeled vehicles are comprehensively processed to generate pre-labeling information of visual 3D targets in the intersection area;

[0013] The image data of multiple roadside cameras are matched with the pre-annotation information of the visual 3D target to obtain the pre-annotation information of the visual 3D target corresponding to the image data of each roadside camera.

[0014] The training data for the visual 3D target detection model is constructed based on image data from multiple roadside cameras and pre-labeled information of corresponding visual 3D targets.

[0015] Optionally, the step of comprehensively processing the vehicle-side perception data of multiple labeled vehicles to generate pre-labeling information for visual 3D targets in the intersection area includes:

[0016] The vehicle-side perception data of each of the labeled vehicles are converted to a predefined local coordinate system for the intersection area;

[0017] The vehicle-side perception data of each labeled vehicle in the predefined local coordinate system of the intersection area are comprehensively processed to generate pre-labeling information of visual 3D targets in the intersection area.

[0018] Optionally, the vehicle-side perception data of the labeled vehicles includes the perception results of LiDAR, and the process of comprehensively processing the vehicle-side perception data of each labeled vehicle in a predefined local coordinate system of the intersection area to generate pre-labeling information for visual 3D targets in the intersection area includes:

[0019] The perception results of the LiDAR of each labeled vehicle in the predefined local coordinate system of the intersection area are merged to obtain the true value of the visual 3D target in the intersection area.

[0020] Optionally, the vehicle-side perception data of the labeled vehicles includes the perception results of the autonomous driving perception system, and the step of comprehensively processing the vehicle-side perception data of multiple labeled vehicles to generate pre-labeling information for visual 3D targets in the intersection area includes:

[0021] The perception results of each labeled vehicle in the predefined local coordinate system of the intersection area are fused to obtain the pre-labeling result of the visual 3D target in the intersection area.

[0022] Optionally, the step of matching the image data of the multiple roadside cameras with the pre-annotation information of the visual 3D target to obtain the pre-annotation information of the visual 3D target corresponding to the image data of each roadside camera includes:

[0023] The image data of each of the roadside cameras are converted to a predefined local coordinate system of the intersection area.

[0024] The image data of each roadside camera in the predefined local coordinate system of the intersection area are matched with the pre-labeled information of the visual 3D target to obtain the pre-labeled information of the visual 3D target corresponding to the image data of each roadside camera.

[0025] Optionally, the pre-annotation information of the visual 3D target includes the ground truth and pre-annotation results of the visual 3D target in the intersection area. The step of matching the image data from multiple roadside cameras with the pre-annotation information of the visual 3D target to obtain the pre-annotation information of the visual 3D target corresponding to the image data of each roadside camera includes:

[0026] The image data from multiple roadside cameras are matched with the ground truth values ​​of the visual 3D targets in the intersection area and the pre-annotation results of the visual 3D targets in the intersection area using timestamps, to obtain the ground truth values ​​and pre-annotation results of the visual 3D targets in the intersection area corresponding to the image data from each roadside camera.

[0027] Optionally, multiple marked vehicles are parked at different locations and in different directions in the intersection area, and the perception range of the lidar on the multiple marked vehicles covers the entire intersection area, as does the perception range of the autonomous driving perception system on the multiple marked vehicles.

[0028] Secondly, embodiments of this application also provide a training data annotation device for a visual 3D object detection model, the training data annotation device for the visual 3D object detection model comprising:

[0029] The acquisition unit is used to acquire image data collected by multiple roadside cameras deployed in the current intersection area, and vehicle-side perception data of multiple labeled vehicles parked at designated locations in the current intersection area, provided that the preset training data labeling conditions are met in the current intersection area.

[0030] The integrated processing unit is used to comprehensively process the vehicle-side perception data of multiple labeled vehicles to generate pre-labeling information of visual 3D targets in the intersection area.

[0031] The matching unit is used to match the image data of the multiple roadside cameras with the pre-labeling information of the visual 3D target to obtain the pre-labeling information of the visual 3D target corresponding to the image data of each roadside camera.

[0032] The construction unit is used to construct training data for the visual 3D target detection model based on image data from multiple roadside cameras and pre-labeled information of corresponding visual 3D targets.

[0033] Thirdly, embodiments of this application also provide an apparatus, comprising:

[0034] A processor; and a memory arranged to store computer-executable instructions, which, when executed, cause the processor to perform the training data annotation method of any of the aforementioned visual 3D object detection models.

[0035] Fourthly, embodiments of this application also provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the training data annotation method for any of the aforementioned visual 3D object detection models.

[0036] The above-mentioned at least one technical solution adopted in the embodiments of this application can achieve the following beneficial effects: The training data annotation method of the visual 3D target detection model in the embodiments of this application firstly acquires image data collected by multiple roadside cameras deployed in the current intersection area, and vehicle-end perception data of multiple labeled vehicles parked at designated locations in the current intersection area, under the condition that the current intersection area meets the preset training data annotation conditions; then, the vehicle-end perception data of the multiple labeled vehicles are comprehensively processed to generate pre-annotation information of visual 3D targets in the intersection area; then, the image data of the multiple roadside cameras are matched with the pre-annotation information of visual 3D targets to obtain the pre-annotation information of visual 3D targets corresponding to the image data of each roadside camera; finally, the training data of the visual 3D target detection model is constructed based on the image data of the multiple roadside cameras and the corresponding pre-annotation information of visual 3D targets. The training data annotation method for the visual 3D target detection model in this application generates pre-annotation information by deploying a mobile annotation vehicle, which provides a reliable annotation basis for the image data of the roadside camera, greatly improving the efficiency and accuracy of training data annotation. It avoids the dependence of existing roadside 3D visual model training data construction schemes on roadside LiDAR, thereby avoiding repeated disassembly and assembly of roadside LiDAR. Attached Figure Description

[0037] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0038] Figure 1 This is a flowchart illustrating a training data annotation method for a visual 3D object detection model according to an embodiment of this application.

[0039] Figure 2 This is a schematic diagram of a data annotation scenario for an intersection area in an embodiment of this application;

[0040] Figure 3 This is a schematic diagram of the structure of a training data annotation device for a visual 3D object detection model according to an embodiment of this application;

[0041] Figure 4 This is a schematic diagram of the structure of a device according to an embodiment of this application. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0043] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0044] Taking a crossroads as an example, a pole is erected on each road segment, and a camera is installed on each pole facing the intersection to cover the intersection area. Current solutions, in order to generate training data for training a visual 3D target detection model, require installing a LiDAR on each pole, also facing the intersection area. During data collection, four images and four LiDAR point clouds are acquired simultaneously. The LiDAR point clouds are used to provide the true 3D position and heading angle of targets within the area for data annotation. However, the size and shape of intersection areas vary greatly; some intersections are large, some are T-junctions, and some are crossroads, and the orientation and angles between each branch road also differ. To enable the visual 3D model to adapt to as many types of intersections as possible, data needs to be collected at various intersections, and LiDAR needs to be installed to obtain the true 3D values ​​of targets. LiDAR equipment is costly, and often it is installed at one intersection, collected data, and then disassembled and moved to another intersection, resulting in low operational efficiency.

[0045] Based on this, embodiments of this application provide a method for annotating training data for a visual 3D object detection model, such as... Figure 1 The diagram illustrates a flowchart of a training data annotation method for a visual 3D object detection model according to an embodiment of this application. The training data annotation method for the visual 3D object detection model includes at least the following steps S110 to S140:

[0046] Step S110: If the current intersection area meets the preset training data labeling conditions, acquire image data collected by multiple roadside cameras deployed in the current intersection area, as well as vehicle-side perception data of multiple labeled vehicles parked at designated locations in the current intersection area.

[0047] When labeling training data for a visual 3D object detection model, it is necessary to first determine whether the intersection area generated by the data to be collected and labeled meets the set training data labeling conditions. In the embodiment of this application, the training data labeling conditions set in the scenario of labeling training data for a roadside 3D visual model are mainly constraints on the data labeling environment of the intersection area.

[0048] like Figure 2 The diagram illustrates a data labeling scenario for an intersection area according to an embodiment of this application. Taking a crossroads as an example, typically each section of a crossroads is equipped with a pole, and each pole has a camera facing the intersection. Unlike existing solutions, this embodiment does not require deploying LiDAR on the poles in each intersection area. Instead, it prepares multiple vehicles equipped with LiDAR, which have undergone vehicle calibration and are equipped with an autonomous driving perception system. Typically, the perception system of an autonomous vehicle can cover at least 100 meters forward and 50 meters to the left, right, and rear. These vehicles serve as the "labeling vehicles" in this embodiment.

[0049] For example, in a crossroads scenario, four marked vehicles can be prepared, all facing the intersection, and parked at... Figure 2 The four black dashed rectangles in the diagram record the location and orientation information of each vehicle. It should be noted that the number and parking positions of vehicles can be flexibly adjusted according to the specific area, shape, and road conditions of the intersection; this embodiment does not impose specific limitations on this.

[0050] Once the above deployment is complete, the current intersection area can be considered to meet the preset training data annotation conditions. Multiple roadside cameras deployed in the current intersection area begin collecting image data and recording timestamp information. This image data contains visual information about various targets (such as vehicles and pedestrians) in the intersection. Simultaneously, multiple labeled vehicles parked at designated locations in the current intersection area also begin collecting vehicle-mounted perception data and recording corresponding timestamp information. These labeled vehicles are equipped with LiDAR and autonomous driving perception systems, enabling them to collect perception data of the surrounding environment, including the 3D position, 3D size, and category of targets.

[0051] To ensure the accuracy of subsequent processing, it is necessary to further synchronize the image data collected by multiple roadside cameras and the vehicle-side perception data collected by multiple labeled vehicles using GPS timing to unify the timestamps.

[0052] Step S120: The vehicle-side perception data of multiple labeled vehicles are comprehensively processed to generate pre-labeling information of visual 3D targets in the intersection area.

[0053] Because multiple labeled vehicles are parked in the intersection area, each with its own field of vision, the vehicle-side perception data they collect may overlap. Therefore, comprehensive processing of this data is necessary. For example, data fusion algorithms can be used to integrate perception data of the same target collected by different labeled vehicles, improving the accuracy and completeness of the data. Simultaneously, point cloud data collected by multiple labeled vehicles at the same time can be merged to obtain a larger and denser point cloud dataset.

[0054] After comprehensive processing, the pre-annotation information of the visual 3D targets in the intersection area can be obtained, which provides the basis for subsequent image data annotation.

[0055] Step S130: Match the image data of the multiple roadside cameras with the pre-annotation information of the visual 3D target to obtain the pre-annotation information of the visual 3D target corresponding to the image data of each roadside camera.

[0056] Since roadside cameras collect two-dimensional image data, while pre-annotation information is based on three-dimensional information obtained from vehicle-mounted perception data, by matching the image data from multiple roadside cameras with the pre-annotation information of visual 3D targets, the image data collected by each roadside camera can be mapped to the corresponding pre-annotation information of a visual 3D target. In this way, annotation information such as 3D position and 3D size is added to the target in each image.

[0057] Step S140: Construct training data for the visual 3D target detection model based on image data from multiple roadside cameras and pre-annotation information of corresponding visual 3D targets.

[0058] After determining the pre-annotated information of the visual 3D targets corresponding to the image data from the roadside cameras, this information can be further provided to annotators for manual verification, thereby improving the accuracy of the annotation results. Finally, the image data from the roadside cameras and the corresponding manually annotated information are combined to form a complete training data sample. For example, for each image captured by a roadside camera, and the final annotation information (including target category, 3D location, 3D size, etc.) corresponding to all targets in that image, a training sample is constituted. By collecting training samples from different time periods and scenarios in the intersection area, a richer and more diverse training dataset can be constructed. This training dataset can be used to train a visual 3D target detection model, enabling the model to learn the ability to accurately detect and locate 3D targets from 2D images.

[0059] In some embodiments of this application, the step of comprehensively processing the vehicle-end perception data of multiple labeled vehicles to generate pre-labeling information of visual 3D targets in the intersection area includes: converting the vehicle-end perception data of each labeled vehicle to a predefined local coordinate system of the intersection area; and comprehensively processing the vehicle-end perception data of each labeled vehicle in the predefined local coordinate system of the intersection area to generate pre-labeling information of visual 3D targets in the intersection area.

[0060] In this embodiment, a three-dimensional coordinate system can be established with the center point of the intersection as the origin, named the local coordinate system (local coordinate system of the intersection area). The x-axis is the positive direction, pointing eastwards towards the road; the y-axis is at a 90-degree angle to the x-axis, pointing northwards towards the road; and the z-axis points upwards. It should be noted that the center point of the intersection can be considered the geometric center of the intersection area. This center point does not need to be particularly precise and does not affect subsequent processing.

[0061] Based on the predefined local coordinate system, the external parameter transformation relationship of each vehicle to the local coordinate system can be further defined. Then, based on the external parameter transformation relationship of each vehicle to the local coordinate system, the vehicle-side perception data collected by each vehicle can be uniformly transformed to the local coordinate system, and then subsequent comprehensive processing can be performed.

[0062] By converting the vehicle-to-vehicle perception data from different labeled vehicles to a unified local coordinate system for the intersection area and then performing data fusion processing, redundant information can be removed, errors reduced, and the accuracy of target location, size, and other information labeling improved. Simultaneously, integrating data from multiple vehicles allows for more comprehensive coverage of targets within the intersection area, increasing data completeness.

[0063] In some embodiments of this application, the vehicle-side perception data of the labeled vehicles includes the perception results of LiDAR. The step of comprehensively processing the vehicle-side perception data of each labeled vehicle in the predefined local coordinate system of the intersection area to generate pre-labeling information of the visual 3D target in the intersection area includes: merging the LiDAR perception results of each labeled vehicle in the predefined local coordinate system of the intersection area to obtain the ground truth of the visual 3D target in the intersection area.

[0064] The perception results (i.e., laser point cloud data) collected by the LiDAR installed on each marked vehicle are based on the vehicle's own coordinate system. Before merging the LiDAR perception results, it is necessary to convert the laser point cloud data of each marked vehicle to a predefined local coordinate system of the intersection area, as described above, to ensure that all point cloud data are unified in the same coordinate system, thus preparing for subsequent merging processing.

[0065] Since the aforementioned embodiments have already performed time synchronization processing on all data from both the vehicle and roadside, multiple laser point cloud files from multiple vehicles at the same time, after time synchronization processing, can be merged into a single laser point cloud file. During the merging process, the point cloud data from each file is directly integrated into the new file without changing the coordinate information of each point (these coordinates have already been unified under the local coordinate system). The number of point clouds in the merged file is the sum of the previous files, meaning that the merged point cloud data contains more extensive intersection area information and a denser point cloud distribution.

[0066] The merged laser point cloud files are processed to extract representative target information (such as vehicles and pedestrians). Point cloud clustering algorithms can be used to group point clouds belonging to the same target together, thereby identifying different targets. Then, based on the clustered point cloud data, the target's category, 3D position in the local coordinate system, shape, size, and heading angle are determined. This information constitutes the ground truth of the visual 3D targets in the intersection area.

[0067] By merging the laser point cloud files of multiple labeled vehicles, the merged point cloud data can cover a wider intersection area. Due to limitations in installation location and field of view, the LiDAR of a single labeled vehicle may not be able to fully cover the entire intersection. Merging the point cloud data of multiple vehicles can overcome this deficiency, achieving effective coverage of all corners and directions of the intersection, providing more comprehensive information for subsequent target detection and labeling.

[0068] The number of points in the merged point cloud file is the sum of the points in the multiple files, meaning that the distribution of points within the same area is denser. Dense point cloud data can more accurately describe the shape and details of the target, thus providing more precise ground truth information about the target.

[0069] Compared to installing LiDAR equipment at every intersection for data collection, using multiple easily movable marking vehicles to collect data at different intersections and merge point clouds significantly reduces equipment deployment and maintenance costs, avoids the inefficiency caused by frequent disassembly and relocation of LiDAR equipment, and improves the efficiency and flexibility of data collection.

[0070] In some embodiments of this application, the vehicle-side perception data of the labeled vehicles includes the perception results of the autonomous driving perception system. The step of comprehensively processing the vehicle-side perception data of multiple labeled vehicles to generate pre-labeling information of visual 3D targets in the intersection area includes: fusing the perception results of the autonomous driving perception system of each labeled vehicle in a predefined local coordinate system of the intersection area to obtain the pre-labeling results of visual 3D targets in the intersection area.

[0071] The autonomous driving perception system equipped in a vehicle typically includes multiple sensors, such as LiDAR, cameras, and millimeter-wave radar. These sensors work together to collect information about the vehicle's surrounding environment and generate perception results. These results can include information such as the 3D position, 3D size, and category of targets around the vehicle.

[0072] As mentioned earlier, since the data collected by the autonomous driving perception system of each labeled vehicle is based on the vehicle's own coordinate system, it is necessary to transform this data to a predefined local coordinate system of the intersection area to ensure that the data collected by different vehicles are processed in the same coordinate system.

[0073] After transforming the perception results of different labeled vehicles into a local coordinate system, it is necessary to correlate these data to determine which data correspond to the same target. This can be done by matching the target's features (such as shape, size, color, and trajectory) with spatiotemporal information. For example, targets with similar positions and motion states within a short period of time can be preliminarily identified as the same target.

[0074] The core of post-fusion methods is to fuse the correlated target information. For the same target, different vehicles may collect different perception information; for example, LiDAR may provide more accurate 3D position information, while cameras may provide richer target appearance feature information. Post-fusion methods integrate information from these different sources to improve the accuracy of target detection and localization. Fusion algorithms can employ weighted averaging, assigning different weights to information from different sources based on the reliability and accuracy of different sensors, and then performing weighted calculations to obtain more accurate target information. Alternatively, methods such as Kalman filtering can be used to estimate and update the target's motion state, combining perception information from different times to obtain a more stable and accurate target state.

[0075] After post-fusion processing, the fused target category, 3D position, and size information are integrated to form a pre-annotation result of the visual 3D targets in the intersection area, serving as the initial annotation result. Further calibration can be achieved through manual annotation. During manual annotation, inaccurate pre-annotation results can be adjusted based on the laser point cloud to obtain more accurate annotation results.

[0076] The post-fusion method comprehensively utilizes the different perception information from the autonomous driving perception systems of multiple labeled vehicles, overcoming the limitations of perception from a single sensor or a single vehicle. Because it integrates the perception results from multiple labeled vehicles, the pre-labeled information contains richer target features and motion state information, providing more accurate pre-labeling information for subsequent manual annotation and reducing the cost of subsequent manual adjustments.

[0077] In some embodiments of this application, the step of matching the image data of the plurality of roadside cameras with the pre-annotation information of the visual 3D target to obtain the pre-annotation information of the visual 3D target corresponding to the image data of each roadside camera includes: converting the image data of each roadside camera to a predefined local coordinate system of the intersection area; and matching the image data of each roadside camera in the predefined local coordinate system of the intersection area with the pre-annotation information of the visual 3D target to obtain the pre-annotation information of the visual 3D target corresponding to the image data of each roadside camera.

[0078] When matching the image data of multiple roadside cameras with the pre-labeled information of visual 3D targets, the image data of each roadside camera can be converted to a predefined local coordinate system first. Here, multiple roadside cameras can be jointly calibrated to the local coordinate system in advance, that is, the camera extrinsic parameters of each of the multiple cameras in the local coordinate system are calibrated, and the above conversion is performed based on the calibrated camera extrinsic parameters.

[0079] After converting the image data to the local coordinate system, it can be matched with the pre-annotated information of the previously generated visual 3D targets. After matching, the image data collected by each roadside camera can be mapped to the corresponding pre-annotated information of the visual 3D targets. In this way, accurate three-dimensional position and size annotation information are added to the targets in each image.

[0080] By jointly calibrating multiple roadside cameras into a local coordinate system and transforming the image data to the same coordinate system based on camera extrinsic parameters, spatial consistency and alignment of image data acquired by different cameras can be achieved. This eliminates coordinate differences between different cameras due to variations in installation position and angle, providing a unified basis for subsequent data processing and analysis.

[0081] In some embodiments of this application, the pre-annotation information of the visual 3D target includes the ground truth and pre-annotation result of the visual 3D target in the intersection area. The step of matching the image data of the multiple roadside cameras with the pre-annotation information of the visual 3D target to obtain the pre-annotation information of the visual 3D target corresponding to the image data of each roadside camera includes: matching the image data of the multiple roadside cameras with the ground truth and the pre-annotation result of the visual 3D target in the intersection area respectively by timestamp to obtain the ground truth and pre-annotation result of the visual 3D target in the intersection area corresponding to the image data of each roadside camera.

[0082] When roadside cameras collect image data, they add a timestamp to each frame, recording the precise moment the image was captured. Simultaneously, the process of generating ground truth and pre-annotation results for the visual 3D targets in the intersection area also records relevant time information, usually also in the form of timestamps. These timestamps are crucial for matching.

[0083] Since the clocks of different devices (cameras, calculation units that generate true values ​​and pre-labeled results, etc.) may deviate, time synchronization is required before timestamp matching. For example, time synchronization can be performed through GPS time synchronization to ensure that the time base of all devices is consistent, so that the timestamps can accurately reflect the correspondence between different data collection and generation times.

[0084] For each roadside camera capturing image data, the timestamp set of ground truth and pre-annotated results for the visual 3D targets in the intersection area is iterated. The timestamp of the image data is compared one by one with the timestamps of the ground truth and pre-annotated results to find the closest match in time.

[0085] Once the ground truth and pre-annotation results of the 3D target that match the timestamp are found, the ground truth and pre-annotation results of the corresponding 3D target are associated with the image data of the roadside camera, thus obtaining the ground truth and pre-annotation results of the visual 3D target in the intersection area corresponding to the image data of each roadside camera.

[0086] By matching timestamps, image data from roadside cameras can be associated with ground truth and pre-annotation results of visual 3D targets. The obtained ground truth and pre-annotation results of visual 3D targets corresponding to image data from each roadside camera provide a more accurate and reliable reference for subsequent tasks such as image annotation and model training.

[0087] In some embodiments of this application, multiple marked vehicles are parked at different locations and in different directions in the intersection area, the perception range of the lidar on the multiple marked vehicles covers the entire intersection area, and the perception range of the autonomous driving perception system on the multiple marked vehicles covers the entire intersection area.

[0088] Based on factors such as the geographical shape, size, and traffic flow distribution of the intersection area, determine the different parking locations for multiple marked vehicles within the intersection area. For example, for a crossroads, marked vehicles can be parked at the right-hand entrance of the road in each of the four directions, a certain distance from the zebra crossing, to ensure that the target in each direction can be effectively perceived; for complex roundabouts, marked vehicles may need to be evenly distributed in different locations within the roundabout to comprehensively cover the traffic situation within the roundabout.

[0089] The parking orientation of the marked vehicles has also been carefully designed to ensure that their LiDAR and autonomous driving perception systems can perceive targets in the intersection area at the optimal angle. For example, vehicles can be parked facing the center of the intersection, or their orientation can be adjusted according to the specific layout of the intersection, so that the perception range of the LiDAR on multiple marked vehicles can cover the entire intersection area, and the perception range of the autonomous driving perception systems on multiple marked vehicles can also cover the entire intersection area.

[0090] Multiple marked vehicles are parked in different locations and directions, ensuring that the LiDAR and autonomous driving perception system cover the entire intersection area. This guarantees that all targets within the intersection area (such as vehicles, pedestrians, and bicycles) can be effectively perceived. Whether it's traffic participants in the center of the intersection or targets in the edge areas, they can all be included in the perception system's field of vision, greatly improving the comprehensiveness and completeness of target perception.

[0091] By having multiple labeled vehicles' perception systems work together, targets can be perceived from different angles and positions, reducing the limitations of single-sensor or single-vehicle perception. Comprehensive and accurate perception data provides a high-quality data foundation for subsequent tasks such as image data annotation and training of visual 3D object detection models.

[0092] This application embodiment also provides a training data annotation device 300 for a visual 3D object detection model, such as... Figure 3 The diagram shows a structural schematic of a training data annotation device for a visual 3D object detection model according to an embodiment of this application. The training data annotation device 300 for the visual 3D object detection model includes: an acquisition unit 310, a comprehensive processing unit 320, a matching unit 330, and a construction unit 340, wherein:

[0093] The acquisition unit 310 is used to acquire image data collected by multiple roadside cameras deployed in the current intersection area, and vehicle-side perception data of multiple labeled vehicles parked at a designated location in the current intersection area, provided that the preset training data labeling conditions are met in the current intersection area.

[0094] The integrated processing unit 320 is used to perform integrated processing on the vehicle-side perception data of multiple labeled vehicles to generate pre-labeling information of visual 3D targets in the intersection area.

[0095] The matching unit 330 is used to match the image data of the multiple roadside cameras with the pre-labeling information of the visual 3D target to obtain the pre-labeling information of the visual 3D target corresponding to the image data of each roadside camera.

[0096] The construction unit 340 is used to construct training data for the visual 3D target detection model based on the image data of the multiple roadside cameras and the pre-labeled information of the corresponding visual 3D targets.

[0097] In some embodiments of this application, the integrated processing unit 320 is specifically used to: convert the vehicle-end perception data of each of the labeled vehicles to a predefined local coordinate system of the intersection area; perform integrated processing on the vehicle-end perception data of each of the labeled vehicles in the predefined local coordinate system of the intersection area to generate pre-labeling information of visual 3D targets in the intersection area.

[0098] In some embodiments of this application, the vehicle-side perception data of the marked vehicles includes the perception results of LiDAR, and the integrated processing unit 320 is specifically used to: merge the perception results of LiDAR of each marked vehicle in a predefined local coordinate system of the intersection area to obtain the true value of the visual 3D target in the intersection area.

[0099] In some embodiments of this application, the vehicle-side perception data of the labeled vehicles includes the perception results of the autonomous driving perception system. The integrated processing unit 320 is specifically used to: fuse the perception results of the autonomous driving perception system of each labeled vehicle in a predefined local coordinate system of the intersection area to obtain the pre-labeling result of the visual 3D target in the intersection area.

[0100] In some embodiments of this application, the matching unit 330 is specifically used to: convert the image data of each of the roadside cameras to a predefined local coordinate system of the intersection area; match the image data of each of the roadside cameras in the predefined local coordinate system of the intersection area with the pre-annotation information of the visual 3D target to obtain the pre-annotation information of the visual 3D target corresponding to the image data of each of the roadside cameras.

[0101] In some embodiments of this application, the pre-annotation information of the visual 3D target includes the ground truth and pre-annotation result of the visual 3D target in the intersection area. The matching unit 330 is specifically used to: match the image data of the multiple roadside cameras with the ground truth and pre-annotation result of the visual 3D target in the intersection area according to the timestamp, so as to obtain the ground truth and pre-annotation result of the visual 3D target in the intersection area corresponding to the image data of each roadside camera.

[0102] In some embodiments of this application, multiple marked vehicles are parked at different locations and in different directions in the intersection area, the perception range of the lidar on the multiple marked vehicles covers the entire intersection area, and the perception range of the autonomous driving perception system on the multiple marked vehicles covers the entire intersection area.

[0103] It is understood that the above-mentioned training data annotation device for the visual 3D object detection model can implement each step of the training data annotation method for the visual 3D object detection model provided in the foregoing embodiments. The relevant explanations regarding the training data annotation method for the visual 3D object detection model are applicable to the training data annotation device for the visual 3D object detection model, and will not be repeated here.

[0104] Figure 4 This is a schematic diagram of the structure of a device according to an embodiment of this application. For example... Figure 4 As shown, the device includes one or more processors (or processing units), and may also include one or more memories coupled to the processors, and may also include a communication module coupled to the processors.

[0105] A communication module can be used to communicate with other devices or apparatuses, such as sending or receiving data and / or signals. A communication module may have at least one communication module for communication. A communication module may include any interface necessary for communicating with other devices. Exemplarily, a communication module may be a transceiver, circuit, bus, module, or other type of communication module.

[0106] The processor may include, but is not limited to, one or more of the following: a general-purpose computer, a special-purpose computer, a microcontroller, a digital signal processor (DSP), or a controller-based multi-core controller architecture. The device may have multiple processors, such as application-specific integrated circuit (ASIC) chips, which are time-dependent on a clock synchronized with the main processor.

[0107] The memory may include one or more non-volatile memories and one or more volatile memories. Examples of non-volatile memories include, but are not limited to, at least one of the following: read-only memory (ROM), electrically programmable read-only memory (EPROM), flash memory, hard disk, compact disc (CD), digital video disc (DVD), or other magnetic and / or optical storage. Examples of volatile memories include, but are not limited to, at least one of the following: random access memory (RAM), or other volatile memories that do not persist during the duration of a power outage.

[0108] A computer program consists of computer-executable instructions that are executed by an associated processor. Programs can be stored in ROM. A processor can perform any appropriate actions and processes by loading the program into RAM.

[0109] Possible implementations of this application can be achieved through a program, enabling the communication device to execute any of the processes discussed in the foregoing embodiments. Possible implementations of this application can also be achieved through hardware or a combination of software and hardware.

[0110] In some implementations, the program may be tangibly contained in a computer-readable storage medium, which may include in a device (such as in memory) or other storage device accessible by the device. The program may be loaded from the computer-readable storage medium into RAM for execution. The computer-readable storage medium may include any type of tangible non-volatile memory, such as ROM, EPROM, flash memory, hard disk, CD, DVD, etc.

[0111] This application also provides a computer-readable storage medium storing computer instructions or program code thereon, which, when executed by a processor, causes the processor to perform the methods and functions involved in any of the above embodiments. A computer-readable medium can be any tangible medium that contains or stores a program for or relating to an instruction execution system, apparatus, or device. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. More detailed examples of computer-readable storage media include electrical connections with one or more wires, magnetic media (e.g., disks, floppy disks, hard disks, magnetic tapes, magnetic storage devices), optical media (e.g., optical storage devices, DVDs), semiconductor media (e.g., solid-state drives), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), or any suitable combination thereof.

[0112] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. Embodiments of this application also provide at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. This computer program product includes one or more computer-executable instructions, such as instructions included in a program module, which execute in a device on a target's real or virtual processor to perform the processes, methods, and functions involved in any of the above embodiments. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0113] This application also proposes a computer program product, including a computer program or instructions that, when run on a computer, cause the computer to perform the processes, methods, and functions described in the above embodiments. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided as needed. The machine-executable instructions for the program modules can be executed locally or in a distributed device. In a distributed device, the program modules can reside in both local and remote storage media.

[0114] Generally, the various embodiments of this application can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software, which can be executed by a controller, microprocessor, or other computing device. Although various aspects of the embodiments of this disclosure are shown and described as block diagrams, flowcharts, or represented using some other illustration, it should be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0115] It should be noted that although embodiments of this application have been described above with reference to the accompanying drawings, these embodiments are not independent of each other, and they can be combined to obtain other embodiments. The methods, situations, categories, and classifications of embodiments in this application are only for the convenience of description and should not constitute a special limitation. Various methods, categories, situations, and features in embodiments can be combined with each other if logically consistent. The various embodiments of this application can be arbitrarily combined to achieve different technical effects. The embodiments of this application will not list various combinations.

[0116] Furthermore, although the operation of the methods of this disclosure is described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Rather, the steps depicted in the flowcharts may be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps. It should also be noted that the features and functions of two or more devices according to this disclosure may be embodied in one device. Conversely, the features and functions of one device described above may be further divided and embodied by multiple devices.

[0117] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0118] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for labeling training data for a visual 3D object detection model, characterized in that, The training data annotation method for the visual 3D object detection model includes: Under the condition that the preset training data labeling conditions are met in the current intersection area, image data collected by multiple roadside cameras deployed in the current intersection area, and vehicle-side perception data of multiple labeled vehicles parked at designated locations in the current intersection area are obtained. The vehicle-side perception data of multiple labeled vehicles are comprehensively processed to generate pre-labeling information of visual 3D targets in the intersection area; The image data of multiple roadside cameras are matched with the pre-annotation information of the visual 3D target to obtain the pre-annotation information of the visual 3D target corresponding to the image data of each roadside camera. The training data for the visual 3D target detection model is constructed based on image data from multiple roadside cameras and pre-labeled information of corresponding visual 3D targets.

2. The training data annotation method for the visual 3D object detection model according to claim 1, characterized in that, The process of comprehensively processing the vehicle-side perception data of multiple labeled vehicles to generate pre-labeling information for visual 3D targets in the intersection area includes: The vehicle-side perception data of each of the labeled vehicles are converted to a predefined local coordinate system for the intersection area; The vehicle-side perception data of each labeled vehicle in the predefined local coordinate system of the intersection area are comprehensively processed to generate pre-labeling information of visual 3D targets in the intersection area.

3. The training data annotation method for the visual 3D object detection model according to claim 2, characterized in that, The vehicle-side perception data of the labeled vehicles includes the perception results of LiDAR. The process of comprehensively processing the vehicle-side perception data of each labeled vehicle in the predefined local coordinate system of the intersection area to generate pre-labeling information for visual 3D targets in the intersection area includes: The perception results of the LiDAR of each labeled vehicle in the predefined local coordinate system of the intersection area are merged to obtain the true value of the visual 3D target in the intersection area.

4. The training data annotation method for the visual 3D object detection model according to claim 2, characterized in that, The vehicle-side perception data of the labeled vehicles includes the perception results of the autonomous driving perception system. The process of comprehensively processing the vehicle-side perception data of multiple labeled vehicles to generate pre-labeling information for visual 3D targets in the intersection area includes: The perception results of each labeled vehicle in the predefined local coordinate system of the intersection area are fused to obtain the pre-labeling result of the visual 3D target in the intersection area.

5. The training data annotation method for the visual 3D object detection model according to claim 1, characterized in that, The step of matching the image data of the multiple roadside cameras with the pre-annotation information of the visual 3D target to obtain the pre-annotation information of the visual 3D target corresponding to the image data of each roadside camera includes: The image data of each of the roadside cameras are converted to a predefined local coordinate system of the intersection area. The image data of each roadside camera in the predefined local coordinate system of the intersection area are matched with the pre-labeled information of the visual 3D target to obtain the pre-labeled information of the visual 3D target corresponding to the image data of each roadside camera.

6. The training data annotation method for the visual 3D object detection model according to claim 1, characterized in that, The pre-annotation information of the visual 3D target includes the ground truth and pre-annotation results of the visual 3D target in the intersection area. Matching the image data from multiple roadside cameras with the pre-annotation information of the visual 3D target to obtain the pre-annotation information of the visual 3D target corresponding to the image data of each roadside camera includes: The image data from multiple roadside cameras are matched with the ground truth values ​​of the visual 3D targets in the intersection area and the pre-annotation results of the visual 3D targets in the intersection area using timestamps, to obtain the ground truth values ​​and pre-annotation results of the visual 3D targets in the intersection area corresponding to the image data from each roadside camera.

7. The training data annotation method for the visual 3D object detection model according to any one of claims 1 to 6, characterized in that, Multiple marked vehicles are parked at different locations and in different directions in the intersection area. The LiDAR on the multiple marked vehicles covers the entire intersection area, and the autonomous driving perception system on the multiple marked vehicles also covers the entire intersection area.

8. A training data annotation device for a visual 3D object detection model, characterized in that, The training data annotation device for the visual 3D object detection model includes: The acquisition unit is used to acquire image data collected by multiple roadside cameras deployed in the current intersection area, and vehicle-side perception data of multiple labeled vehicles parked at designated locations in the current intersection area, provided that the preset training data labeling conditions are met in the current intersection area. The integrated processing unit is used to comprehensively process the vehicle-side perception data of multiple labeled vehicles to generate pre-labeling information of visual 3D targets in the intersection area. The matching unit is used to match the image data of the multiple roadside cameras with the pre-labeling information of the visual 3D target to obtain the pre-labeling information of the visual 3D target corresponding to the image data of each roadside camera. The construction unit is used to construct training data for the visual 3D target detection model based on image data from multiple roadside cameras and pre-labeled information of corresponding visual 3D targets.

9. An apparatus comprising: processor; And a memory arranged to store computer-executable instructions, which, when executed, cause the processor to perform the training data annotation method of any one of the visual 3D object detection models of claims 1 to 7.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the training data annotation method of any one of the visual 3D object detection models according to claims 1 to 7.