A dynamic object labeling method and apparatus
By acquiring multiple sets of data to be labeled and combining data from different acquisition locations and perspectives, dynamic object labeling is performed using multi-sensor fusion and SLAM algorithms. This solves the problem of missing or incorrect labeling in existing technologies and achieves efficient and high-quality dynamic object labeling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING VOYAGER TECH CO LTD
- Filing Date
- 2024-12-02
- Publication Date
- 2026-06-02
AI Technical Summary
Existing dynamic object annotation methods suffer from omissions and errors, and manual annotation is time-consuming and labor-intensive. Automatic annotation methods are limited by factors such as observation distance and target object occlusion, making it difficult to achieve high-quality and efficient dynamic object annotation.
By acquiring multiple sets of data to be labeled, and combining data from different acquisition locations and perspectives, data labeling and fusion are performed. By utilizing multi-sensor fusion, SLAM algorithms, and pose graph optimization techniques, global consistency reconstruction and fusion labeling of dynamic objects are achieved, reducing missed and incorrect labeling.
It achieves higher quality automatic annotation of dynamic objects, improves annotation efficiency and accuracy, and provides more comprehensive observation of dynamic objects.
Smart Images

Figure CN122135362A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data annotation technology, and more specifically, to a method and apparatus for annotating dynamic objects. Background Technology
[0002] With the rapid development of autonomous driving and assisted driving technologies, the perception of the vehicle's surrounding environment has become particularly important, especially the perception of dynamic objects around the vehicle. This perception capability largely depends on a large amount of high-quality data annotation and training. However, existing dynamic object annotation methods have significant shortcomings.
[0003] Traditional manual annotation methods involve manually annotating dynamic objects frame by frame. While manual annotation can provide relatively accurate results, it is time-consuming and labor-intensive. Although existing automatic annotation methods can significantly improve annotation efficiency, they are prone to omissions and errors due to limitations such as observation distance and target object occlusion.
[0004] Therefore, overcoming the problems of missing or incorrect labeling in existing technologies and improving the quality and efficiency of dynamic object labeling has become a key issue that urgently needs to be addressed in the development of autonomous driving technology. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a dynamic object annotation method and apparatus to combine data to be annotated from different acquisition positions and / or viewpoints to obtain more comprehensive dynamic object observations, thereby reducing omissions and mis-annotations and achieving higher quality automatic annotation of dynamic objects.
[0006] In a first aspect, embodiments of the present invention provide a dynamic object annotation method, the method comprising:
[0007] Multiple sets of data to be labeled are acquired, and the acquisition time and acquisition range of the multiple sets of data to be labeled are at least partially overlapping. Each set of data to be labeled contains data to be labeled collected by at least one sensor on a corresponding data acquisition vehicle. The data acquisition positions and / or data acquisition perspectives of different data acquisition vehicles are different in the same time and space.
[0008] Each of the data groups to be labeled is labeled to obtain the corresponding single-group labeling results;
[0009] Multiple single-group annotation results are fused to obtain corresponding fused annotation results, which include fused temporal tracking results corresponding to multiple dynamic objects respectively.
[0010] The fusion time-series tracking results of each dynamic object are mapped to a single set of annotation results that meet the mapping conditions to obtain the target annotation results corresponding to each set of data to be annotated.
[0011] Secondly, embodiments of the present invention provide a dynamic object annotation device, the device comprising:
[0012] The acquisition module is used to acquire multiple data groups to be labeled. The acquisition time and acquisition range of the multiple data groups to be labeled are at least partially overlapping. The data groups to be labeled contain data to be labeled collected by at least one sensor on the corresponding data acquisition vehicle. The data acquisition positions and / or data acquisition angles of different data acquisition vehicles are different in the same time and space.
[0013] The annotation module is used to annotate each of the data groups to be annotated, and obtain the corresponding single-group annotation results;
[0014] The fusion module is used to fuse multiple single-group annotation results to obtain corresponding fused annotation results, which include fused temporal tracking results corresponding to multiple dynamic objects respectively;
[0015] The mapping module is used to map the fusion time-series tracking results of each dynamic object to a single set of annotation results that meet the mapping conditions, so as to obtain the target annotation results corresponding to each set of data to be annotated.
[0016] Thirdly, an electronic device is provided, including a memory and a processor, the memory being used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in the first aspect above.
[0017] Fourthly, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when executed by a processor, the computer program implements the method described in the first aspect.
[0018] Fifthly, a computer program product is provided, comprising a computer program / instructions that, when executed by a processor, implement the method described in the first aspect above.
[0019] This invention acquires multiple sets of data to be labeled, with at least partial overlap in acquisition time and range. Each set contains data collected by at least one sensor on a corresponding data acquisition vehicle. Different vehicles have different acquisition positions and / or viewing angles within the same time and space. Each set is labeled to obtain a single-set labeling result. These single-set labels are then fused to obtain a fused labeling result, which includes fused temporal tracking results for multiple dynamic objects. The fused temporal tracking results for each dynamic object are mapped to single-set labels that meet the mapping conditions, resulting in the target labeling result for each set of data. This embodiment combines data from different acquisition positions and / or viewing angles to obtain more comprehensive observations of dynamic objects, thereby reducing missed and incorrect labeling and achieving higher-quality automatic labeling of dynamic objects. Attached Figure Description
[0020] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which:
[0021] Figure 1 This is a schematic diagram of the dynamic object annotation system according to an embodiment of the present invention;
[0022] Figure 2 This is a schematic diagram of the pose graph optimization results according to an embodiment of the present invention;
[0023] Figure 3 This is a flowchart of the dynamic object annotation method according to an embodiment of the present invention;
[0024] Figure 4 This is a flowchart of a method for determining a single set of annotation results according to an embodiment of the present invention;
[0025] Figure 5 This is a flowchart of the method for determining the fusion timing tracking results according to an embodiment of the present invention;
[0026] Figure 6 This is a flowchart of the target annotation box determination method according to an embodiment of the present invention;
[0027] Figure 7 This is a flowchart of the fusion timing tracking result mapping method according to an embodiment of the present invention;
[0028] Figure 8 This is a schematic diagram of the dynamic object annotation device according to an embodiment of the present invention;
[0029] Figure 9 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0030] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.
[0031] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.
[0032] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".
[0033] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0034] The solutions described in this specification and embodiments, if involving the processing of personal information, will be processed only under the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be processed within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.
[0035] Figure 1 This is a schematic diagram of a dynamic object annotation system according to an embodiment of the present invention. Figure 1 As shown, the dynamic object annotation system 100 includes an annotation server 101. In this embodiment, the annotation server 101 is a general-purpose data processing device that provides computing or application services for the dynamic object annotation system. It can be a single computer, a cluster of multiple computers, or a cloud server that can flexibly adjust computing resources through cloud technology. The annotation server 101 is used to acquire data to be annotated and to annotate the data with dynamic objects. Dynamic objects refer to dynamic traffic participants such as vehicles, pedestrians, motorcycles / bicycles, etc.
[0036] The acquisition of data to be labeled can be divided into direct acquisition and indirect acquisition. Direct acquisition refers to acquiring the data to be labeled from the data acquisition device 103. The data acquisition device 103 is a device that provides data acquisition services for the dynamic object labeling system. The data acquisition device 103 can be a general-purpose vehicle such as a vehicle, electric vehicle, or bicycle equipped with at least one sensor. The sensor may include a camera sensor, radar sensor, lidar sensor, and / or inertial measurement unit (IMU), etc. The data acquisition device 103 can also be a roadside device using Vehicle-to-Everything (V2X) technology. The roadside device includes an on-board unit and a roadside unit. The on-board unit is installed on a general-purpose vehicle such as a vehicle, electric vehicle, or bicycle, while the roadside unit is installed on a fixed object on the road. The fixed object may be a traffic light, traffic camera, and / or intelligent parking equipment, etc. The V2X roadside device communicates with other devices and systems through a wireless communication module to achieve comprehensive traffic perception and collaborative control.
[0037] Indirect acquisition refers to obtaining the data to be labeled from data storage source 102, which includes networks, databases, data warehouses, public datasets and / or third-party data providers.
[0038] In one possible implementation, the data to be labeled obtained by the direct acquisition method described above can also be stored in the data storage source 102 for later retrieval.
[0039] The dynamic object annotation process of the dynamic object annotation system 100 in this embodiment can be divided into a data mining stage, a multi-source data alignment stage, and an automatic annotation stage.
[0040] During the data mining phase, the annotation server 101 interacts with the data storage source 102 and / or the data acquisition device 103 via the network to obtain a large amount of data to be labeled. The acquisition methods are the direct acquisition methods and / or indirect acquisition methods described above.
[0041] In one possible implementation, the data to be labeled is multi-source collected data, wherein the multiple data sources can be multiple general-purpose vehicles of the same type, or multiple general-purpose vehicles of different types, or multiple roadside devices in the same area.
[0042] In one possible implementation, the data to be labeled can also be obtained from data storage sources 102 such as web crawlers and public datasets.
[0043] After acquiring a large amount of data to be labeled, data mining should be performed based on spatiotemporal information to obtain multiple sets of data to be labeled at the same location and at similar times. During data mining, the data to be labeled should meet spatial constraints, temporal constraints, and richness constraints.
[0044] Spatial constraints refer to the requirement that the regions corresponding to multiple sets of data to be labeled should be located near the same location, their acquisition trajectories should be close in distance, and the visible observation ranges of the data should have a certain degree of spatial overlap. The visible observation range can be calculated by superimposing the acquisition trajectories with the observation range of the data acquisition device 103.
[0045] Time constraints refer to the fact that multiple groups of data to be labeled have similar collection time ranges and that there is a certain overlap rate in the ranges.
[0046] Richness constraints refer to the requirement that multiple sets of data to be labeled must meet a certain level of richness. This could be observations from different lanes, different road directions, or different perspectives, etc. The purpose of richness constraints is to obtain a more comprehensive and complete dynamic environmental observation at a given time and location.
[0047] Through the above data mining, multiple data groups to be labeled can be obtained, wherein the collection time and collection range of the multiple data groups to be labeled overlap at least partially.
[0048] Taking direct data acquisition as an example, with data acquisition device 103 being a data acquisition vehicle, there is a one-to-one mapping relationship between the data group to be labeled and the data acquisition vehicle. That is, all data to be labeled in the data group is collected by the same data acquisition vehicle. The data groups to be labeled collected by multiple data acquisition vehicles have a spatiotemporal overlap relationship. Different data acquisition vehicles have different data acquisition positions and / or data acquisition perspectives in the same spatiotemporal space. Each group of data to be labeled contains at least one data to be labeled. Each data to be labeled is data collected by a sensor on the corresponding data acquisition vehicle. If the data acquisition vehicle contains multiple sensors, such as one laser sensor and two camera sensors, then the data group to be labeled corresponding to the data acquisition vehicle contains three data to be labeled.
[0049] During the multi-source data alignment stage, the annotation server 101 aligns the trajectories corresponding to each data to be labeled. The alignment process can be divided into intra-group alignment and inter-group alignment. If the data group to be labeled contains only one data to be labeled, intra-group alignment is not required, and inter-group alignment is performed directly. If the data group to be labeled contains multiple data to be labeled, intra-group alignment is performed first, and then inter-group alignment is performed.
[0050] Intra-group alignment refers to aligning the trajectories corresponding to multiple unlabeled data points within a group of data to be labeled. Specifically, high-precision 3D reconstruction is performed within the group of data to be labeled to ensure global consistency in vehicle pose across multiple unlabeled data points. Each unlabeled data point contains multiple data frames to be identified, and the sensor acquires the corresponding data frame at a predetermined acquisition time interval.
[0051] In one possible implementation, multi-sensor fusion can be used to achieve intra-group alignment. For example, tightly coupled Lidar Odometry and / or Lidar-Inertial-Odometry (LIO) can be used to align the point clouds in the data frames to be labeled from multiple sensors such as laser sensors and IMU sensors, ensuring global consistency within the group.
[0052] In one possible implementation, Simultaneous Localization and Mapping (SLAM) algorithms can be used for high-precision reconstruction of 3D maps. SLAM technology combines localization and mapping tasks, enabling the determination of one's own position and the construction of a map of the surrounding environment using sensor data without prior environmental information. SLAM typically relies on various sensors, such as LiDAR, cameras, and inertial measurement units, to acquire environmental information. The data provided by these sensors is used for feature extraction, matching, and map construction. The SLAM process includes data acquisition, feature extraction, data association, pose estimation, map construction, and loop closure detection. Data acquisition refers to acquiring data about the surrounding environment (i.e., the data to be identified) through sensors, such as laser scan data and image data. Feature extraction refers to extracting feature points or feature lines that can be used for matching from the sensor data (i.e., the data to be identified). These features are usually stable structures in the environment, such as streetlights or routes. Data association refers to matching the features acquired at the current moment with features from previous moments to determine the movement of dynamic objects relative to the environment. Pose estimation refers to estimating the current position and orientation of a dynamic object based on data association results, using geometric constraints and probabilistic models. Map building refers to adding new information to a map after the pose of the dynamic object has been determined, thereby gradually constructing a complete map. Loop closure detection refers to detecting and correcting positioning errors, thereby improving the accuracy and consistency of the map.
[0053] Inter-group alignment refers to aligning multiple groups of data to be labeled that have already been aligned within a group. Specifically, pose graph optimization can be used to jointly optimize multiple groups of data to be labeled to achieve overall alignment.
[0054] Pose Graph Optimization (PGraph Optimization) is a graph-building method based on least-squares nonlinear optimization. It uses a vertex-edge pose graph to model the Simultaneous Localization and Mapping (SLAM) problem for dynamic objects. Pose Graph Optimization includes pose graph construction and pose graph optimization. Pose graph construction refers to calculating the relative motion constraints between adjacent nodes and the motion constraints between non-adjacent nodes. In laser localization mapping, the construction of motion constraints is mainly achieved through scan matching, that is, matching two frames of laser data together so that the end point of one frame of laser beam can match the corresponding end point of another frame of laser beam corresponding to the same object in space. Pose graph optimization involves solving a nonlinear least squares optimization problem, similar to solving graph optimization problems. It uses an iterative approach, calculating the increment of an optimization variable each time and adding it to the current optimization variable to trigger the next iteration. During iteration, an equivalent linear solution equation is first constructed, and then the increment of the optimization variable is calculated. Optimization convergence is considered achieved when the number of iterations reaches an upper limit or the objective function corresponding to the calculated optimization variable is less than a certain threshold. The value of the optimization variable at this point is the solved value. Optimization algorithms can include Gauss-Newton's method, the Levenberg-Marquardt algorithm, and the LM optimization method.
[0055] Figure 2 This is a schematic diagram illustrating the pose graph optimization results of an embodiment of the present invention. Figure 2 As shown, in the pose graph optimization results, the pose graphs corresponding to multiple groups of data to be labeled are similar.
[0056] In the automatic annotation stage, the annotation server 101 annotates each of the data groups to be annotated, obtaining corresponding single-group annotation results. Then, it fuses multiple single-group annotation results to obtain corresponding fused annotation results. The fused annotation results include fused temporal tracking results corresponding to multiple dynamic objects. Finally, the fused temporal tracking results of each dynamic object are mapped to single-group annotation results that meet the mapping conditions, obtaining the target annotation results corresponding to each of the data groups to be annotated. Specific annotation results are detailed in the following embodiments, and will not be described in detail here.
[0057] The system of this invention acquires multiple sets of data to be labeled, with the acquisition time and acquisition range of the multiple sets of data to be labeled overlapping at least partially. Each set of data to be labeled contains data collected by at least one sensor on a corresponding data acquisition vehicle. Different data acquisition vehicles have different data acquisition positions and / or data acquisition perspectives in the same space-time. Each set of data to be labeled is labeled to obtain a corresponding single-set labeling result. Multiple single-set labeling results are fused to obtain a corresponding fused labeling result. The fused labeling result includes fused temporal tracking results corresponding to multiple dynamic objects. The fused temporal tracking results of each dynamic object are mapped to single-set labeling results that meet the mapping conditions to obtain the target labeling result corresponding to each set of data to be labeled. This embodiment combines the data to be labeled from different acquisition positions and / or perspectives to obtain a more comprehensive observation of dynamic objects, thereby reducing the occurrence of missed and incorrect labels and achieving higher quality automatic labeling of dynamic objects.
[0058] Figure 3 This is a flowchart of the dynamic object annotation method according to an embodiment of the present invention. Figure 3 As shown, the dynamic object annotation method includes the following steps:
[0059] Step S301: Obtain multiple groups of data to be labeled.
[0060] In this context, the acquisition time and acquisition range of multiple sets of data to be labeled overlap at least partially. Each set of data to be labeled contains data collected by at least one sensor on a corresponding data acquisition vehicle. Different data acquisition vehicles have different data acquisition positions and / or data acquisition perspectives in the same time and space. The data to be labeled refers to data containing dynamic objects, which are dynamic traffic participants such as vehicles, pedestrians, motorcycles / bicycles, etc.
[0061] In other words, multiple data acquisition vehicles collect data from the same location at similar times. Each data acquisition vehicle corresponds to a set of data to be labeled, and each set of data to be labeled contains all the data collected by that vehicle. Each set of data to be labeled contains at least one data point to be labeled, and these data points are mapped to the sensors on the vehicle; that is, each sensor on the data acquisition vehicle corresponds to one data point to be labeled. Each data point to be labeled contains multiple data frames to be labeled, and each data frame to be labeled corresponds to a specific time point. In other words, the sensors collect the corresponding data frames to be labeled according to a preset data collection time interval.
[0062] It is worth noting that the aforementioned data collection vehicles can be replaced with any general-purpose vehicle capable of data collection.
[0063] Step S302 involves labeling each of the data groups to be labeled to obtain the corresponding single-group labeling results.
[0064] Specifically, the data annotation for dynamic objects includes not only object detection but also object tracking.
[0065] In one possible implementation, the data annotation of dynamic objects also includes target shape optimization and target orientation optimization.
[0066] Figure 4 This is a flowchart of a method for determining a single set of annotation results according to an embodiment of the present invention. Figure 4 As shown, the method for determining a single set of annotation results includes the following steps:
[0067] Step S401: Input each of the data frames to be labeled into the target detection model to obtain the corresponding single-frame target detection result, wherein the single-frame target detection result includes multiple dynamic objects.
[0068] The object detection model is a pre-trained offline detection model capable of accurately identifying and locating various objects on the road, such as vehicles, pedestrians, and traffic signs. The object detection model can be a YOLO series model, a SingleShot MultiBox Detector model, or a Faster R-CNN (Regions with Convolutional Neural Networks) model, etc.
[0069] The input data for the target detection model is the data to be labeled, which can be various types of data such as laser point cloud data or image data. In one possible implementation, the target detection model is input with one or more consecutive data frames to be labeled, as well as the image data corresponding to each data frame to be labeled.
[0070] The output data of the object detection model is the single-frame object detection result corresponding to each data frame to be labeled. The single-frame object detection result includes multiple dynamic objects (i.e., the object detection results corresponding to the dynamic objects). The object detection result corresponding to each dynamic object includes the acquisition time, bounding box, prediction category, prediction speed, confidence level, etc. of the dynamic object.
[0071] Step S402: Input the multiple single-frame target detection results corresponding to each of the data to be labeled into the target tracking model to obtain the temporal tracking results of each of the dynamic objects.
[0072] The temporal tracking results include bounding boxes corresponding to different time points of the dynamic object and the point clouds corresponding to the bounding boxes. The target tracking model is a model based on a target tracking algorithm. It identifies dynamic objects and predicts their positions in consecutive data frames to be labeled, and then searches and matches them in the next data frame to be labeled based on the prediction results, thereby achieving continuous tracking of dynamic objects. The matching process can use various distance metrics, such as Euclidean distance and similarity matching. Target tracking algorithms can be broadly divided into two categories: online learning-based tracking algorithms and detection-based tracking algorithms (Tracking-by-Detection). Online learning-based tracking algorithms continuously update the model during the tracking process to adapt to changes in the appearance of the target, such as KCF (Kernelized Correlation Filters), TLD (Tracking-Learning-Detection), and MOSSE (Minimum Output Sum of Squared Error). Detection-based tracking algorithms first run a target detection algorithm in each frame, and then use data association techniques (such as the Hungarian algorithm, Kalman filtering, etc.) to match the detected targets with existing tracking targets, such as SORT (SimpleOnline and Realtime Tracking) and DeepSORT.
[0073] The target tracking model takes multiple consecutive frames of data to be labeled as input and outputs the time-series tracking results of each dynamic object.
[0074] In one possible implementation, the time-series tracking result is a set of time-continuous target detection results for the corresponding dynamic object, that is, the target detection results of the dynamic object at time points T0, T1, T2, ... TN.
[0075] In one possible implementation, the velocity information of the dynamic object can be predicted based on the target detection results of the dynamic object in two adjacent data frames to be labeled, so that the target detection results at any time point within the start and end time of the time-series tracking results can be interpolated.
[0076] Step S403: Input the temporal tracking results of each dynamic object into the object optimization model to correct the shape of the corresponding bounding box based on the point cloud.
[0077] Specifically, after obtaining the temporal tracking results of the dynamic object, the corresponding point cloud can be cropped from the original point cloud at the corresponding time point based on the corresponding bounding boxes at each time point, and input into the object optimization model. The object optimization model outputs the corrected bounding boxes of the dynamic object for each frame. In other words, the input data of the object optimization model are the bounding boxes and point clouds of the dynamic object in each frame to be labeled, and the output data are the bounding boxes whose shapes are corrected according to the point cloud layout.
[0078] In one possible implementation, the object optimization model is a Refiner model, also known as a refiner or polisher. A Refiner model is typically a deep learning model.
[0079] Step S404: Correct the orientation of each annotation box according to the relative positional relationship between the multiple annotation boxes in each of the time-series tracking results.
[0080] Step S405: Determine the single set of annotation results based on the time-series tracking results corresponding to the multiple dynamic objects.
[0081] In other words, the time-series tracking results corresponding to all dynamic objects appearing in the data group to be labeled constitute a single set of labeling results.
[0082] It should be noted that in the process of determining the single-group annotation results, the above steps S403 and S404 are not necessary steps, but are post-processing steps used to correct the single-group annotation results to make the single-group annotation results more accurate.
[0083] In one possible implementation, the post-processing step may further include high-confidence bounding box filtering. Specifically, it may combine the original point cloud information to filter out low-confidence bounding boxes and their corresponding target detection results.
[0084] Step S303 involves fusing multiple single-group annotation results to obtain corresponding fused annotation results, which include fused temporal tracking results corresponding to multiple dynamic objects.
[0085] Specifically, the data fusion process can be divided into multi-source data clustering and multi-source data fusion.
[0086] Figure 5 This is a flowchart illustrating the method for determining the fusion timing tracking results according to an embodiment of the present invention. Figure 5 As shown, the method for determining the fusion timing tracking result includes the following steps:
[0087] Step S501: Determine multiple time-series tracking results belonging to the same dynamic object from different single-group annotation results.
[0088] Specifically, dynamic objects in multiple single-group annotation results can be matched based on factors such as the predicted category, trajectory, acquisition time, acquisition viewpoint, shape of the annotation box and / or orientation of the annotation box of each dynamic object in a single group annotation result, thereby determining multiple time-series tracking results belonging to the same dynamic object.
[0089] For example, the trajectory of a corresponding dynamic object can be determined based on the temporal tracking results. Based on the trajectory and the predicted category, multiple temporal tracking results belonging to the same dynamic object can be determined from different single-group annotation results. Specifically, the complete trajectory of the dynamic object within the tracking time range and the corresponding bounding box target detection results are obtained based on the temporal tracking results. Then, the temporal tracking results from different groups are matched based on the trajectory and bounding box target detection results. A successful match indicates that they belong to the same dynamic object.
[0090] Step S502: Cluster the multiple time-series tracking results to obtain the clustering results of the corresponding dynamic objects.
[0091] Specifically, the screening condition for determining whether different time-series tracking results can be clustered is that if the acquisition trajectories of two different time-series tracking results overlap within the same time period, and the intersection over union (IOU) of their respective bounding boxes at each moment within the overlapping time range is greater than a certain threshold, then the two are considered to be clustered.
[0092] In one possible implementation, clustering results are determined based on the trajectory overlap rate corresponding to multiple time-series tracking results and the overlap rate of bounding boxes with similar spatiotemporal information. The similar spatiotemporal information refers to the time difference between the acquisition of the bounding boxes being within a preset aggregation time difference threshold, and the position difference between the bounding boxes being within a preset position difference threshold. The preset aggregation time difference threshold and the preset position difference threshold can be set and adjusted according to actual conditions.
[0093] Step S503: Determine the labeling time period corresponding to the clustering result based on the collection time of all the label boxes in the clustering result.
[0094] Specifically, determine the earliest and latest collection times corresponding to all the labeled boxes in the clustering results, and use them as the start and end times of the labeled time periods, respectively.
[0095] Step S504: Determine multiple collection time points from the labeled time period at preset time intervals.
[0096] Step S505: Determine the target annotation box corresponding to each of the acquisition time points.
[0097] Figure 6This is a flowchart illustrating the target bounding box determination method according to an embodiment of the present invention. Figure 6 As shown, the target bounding box determination method includes the following steps:
[0098] Step S601: Determine the alternative time periods corresponding to each of the aforementioned collection time points.
[0099] Specifically, alternative time periods are constructed according to a pre-set alternative time period construction method. For example, the data collection time point can be used as the starting point of the alternative time period and the alternative time period can be determined by a preset time length; or the data collection time point can be used as the ending point of the alternative time period and the alternative time period can be determined by a preset time length; or the data collection time point can be used as the midpoint of the alternative time period and the alternative time period can be determined by a preset time length. The above are just examples, and this embodiment does not limit the method for determining alternative time periods.
[0100] Step S602: The annotation boxes that are within the candidate time period are determined as candidate annotation boxes.
[0101] Step S603: Determine whether the number of candidate annotation boxes is greater than 1.
[0102] If not, proceed to step S604; if yes, proceed to step S605.
[0103] Step S604: The candidate annotation box is determined as the target annotation box.
[0104] In other words, if there is a candidate label box, it is directly used as the target label box.
[0105] Step S605: In response to the existence of multiple candidate label boxes, the candidate label box with the highest confidence level is determined as the target label box.
[0106] Specifically, in response to the existence of multiple candidate bounding boxes, the corresponding confidence level is determined based on the detection score (i.e., the confidence level of the bounding box), the sparsity of the point cloud, and / or the image quality of the candidate bounding boxes, and the candidate bounding box with the highest confidence level is determined as the target bounding box.
[0107] Step S506: Input the multiple target annotation boxes corresponding to the clustering results into the target tracking model to obtain the corresponding fusion annotation results.
[0108] In other words, based on the selected high-confidence bounding boxes at each time step, the target tracking algorithm is rerun to obtain the tracking results of the fused dynamic object, which is the fused annotation result.
[0109] Step S304: Map the fusion time-series tracking results of each dynamic object to a single set of annotation results that meet the mapping conditions to obtain the target annotation results corresponding to each set of data to be annotated.
[0110] The mapping conditions include the existence of point cloud at the corresponding position of the bounding box to be mapped in the single set of annotation results, and / or, the bounding box to be mapped is not occluded in the single set of annotation results. A method to determine whether it is occluded can be to emit a ray from the current acquisition position toward the target dynamic object and detect whether it is occluded by other objects.
[0111] Figure 7 This is a flowchart of the fusion timing tracking result mapping method according to an embodiment of the present invention. Figure 7 As shown, the fusion timing tracking result mapping method includes the following steps:
[0112] Step S701: Identify the dynamic objects that are the same as those in the fused annotation results and the single annotation results as the target objects.
[0113] Step S702: Determine the bounding box corresponding to the target object in the fusion annotation result as the bounding box to be mapped.
[0114] Step S703: Based on the time point corresponding to the annotation box to be mapped, replace the annotation box corresponding to the time point in the single set of annotation results with the annotation box to be mapped to obtain the target annotation result.
[0115] In one possible implementation, the data to be labeled should be aligned before data annotation. The alignment process can be divided into intra-group alignment and inter-group alignment. If the data group to be labeled contains only one data point, intra-group alignment is not required, and inter-group alignment is performed directly. If the data group to be labeled contains multiple data points, intra-group alignment is performed first, followed by inter-group alignment. Specifically, in response to the data group to be labeled containing multiple data points, the multiple data points are reconstructed in 3D to align them. The aligned data groups are then subjected to pose graph optimization to align them. Specific alignment methods are detailed in the above embodiments and will not be repeated here.
[0116] The method of this invention involves acquiring multiple sets of data to be labeled, where the acquisition time and acquisition range of the multiple sets of data to be labeled at least partially overlap. Each set of data to be labeled contains data acquired by at least one sensor on a corresponding data acquisition vehicle. Different data acquisition vehicles have different data acquisition positions and / or data acquisition perspectives in the same time and space. Each set of data to be labeled is labeled to obtain a corresponding single-set labeling result. Multiple single-set labeling results are fused to obtain a corresponding fused labeling result. The fused labeling result includes fused temporal tracking results corresponding to multiple dynamic objects. The fused temporal tracking results of each dynamic object are mapped to single-set labeling results that meet the mapping conditions to obtain the target labeling result corresponding to each set of data to be labeled. This embodiment combines the data to be labeled from different acquisition positions and / or perspectives to obtain a more comprehensive observation of dynamic objects, thereby reducing the occurrence of missed and incorrect labels and achieving higher quality automatic labeling of dynamic objects.
[0117] Figure 8 This is a schematic diagram of a dynamic object annotation device according to an embodiment of the present invention. Figure 8 As shown, the dynamic object annotation device includes:
[0118] The acquisition module 801 is used to acquire multiple data groups to be labeled. The acquisition time and acquisition range of the multiple data groups to be labeled are at least partially overlapping. The data groups to be labeled include data to be labeled collected by at least one sensor on the corresponding data acquisition vehicle. The data acquisition positions and / or data acquisition angles of different data acquisition vehicles are different in the same time and space.
[0119] The annotation module 802 is used to annotate each of the data groups to be annotated, and obtain the corresponding single-group annotation results.
[0120] The fusion module 803 is used to fuse multiple single-group annotation results to obtain corresponding fused annotation results, which include fused temporal tracking results corresponding to multiple dynamic objects respectively.
[0121] The mapping module 804 is used to map the fusion time-series tracking results of each dynamic object to a single set of annotation results that meet the mapping conditions, so as to obtain the target annotation results corresponding to each set of data to be annotated.
[0122] The apparatus of this invention acquires multiple sets of data to be labeled, with the acquisition time and acquisition range of the multiple sets of data to be labeled overlapping at least partially. Each set of data to be labeled contains data collected by at least one sensor on a corresponding data acquisition vehicle. Different data acquisition vehicles have different data acquisition positions and / or data acquisition perspectives in the same time and space. Each set of data to be labeled is labeled to obtain a corresponding single-set labeling result. The multiple single-set labeling results are fused to obtain a corresponding fused labeling result. The fused labeling result includes fused temporal tracking results corresponding to multiple dynamic objects. The fused temporal tracking results of each dynamic object are mapped to the single-set labeling results that meet the mapping conditions to obtain the target labeling result corresponding to each set of data to be labeled. This embodiment combines the data to be labeled from different acquisition positions and / or perspectives to obtain a more comprehensive observation of dynamic objects, thereby reducing the occurrence of missed and incorrect labels and achieving higher quality automatic labeling of dynamic objects.
[0123] Figure 9 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Figure 9 As shown, Figure 9 The illustrated electronic device is a dynamic object annotation device, comprising a general computer hardware architecture, including at least a processor 901 and a memory 902. The processor 901 and memory 902 are connected via a bus 903. The memory 902 is adapted to store instructions or programs executable by the processor 901. The processor 901 can be a standalone microprocessor or a collection of one or more microprocessors. Thus, the processor 901 executes the instructions stored in the memory 902, thereby performing the method flow of the embodiments of the present invention as described above to process data and control other devices. The bus 903 connects the aforementioned components together, and also connects these components to a display controller 904, a display device, and an input / output (I / O) device 905. The input / output (I / O) device 905 can be a mouse, keyboard, modem, network interface, touch input device, motion-sensing input device, printer, and other devices known in the art. Typically, the input / output device 905 is connected to the system via an input / output (I / O) controller 906.
[0124] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus (devices), or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0125] This application is described with reference to flowchart illustrations of methods, apparatus (devices), and computer program products according to embodiments of this application. It should be understood that each step in the flowchart can be implemented by computer program instructions.
[0126] These computer program instructions may be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction means, the implementation process of which is described in the instruction means. Figure 1 The function specified in one or more processes.
[0127] These computer program instructions may also be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, produce instructions for implementing processes. Figure 1 A device for a function specified in one or more processes.
[0128] This invention acquires multiple sets of data to be labeled, with at least partial overlap in acquisition time and range. Each set contains data collected by at least one sensor on a corresponding data acquisition vehicle. Different vehicles have different acquisition positions and / or viewing angles within the same time and space. Each set is labeled to obtain a single-set labeling result. These single-set labels are then fused to obtain a fused labeling result, which includes fused temporal tracking results for multiple dynamic objects. The fused temporal tracking results for each dynamic object are mapped to single-set labels that meet the mapping conditions, resulting in the target labeling result for each set of data. This embodiment combines data from different acquisition positions and / or viewing angles to obtain more comprehensive observations of dynamic objects, thereby reducing missed and incorrect labeling and achieving higher-quality automatic labeling of dynamic objects.
[0129] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program for use by a computer to execute some or all of the above-described method embodiments.
[0130] Another embodiment of the present invention relates to a computer program product, including a computer program / instructions that, when executed by a processor, implement some or all of the above-described method embodiments.
[0131] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program specifying the relevant hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0132] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A dynamic object annotation method, characterized in that, The method includes: Multiple sets of data to be labeled are acquired, and the acquisition time and acquisition range of the multiple sets of data to be labeled are at least partially overlapping. Each set of data to be labeled contains data to be labeled collected by at least one sensor on a corresponding data acquisition vehicle. The data acquisition positions and / or data acquisition perspectives of different data acquisition vehicles are different in the same time and space. Each of the data groups to be labeled is labeled to obtain the corresponding single-group labeling results; Multiple single-group annotation results are fused to obtain corresponding fused annotation results, which include fused temporal tracking results corresponding to multiple dynamic objects respectively. The fusion time-series tracking results of each dynamic object are mapped to a single set of annotation results that meet the mapping conditions to obtain the target annotation results corresponding to each set of data to be annotated.
2. The method according to claim 1, characterized in that, Before labeling each of the data groups to be labeled to obtain the corresponding single-group labeling results, the method further includes: In response to the fact that the data set to be labeled contains multiple data sets to be labeled, the multiple data sets to be labeled are reconstructed in three dimensions to align the multiple data sets to be labeled. The pose graph of the aligned data sets is optimized to align the data sets.
3. The method according to claim 1, characterized in that, Each data set to be labeled contains multiple data frames to be labeled. The step of labeling each of the data groups to be labeled to obtain the corresponding single-group labeling results includes: Each of the data frames to be labeled is input into the target detection model to obtain the corresponding single-frame target detection result, wherein the single-frame target detection result includes multiple dynamic objects; Input the target detection results of multiple single frames corresponding to each of the data to be labeled into the target tracking model to obtain the temporal tracking results of each of the dynamic objects. The temporal tracking results include the bounding boxes corresponding to different time points of the dynamic objects and the point clouds corresponding to the bounding boxes. The single set of annotation results is determined based on the time-series tracking results corresponding to the multiple dynamic objects.
4. The method according to claim 3, characterized in that, Before determining the single set of annotation results based on the temporal tracking results corresponding to the multiple dynamic objects, the method further includes: The temporal tracking results of each dynamic object are input into the object optimization model to correct the shape of the corresponding bounding box based on the point cloud. The orientation of each of the annotation boxes is corrected based on the relative positional relationship between the multiple annotation boxes in each of the time-series tracking results.
5. The method according to claim 1, characterized in that, The step of fusing multiple single-group annotation results to obtain a corresponding fused annotation result includes: Identify multiple temporal tracking results belonging to the same dynamic object from different single-group annotation results; The multiple time-series tracking results are clustered to obtain the clustering results of the corresponding dynamic objects; Based on the collection time of all the labeled boxes in the clustering results, determine the labeled time period corresponding to the clustering results; Multiple collection time points are determined from the labeled time period at preset time intervals; Determine the target bounding box corresponding to each of the aforementioned data collection time points; The clustering results are input into the target tracking model to obtain the corresponding fusion annotation results.
6. The method according to claim 5, characterized in that, The step of clustering the multiple time-series tracking results to obtain the clustering results for the corresponding dynamic objects includes: The clustering result is determined based on the trajectory overlap rate corresponding to multiple time-series tracking results and the overlap rate of the bounding boxes with similar spatiotemporal information. The similar spatiotemporal information refers to the time difference of the bounding boxes being within a preset aggregation time difference threshold and the position difference of the bounding boxes being within a preset position difference threshold.
7. The method according to claim 5 or 6, characterized in that, The time-series tracking results include the predicted category of the corresponding dynamic object; The determination of multiple temporal tracking results belonging to the same dynamic object from different single-group annotation results includes: The trajectory of the corresponding dynamic object is determined based on the time-series tracking results; Based on the trajectory and the predicted category, multiple temporal tracking results belonging to the same dynamic object are determined from different single-set annotation results.
8. The method according to claim 5, characterized in that, Determining the target bounding box corresponding to each of the acquisition time points includes: Determine the alternative time periods corresponding to each of the aforementioned data collection time points; The annotation boxes that fall within the aforementioned candidate time period are identified as candidate annotation boxes; In response to the existence of a candidate label box, the candidate label box is determined as the target label box; In response to the existence of multiple candidate bounding boxes, the candidate bounding box with the highest confidence level is determined as the target bounding box.
9. The method according to claim 8, characterized in that, The step of determining the candidate bounding box with the highest confidence level as the target bounding box in response to the existence of multiple candidate bounding boxes includes: In response to the existence of multiple candidate bounding boxes, a corresponding confidence level is determined based on the detection score, point cloud sparsity, and / or image quality of the candidate bounding boxes; The candidate label box with the highest confidence level is determined as the target label box.
10. The method according to claim 1, characterized in that, The step of mapping the fused temporal tracking results of each of the dynamic objects to a single set of annotation results that satisfy the mapping results, to obtain the target annotation results corresponding to each set of data to be annotated, includes: Dynamic objects that are identical in the fused annotation results and the single set of annotation results are identified as target objects; The bounding boxes corresponding to the target objects in the fusion annotation results are determined as the bounding boxes to be mapped. Based on the time point corresponding to the annotation box to be mapped, the annotation box corresponding to the time point in the single set of annotation results is replaced with the annotation box to be mapped to obtain the target annotation result.
11. The method according to claim 10, characterized in that, The mapping conditions include the presence of point cloud at the corresponding position of the bounding box to be mapped in the single set of annotation results, and / or the bounding box to be mapped is not occluded in the single set of annotation results.
12. A dynamic object labeling device, characterized in that, The device includes: The acquisition module is used to acquire multiple data groups to be labeled. The acquisition time and acquisition range of the multiple data groups to be labeled are at least partially overlapping. The data groups to be labeled contain data to be labeled collected by at least one sensor on the corresponding data acquisition vehicle. The data acquisition positions and / or data acquisition angles of different data acquisition vehicles are different in the same time and space. The annotation module is used to annotate each of the data groups to be annotated, and obtain the corresponding single-group annotation results; The fusion module is used to fuse multiple single-group annotation results to obtain corresponding fused annotation results, which include fused temporal tracking results corresponding to multiple dynamic objects respectively; The mapping module is used to map the fusion time-series tracking results of each dynamic object to a single set of annotation results that meet the mapping conditions, so as to obtain the target annotation results corresponding to each set of data to be annotated.
13. An electronic device comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of claims 1-11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-11.
15. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method as described in any one of claims 1-11.