Image data annotation method fusing pseudo 3D manual annotation and 2D target tracking

By integrating pseudo-3D manual labeling and 2D target tracking, the pseudo-3D annotation frame is derived using YOLOv5 and ByteTrack algorithms, the pseudo-3D annotation frame is solved, and the pseudo-3D manual labeling in the existing technology is low efficiency and high cost, and efficient and low-cost image data labeling is achieved.

CN119942399AActive Publication Date: 2025-05-06SHANGHAI XIHONGQIAO NAVIGATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411864866.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-06
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

In the prior art, pseudo-3D manual labeling is inefficient and costly, making it difficult to meet the needs of massive data labeling.

Method used

An image data labeling method that combines pseudo-3D manual annotation and 2D target tracking is adopted. By extracting frames on video data, using YOLOv5 multi-object detection algorithm and ByteTrack multi-object tracking algorithm, the tracking prediction box and unique index of the target object are determined, and then the pseudo-3D labeling box is derived and the attribute data in the radar point cloud is bound.

Benefits of technology

It effectively improves image labeling efficiency, reduces labeling costs, and reduces the number and complexity of manual labeling through automated means.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942399A_ABST
    Figure CN119942399A_ABST
Patent Text Reader

Abstract

The invention relates to an image data annotation method fusing pseudo 3D manual annotation and 2D target tracking. The method comprises the following steps: acquiring an annotation data video; performing frame extraction on the annotated data video, and setting an extracted image frame as an associated frame and other image frames as derivation frames; performing pseudo 3D manual labeling on a plurality of target objects in the associated frame to obtain a manual labeling box; sequentially identifying target objects in all image frames of the annotated data video by using a target detection model, and obtaining a plurality of target detection frames and confidence degrees thereof; determining a tracking prediction frame and a unique index of each target object by using a multi-target tracking algorithm based on the target detection frame and the confidence thereof; and according to the offset of the centroid of the tracking prediction frame in the associated frame relative to the centroid of the manual annotation frame, determining the associated manual annotation frame, and further deducing the pseudo 3D annotation frame of each deduced frame. The pseudo 3D labeling efficiency can be effectively improved, and the labeling cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image annotation, and in particular to an image data annotation method that integrates pseudo 3D manual annotation and 2D target tracking. Background Art

[0002] In recent years, AI algorithms based on the fusion of radar and visual sensors have played an important role in the field of autonomous driving due to their high environmental perception capabilities, in which the fusion of radar point cloud data and image pseudo 3D data is the key. However, training AI models requires massive data annotation support, and manual pseudo 3D image annotation is inefficient and costly. Summary of the invention

[0003] The technical problem to be solved by the present invention is to provide a method that can effectively improve the efficiency of pseudo 3D annotation and reduce the annotation cost.

[0004] The technical solution adopted by the present invention to solve the technical problem is: to provide an image data annotation method integrating pseudo 3D manual annotation and 2D target tracking, comprising the following steps:

[0005] Get the labeled data video;

[0006] Extract frames from the annotated data video, set the extracted image frames as associated frames, and other image frames as derived frames;

[0007] Perform pseudo 3D manual annotation on several target objects in the associated frames to obtain manual annotation boxes;

[0008] The target detection model is used to sequentially identify the target objects in all image frames of the labeled data video, and multiple target detection frames and their confidence levels are obtained;

[0009] Based on the target detection frame and its confidence, a multi-target tracking algorithm is used to determine a tracking prediction frame and a unique index of each target object;

[0010] The associated manually labeled frames are determined according to the offset of the centroid of each tracking prediction frame in the associated frame relative to the centroid of the manually labeled frame, and then the pseudo 3D labeled frames of each derived frame are derived.

[0011] Furthermore, determining the associated manually labeled frame according to the offset of the centroid of the tracking prediction frame in the associated frame relative to the centroid of the manually labeled frame includes:

[0012] Constructing an association derivation device, wherein the association derivation device includes a first association frame, a second association frame, and all derivation frames between the two association frames;

[0013] The first centroid offset of each tracking prediction frame in the first associated frame and the second associated frame is calculated respectively, and the manually labeled frame associated with the current tracking prediction frame is determined according to the size of the first centroid offset, wherein the first centroid offset is the pixel interval between the centroids of the tracking prediction frame and the manually labeled frame in the same associated frame.

[0014] Further, the step of determining the manually labeled frame associated with the current tracking prediction frame according to the size of the first centroid offset includes:

[0015] The first centroid offset is compared with a set threshold value. If the offset is smaller than the set threshold value, it is considered that the manually marked frame corresponding to the first centroid offset is associated with the current tracking prediction frame.

[0016] Furthermore, deriving the pseudo 3D annotation frame of each derived frame includes:

[0017] Based on the unique index, determine the manually annotated box associated with each tracking prediction box in the derived frame as the associated annotated box;

[0018] The second centroid offset of each tracking prediction frame in the derivation frame is calculated respectively, and the pseudo 3D annotation frame of the current tracking prediction frame is derived using the second centroid offset, where the second centroid offset is the pixel interval between the centroid of the tracking prediction frame in the derivation frame and the associated annotation frame.

[0019] Furthermore, the step of determining, based on the unique index, a manually annotated frame associated with each tracking prediction frame in the derived frame as an associated annotated frame includes:

[0020] Setting a set number of derivation frames close to the first associated frame in the associated derivation device as first derivation frames, and setting the other derivation frames as second derivation frames;

[0021] Setting the associated annotation frame of each tracking prediction frame in the first derived frame to be the manually annotated frame associated with the tracking prediction frame of the first associated frame having the same unique index;

[0022] The associated annotation box of each tracking prediction box in the second derived frame is set to be a manually annotated box associated with the tracking prediction box of the second associated frame having the same unique index.

[0023] Furthermore, the set number is the total number of derived frames in the associated derivation device divided by 2 and then rounded up.

[0024] Furthermore, the method also includes a step of binding the derived pseudo 3D annotation box with the attribute data of the corresponding target object in the radar point cloud.

[0025] Further, the first associated frame and the second associated frame are sequentially adjacent.

[0026] Furthermore, the determining of the tracking prediction frame of each target object by using a multi-target tracking algorithm based on the target detection frame and its confidence is achieved by using a ByteTrack target tracking algorithm to perform trajectory matching on the tracking prediction frames in a set order.

[0027] Furthermore, the ByteTrack target tracking algorithm is used to perform trajectory matching on the tracking prediction box in a set order, including:

[0028] Sort the tracking prediction boxes in each video frame according to the frequency of the corresponding target object appearing in the current video frame scene;

[0029] The ByteTrack target tracking algorithm is used to match and track the sorted tracking prediction box sequence in sequence.

[0030] Furthermore, the target detection model is built based on the YOLO algorithm.

[0031] Beneficial Effects

[0032] Due to the adoption of the above technical scheme, the present invention has the following advantages and positive effects compared with the prior art: the present invention reduces the number of manual annotations by extracting frames from massive data, tracks all detected targets in each frame based on the YOLOv5 multi-target detection algorithm and the ByteTrack multi-target tracking algorithm, associates the tracking results with the manually annotated pseudo 3D frame, derives the pseudo 3D frame, and binds attribute data in other radar point clouds, thereby effectively improving the image annotation efficiency and reducing the annotation cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a schematic diagram of a process module of an embodiment of the present invention;

[0034] Figure 2 is a flow chart of an embodiment of the present invention;

[0035] Figure 3 It is a flow chart of a multi-target tracking algorithm according to an embodiment of the present invention. DETAILED DESCRIPTION

[0036] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall within the scope limited by the appended claims of the application equally.

[0037] The embodiments of the present invention relate to an image data annotation method that integrates pseudo 3D manual annotation and 2D target tracking. Figure 1As shown, it includes the following process modules:

[0038] (1) 3D data processing module: Its main function is to extract frames from massive annotated data to reduce the number of manual annotations. Each image corresponds to a labeling file, including the pseudo 3D information box of the target to be annotated in the current image and other attribute data of the target in the radar point cloud. Secondly, the data video before frame extraction is serialized, and the manually annotated image frames are defined as associated frames. The associated frames are bound to the contents of the manually annotated files, and the other frames are derived frames.

[0039] (2) 2D target detection and tracking module: Its main function is to track all detected targets in each frame based on the YOLOv5 multi-target detection algorithm and the ByteTrack multi-target tracking algorithm. Each tracked target has a unique index.

[0040] (3) The association derivation module has the main function of calculating the centroid of the manually annotated pseudo 3D frame and the centroid of the 2D target tracking frame, determining whether they are associated based on the centroid pixel distance, and calculating the centroid offset between the derived frame and the associated frame, thereby deriving the pseudo 3D frame and binding other attribute data of the target in the radar point cloud.

[0041] (4) Output annotation file module, whose main function is to automatically generate a new annotation file based on the derived frame calculation results.

[0042] The present embodiment will be further described below in conjunction with the accompanying drawings. Figure 2 As shown, the specific steps include:

[0043] Step 1: Serialize the data video before frame extraction, and bind the manual annotation file when the associated frame is found.

[0044] Step 2: Initialize the target detector, load the YOLOv5 multi-target detection model, and return the detection boxes of all targets in the image.

[0045] Step 3: Obtain the target detection frame through the detector, convert the detection frame into the format required by the ByteTrack target tracker, and divide the detection frame into high-confidence detection frame and low-confidence detection frame according to the confidence of the target detection frame.

[0046] Step 4: Use the improved ByteTrack target tracking algorithm to create a tracking track for each target. The algorithm uses Kalman filtering to predict the next bounding box of each tracking track. Since the tracking results returned by the original ByteTrack algorithm do not have target category information, the motion trajectories of targets of the same category have certain correlations. Each Kalman filter prediction is traversed in the order of trajectory discovery, which has no category correlation and will reduce the algorithm speed. Therefore, add target category attributes to each trajectory, and pre-sort the detection boxes according to the frequency of the target object in the corresponding scene. Then use ByteTrack to track the trajectory in the order of category to improve the algorithm speed. The specific process is as follows: Figure 3 shown.

[0047] Step 5: For the high-score confidence detection box, calculate its IOU with the predicted box, use the Hungarian algorithm to match the IOU, and obtain three results: matched trajectory, unmatched trajectory, and unmatched high-score detection box. After matching the trajectory, update the box in the trajectory to the detection box.

[0048] Step 6: For the low-confidence detection box, calculate its IOU with the unmatched trajectory prediction box in the previous step, use the Hungarian algorithm to match the IOU, and obtain three results: matched trajectory, unmatched trajectory, and unmatched low-score detection box. After matching the trajectory, update the box in the trajectory to the detection box.

[0049] Step 7: For the unmatched high-score detection box, match it with the inactive track (new track in track prediction) to obtain three results: match, unmatched track, and unmatched detection box. For the match, update the status, mark the unmatched track as deleted, and for the unmatched detection box, create a new tracking track if the confidence is greater than the threshold, otherwise delete it.

[0050] Step 8: Return the tracking prediction box of each target and the target unique index.

[0051] Step 9: Construct two associated frames and all derived frames in between as an associated derivation. If the frame extraction interval is n, the length of an associated derivation is n+2. Each element includes the current frame, the current frame pseudo-3D manual annotation information, radar point cloud information, and the current frame target tracking information. The current frame pseudo-3D information refers to a total of 16 pseudo-3D coordinates, each coordinate value is the normalized value of the 2D image pixel coordinate, and radar point cloud information such as timestamp information, lane information, direction information, etc.

[0052] Step 10: An association deriver is constructed. The centroid of the tracking prediction box and the centroid of the manually labeled box of each target in the associated frame are calculated respectively, and their pixel intervals are calculated. If the interval is less than the threshold, they are associated, that is, they are the same target.

[0053] Step 11: Derive the pseudo 3D coordinate value based on the offset between the centroid of the tracking frame in the associated frame and the derived frame. If the frame extraction interval is 10, the associated deriver derives the next 5 frames for the first associated frame, and the associated deriver derives the first 5 frames for the last associated frame.

[0054] Step 12: When an association derivation is completed, the derivation annotation file is output, with each target as a line of data, including pseudo 3D coordinates and other attribute data in the radar point cloud, such as timestamp information, lane information, direction information, etc.

[0055] Step 13: Pop out the last frame of associated frame information as the first frame information of the next associated derivation device, continue to append new derivation frames, and repeat the updating of the associated derivation device.

Claims

1. An image data annotation method integrating pseudo 3D manual annotation and 2D target tracking, characterized in that: The following steps are involved: Get the labeled data video; Extract frames from the annotated data video, set the extracted image frames as associated frames, and other image frames as derived frames; Perform pseudo 3D manual annotation on several target objects in the associated frames to obtain manual annotation boxes; The target detection model is used to sequentially identify the target objects in all image frames of the labeled data video, and multiple target detection frames and their confidence levels are obtained; Based on the target detection frame and its confidence, a multi-target tracking algorithm is used to determine a tracking prediction frame and a unique index of each target object; According to the offset of the centroid of the tracking prediction box in the associated frame relative to the centroid of the manually labeled box, the associated manually labeled box is determined, and then the pseudo 3D labeled boxes of each derived frame are derived.

2. The method according to claim 1, characterized in that The step of determining the associated manually labeled frame according to the offset of the centroid of the tracking prediction frame in the associated frame relative to the centroid of the manually labeled frame includes: Constructing an association derivation device, wherein the association derivation device includes a first association frame, a second association frame, and all derivation frames between the two association frames; The first centroid offset of each tracking prediction frame in the first associated frame and the second associated frame is calculated respectively, and the manually labeled frame associated with the current tracking prediction frame is determined according to the size of the first centroid offset, wherein the first centroid offset is the pixel interval between the centroids of the tracking prediction frame and the manually labeled frame in the same associated frame.

3. The method according to claim 2, characterized in that The step of determining the manually labeled frame associated with the current tracking prediction frame according to the magnitude of the first centroid offset includes: The first centroid offset is compared with a set threshold value. If the offset is smaller than the set threshold value, it is considered that the manually marked frame corresponding to the first centroid offset is associated with the current tracking prediction frame.

4. The method according to claim 2, characterized in that: The deriving of the pseudo 3D annotation frame of each derived frame includes: Based on the unique index, the manually labeled frame associated with each tracking prediction frame in the derived frame is determined as the associated labeled frame; the second center of mass offset of each tracking prediction frame in the derived frame is calculated respectively, and the pseudo 3D labeled frame of the current tracking prediction frame is derived using the second center of mass offset, wherein the second center of mass offset is the pixel interval between the center of mass of the tracking prediction frame in the derived frame and the associated labeled frame.

5. The method according to claim 4, characterized in that The step of determining, based on the unique index, a manually labeled frame associated with each tracking prediction frame in the derived frame as an associated labeled frame includes: Setting a set number of derivation frames close to the first associated frame in the associated derivation device as first derivation frames, and setting the other derivation frames as second derivation frames; Setting the associated annotation frame of each tracking prediction frame in the first derived frame to be the manually annotated frame associated with the tracking prediction frame of the first associated frame having the same unique index; The associated annotation box of each tracking prediction box in the second derived frame is set to be a manually annotated box associated with the tracking prediction box of the second associated frame having the same unique index.

6. The method according to claim 5, characterized in that The set number is the total number of derived frames in the associated derivation unit divided by 2 and then rounded to the nearest integer.

7. The method according to claim 1, characterized in that The method also includes a step of binding the derived pseudo 3D annotation box to the attribute data of the corresponding target object in the radar point cloud.

8. The method according to claim 1, characterized in that: The first associated frame and the second associated frame are sequentially adjacent.

9. The method according to claim 1, characterized in that: The method of determining the tracking prediction frame of each target object based on the target detection frame and its confidence using a multi-target tracking algorithm is implemented by matching the tracking prediction frames with ByteTrack target tracking algorithm in a set order.

10. The method according to claim 1, characterized in that The ByteTrack target tracking algorithm is used to match the track prediction box according to the set order, including: Sort the tracking prediction boxes in each video frame according to the frequency of the corresponding target object appearing in the current video frame scene; The ByteTrack target tracking algorithm is used to match and track the sorted tracking prediction box sequence in sequence.

Citation Information

Patent Citations

  • Video semi-automatic target labeling method integrating target detection and tracking

    CN110929560A

  • Live pig multi-target tracking and behavior recognition method, computer equipment and storage medium

    CN115830078A

  • Classroom behavior dynamic labeling method and device based on dense character detection and medium

    CN118135453A

  • Method for generating improved training data video and apparatus thereof

    US20240282072A1

  • Multi-object tracking method and apparatus, computer device, and storage medium

    WO2022135027A1