Data annotation method, device, equipment and storage medium

The fusion and cross-verification integration of data annotation through the automatic annotation model solves the problem that existing annotation tasks cannot verify complementary each other, and improves the integrity of the data set and the accuracy of the labeling.

CN115311512BActive Publication Date: 2025-06-24SHANGHAI WERIDE AUTOMOTIVE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210750667.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-28
Publication Date
2025-06-24
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

Existing labeling tasks are not validated separately and complement each other, resulting in different data set sizes and different coverage ranges.

Method used

By entering the preset automatic labeling model to be labeled, perform fusion and completion between labels, determine the labeling data with redundant labels, and perform cross-verification integration to generate a complete data set.

Benefits of technology

It realizes mutual verification and completion between different annotated data sets, improves the integrity and consistency of the data set, and enhances the accuracy and coverage of the annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115311512B_ABST
    Figure CN115311512B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence, and discloses a data annotation method, device, equipment and storage medium. The method includes: inputting a dataset to be annotated into a preset automatic annotation model to obtain a first dataset; performing fusion and complementation between the annotation data of the same frame in the first dataset to obtain a second dataset; determining the annotation data with redundant annotations in the second dataset, and performing cross-validation integration on the second dataset according to the annotation data with redundant annotations to obtain a complemented dataset. In the technical solution of this application, for the way that the annotation data is correlated, point cloud and image data are acquired through a collector, multi-modal annotation cross-validation is used to enrich semantic information, generate temporal information such as target speed and acceleration, and fuse the multi-modal annotations, enriching the annotation semantic information while smoothing the annotations and increasing the annotation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular, to a data annotation method, apparatus, device, and storage medium. Background Art

[0002] The perception tasks of driverless vehicles usually use multiple sensor inputs, such as images, laser point clouds, etc. After these input information are annotated, they are used for the training of downstream object detection, semantic segmentation and other models to complete various goals of the perception module, such as obstacle detection, object tracking, speed estimation, drivable area prediction, etc. Due to the limitations of data annotation requirements and annotation costs, different tasks will perform single-frame annotation for relevant valuable data, and the annotation specifications are also different, resulting in different sizes of data sets with different annotations and different coverage ranges. Summary of the Invention

[0003] The main purpose of the present invention is to solve the technical problem that the existing annotation tasks cannot be mutually verified and complemented when performed separately.

[0004] The first aspect of the present invention provides a data annotation method, which includes: inputting a data set to be annotated into a preset automatic annotation model to obtain a first data set, where the first data set includes at least one frame of annotated data, and the annotated data contains one or more annotations; performing fusion and complementation between the annotations of the annotated data of the same frame in the first data set to obtain a second data set; determining the annotated data with redundant annotations in the second data set, and performing cross-verification integration on the second data set according to the annotated data with redundant annotations to obtain a complemented data set.

[0005] Optionally, in the first implementation manner of the first aspect of the present invention, the annotation includes object detection annotation, and the object detection annotation includes point cloud object detection annotation and image object detection annotation. The performing fusion and complementation between the annotations of the annotated data of the same frame in the first data set to obtain a second data set includes: determining whether both the point cloud object detection annotation and the image object detection annotation exist in the annotated data of the same frame; if so, converting the point cloud object detection annotation included in the annotated data of the same frame into a 2D detection box based on a preset coordinate system; determining whether the overlap rate between the 2D detection box and the image object detection annotation is higher than a preset threshold; if it is higher, marking the object detected by the point cloud object detection annotation and the target detected by the image object detection annotation as the same target to obtain the second data set.

[0006] Optionally, in the second implementation manner of the first aspect of the present invention, the coordinate system includes a three-dimensional coordinate system and a planar coordinate system; the conversion of the point cloud object detection annotation included in the annotation data of the same frame into a 2D detection box based on the preset coordinate system includes: obtaining the camera pose information, lidar pose information, and one or more 3D detection boxes in the point cloud object detection annotation corresponding to each frame of annotation data in the first dataset; constructing a stereo-plane mapping relationship between the three-dimensional coordinate system and the planar coordinate system in the same frame according to the camera pose information and the lidar pose information; and converting the 3D detection box into the corresponding 2D detection box according to the stereo-plane mapping relationship.

[0007] Optionally, in the third implementation manner of the first aspect of the present invention, the annotation further includes a semantic segmentation annotation, and the semantic segmentation annotation at least includes a point cloud semantic segmentation annotation, a point cloud semantic map annotation, and an image semantic segmentation annotation. The step of fusing and complementing the annotations in the annotation data of the same frame in the first dataset to obtain a second dataset further includes: determining whether at least two of the point cloud semantic segmentation annotation, the point cloud semantic map annotation, and the image semantic segmentation annotation exist simultaneously in the annotation data of the same frame; if the point cloud semantic segmentation annotation and the image semantic segmentation annotation exist simultaneously, constructing a mapping relationship between the point cloud laser points corresponding to the point cloud semantic segmentation annotation and the camera pixels corresponding to the image semantic segmentation annotation according to the stereo-plane mapping relationship; performing annotation complementation between the point cloud laser points and the camera pixels according to the mapping relationship to obtain a second dataset; if the point cloud semantic segmentation annotation and the point cloud semantic map annotation exist simultaneously, optimizing the semantic map annotation according to the point cloud semantic segmentation annotation to obtain a second dataset.

[0008] Optionally, in the fourth implementation manner of the first aspect of the present invention, the annotation data further includes a motion annotation. Before determining the annotation data with redundant annotations in the second dataset and cross-validating and integrating the second dataset according to the annotation data with redundant annotations to obtain a complemented dataset, it includes: constructing, frame by frame, temporal information corresponding to the target for the first dataset including at least the point cloud object detection annotation and the image object detection annotation; fusing the temporal information frame by frame to obtain a discrete trajectory of the target; and obtaining the motion annotation of the target by performing differential calculation on the discrete trajectory.

[0009] Optionally, in the fifth implementation manner of the first aspect of the present invention, determining the annotation data with redundant annotations in the second data set, and cross-validating and integrating the second data set according to the annotation data with redundant annotations to obtain a complemented data set, includes: cross-validating the annotations in the first data set based on the connection relationship and annotation specifications between upper and lower frames to obtain point cloud class verification information and image class verification information; smoothing the annotation data in consecutive frames by using a weighted algorithm through the point cloud class verification information and the image class verification information to obtain processed annotations; traversing frame by frame the missing annotation data in the second data set by using the processed annotations after smoothing to obtain a complemented data set.

[0010] Optionally, in the sixth implementation manner of the first aspect of the present invention, after determining the annotation data with redundant annotations in the second data set, and cross-validating and integrating the second data set according to the annotation data with redundant annotations to obtain a complemented data set, it includes: collecting a preprocessed sample data set, inputting the sample data set with some sample semantic segmentation annotations and / or sample object detection annotations and / or sample object tracking annotations randomly masked into the to-be-trained automatic annotation model to obtain prediction annotation data of the to-be-trained screening model, where the prediction annotation data includes prediction semantic segmentation annotations and / or prediction object detection annotations and / or prediction object tracking annotations, and the sample data set further includes the complemented data set obtained by executing the complementing method this time; the to-be-trained automatic annotation model further includes the automatic annotation model obtained by executing the complementing method this time; calculating a preset semantic segmentation cross-entropy loss function based on the prediction semantic segmentation annotations and the sample semantic segmentation annotations to obtain a semantic segmentation cross-entropy loss function value; calculating a preset object detection cross-entropy loss function based on the prediction object detection annotations and / or the prediction object tracking annotations and the sample object detection annotations and / or the sample object tracking annotations to obtain an object detection cross-entropy loss function value; determining whether the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value is less than a preset threshold; if not, adjusting the model parameters of the to-be-trained automatic annotation model based on the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value, and re-training the model with the sample data set with some sample semantic segmentation annotations and / or sample object detection annotations and / or sample object tracking annotations randomly masked until the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value is less than the preset threshold; if so, updating the to-be-trained screening model to an automatic annotation model.

[0011] In a second aspect of the present invention, a data annotation device is provided, including: an automatic annotation module for inputting a dataset to be annotated into a preset automatic annotation model to obtain a first dataset, where the first dataset includes at least one frame of annotated data, and the annotated data contains one or more annotations; an annotation fusion and completion module for fusing and completing the annotations among the annotated data of the same frame in the first dataset to obtain a second dataset; an annotation verification and integration module for determining the annotated data with redundant annotations in the second dataset and cross-verifying and integrating the second dataset according to the annotated data with redundant annotations to obtain a completed dataset.

[0012] Optionally, in a first implementation manner of the second aspect of the present invention, the annotation fusion and completion module is specifically configured to: a target detection judgment unit for judging whether both point cloud target detection annotations and image target detection annotations exist in the annotated data of the same frame; a plane detection box conversion unit for, if so, converting the point cloud target detection annotations included in the annotated data of the same frame into 2D detection boxes based on a preset coordinate system; an overlap rate judgment unit for judging whether the overlap rate between the 2D detection box and the image target detection annotation is higher than a preset threshold; a second dataset acquisition unit for, if higher, marking the objects detected by the point cloud target detection annotations and the targets detected by the image target detection annotations as the same target to obtain the second dataset.

[0013] Optionally, in a second implementation manner of the second aspect of the present invention, the plane detection box conversion unit is specifically configured to: obtain one or more 3D detection boxes corresponding to each frame of annotated data in the first dataset, the camera pose information, the lidar pose information, and the point cloud target detection annotations; construct a stereo-plane mapping relationship between the three-dimensional coordinate system and the plane coordinate system in the same frame according to the camera pose information and the lidar pose information; and convert the 3D detection box into the corresponding 2D detection box according to the stereo-plane mapping relationship.

[0014] Optionally, in a third implementation manner of the second aspect of the present invention, the annotation fusion and completion module is further specifically configured to: judge whether at least two of point cloud semantic segmentation annotations or point cloud semantic map annotations or image semantic segmentation annotations exist simultaneously in the annotated data of the same frame; if both point cloud semantic segmentation annotations and image semantic segmentation annotations exist simultaneously, construct a mapping relationship between the point cloud laser points corresponding to the point cloud semantic segmentation annotations and the camera pixels corresponding to the image semantic segmentation annotations according to the stereo-plane mapping relationship; perform annotation completion between the point cloud laser points and the camera pixels according to the mapping relationship to obtain the second dataset; if both point cloud semantic segmentation annotations and point cloud semantic map annotations exist simultaneously, optimize the semantic map annotations according to the point cloud semantic segmentation annotations to obtain the second dataset.

[0015] Optionally, in the fourth implementation manner of the second aspect of the present invention, the data annotation device further includes a motion annotation acquisition module, and the motion annotation acquisition module is specifically configured to: construct, frame by frame, the time series information corresponding to the target for the first data set including at least the point cloud target detection annotation and the image target detection annotation; combine the frame-by-frame time series information to fuse and obtain the discrete trajectory of the target; and obtain the motion annotation of the target by performing differential calculation on the discrete trajectory.

[0016] Optionally, in the fifth implementation manner of the second aspect of the present invention, the annotation verification and integration module is specifically configured to: perform cross-verification on the annotations in the first data set based on the connection relationship and annotation specifications between adjacent frames to obtain point cloud type verification information and image type verification information; perform smoothing processing on the annotation data in consecutive frames by adopting a weighted algorithm through the point cloud type verification information and the image type verification information to obtain processed annotations; and traverse and complete the missing annotation data in the second data set frame by frame by adopting the processed annotations after smoothing processing to obtain a completed data set.

[0017] Optionally, in the sixth implementation manner of the second aspect of the present invention, the data annotation device further includes a model training module, and the model training module is specifically configured to: collect the preprocessed sample data set, input the sample data set with some sample semantic segmentation annotations and / or sample object detection annotations and / or sample object tracking annotations randomly masked into the automatic annotation model to be trained, and obtain the predicted annotation data of the screening model to be trained, where the predicted annotation data includes predicted semantic segmentation annotations and / or predicted object detection annotations and / or predicted object tracking annotations, and the sample data set further includes the completion data set after the completion method is executed for the current time; the automatic annotation model to be trained further includes the automatic annotation model after the completion method is executed for the current time; calculate a preset semantic segmentation cross-entropy loss function based on the predicted semantic segmentation annotations and the sample semantic segmentation annotations to obtain a semantic segmentation cross-entropy loss function value; calculate a preset object detection cross-entropy loss function based on the predicted object detection annotations and / or the predicted object tracking annotations and the sample object detection annotations and / or the sample object tracking annotations to obtain an object detection cross-entropy loss function value; determine whether the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value is less than a preset threshold; if not, adjust the model parameters of the automatic annotation model to be trained based on the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value, and re-train the model with the sample data set with some sample semantic segmentation annotations and / or sample object detection annotations and / or sample object tracking annotations randomly masked until the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value is less than the preset threshold; if so, update the screening model to be trained to an automatic annotation model.

[0018] The third aspect of the present invention provides a data annotation device, including: a memory and at least one processor, where instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line; the at least one processor invokes the instructions in the memory to cause the data annotation device to execute the steps of the above data annotation method.

[0019] The fourth aspect of the present invention provides a computer-readable storage medium, in which instructions are stored, and when the instructions run on a computer, the computer is caused to execute the steps of the above data annotation method.

[0020] In the technical solution of the present invention, an unlabeled data set is input into a preset automatic labeling model to obtain a first data set. Among them, the first data set includes at least one frame of labeled data, and the labeled data contains one or more labels; the labels in the labeled data of the same frame in the first data set are fused and complemented to obtain a second data set; the labeled data with redundant labels in the second data set is determined, and the second data set is cross-validated and integrated according to the labeled data with redundant labels to obtain a complemented data set. In this method, for the way that the labeled data is correlated, point cloud and image data are acquired by a collector, multi-modal labeling cross-validation is used to enrich semantic information, generate temporal information such as target speed and acceleration, and fuse multi-modal labels, enriching the labeled semantic information while smoothing the labels and increasing the labeling accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 Schematic diagram of the first embodiment of the data labeling method in the embodiment of the present invention;

[0022] Figure 2 Schematic diagram of the second embodiment of the data labeling method in the embodiment of the present invention;

[0023] Figure 3 Schematic diagram of the third embodiment of the data labeling method in the embodiment of the present invention;

[0024] Figure 4 Schematic diagram of the fourth embodiment of the data labeling method in the embodiment of the present invention;

[0025] Figure 5 Schematic diagram of an embodiment of the data labeling device in the embodiment of the present invention;

[0026] Figure 6 Schematic diagram of another embodiment of the data labeling device in the embodiment of the present invention;

[0027] Figure 7 Schematic diagram of an embodiment of the data labeling device in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] In the technical solution of the present invention, an unlabeled data set is input into a preset automatic labeling model to obtain a first data set. Among them, the first data set includes at least one frame of labeled data, and the labeled data contains one or more labels; the labels of the same frame of labeled data in the first data set are fused and complemented to obtain a second data set; the labeled data with redundant labels in the second data set is determined, and the second data set is cross-validated and integrated according to the labeled data with redundant labels to obtain a complemented data set. In this method, for the way that the labeled data is correlated, point cloud and image data are acquired by a collector, multi-modal labeling cross-validation is used to enrich semantic information, generate temporal information such as target speed and acceleration, and fuse multi-modal labels, enriching the labeled semantic information while smoothing the labels and increasing the labeling accuracy.

[0029] In the description and claims of the present invention and the above drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" or "having" and any variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0030] For ease of understanding, the specific process of the embodiment of the present invention is described below. Please refer to Figure 1 , the first embodiment of the data annotation method in the embodiment of the present invention includes:

[0031] 101. Input the unlabeled data set into a preset automatic labeling model to obtain a first data set;

[0032] In this embodiment, the unlabeled data set includes at least one frame of labeled data. The labeled data includes original labeled data and automatically labeled data. The original labeled data includes at least any one of point cloud object detection labeling, point cloud object tracking labeling, point cloud semantic segmentation labeling, point cloud semantic map labeling, image object detection labeling, image object tracking labeling, and image semantic segmentation labeling;

[0033] Specifically, before inputting the unlabeled data set into the preset automatic labeling model, the unlabeled data set should be screened so that the unlabeled data set must include lidar pose, lidar point cloud, camera pose, and camera image.

[0034] Specifically, the automatic annotation model is a neural network model that has been trained. Since there are no requirements for the time of automatic annotation and multi-task balance, the main optimization point is biased towards the prediction accuracy of the automatic annotation model. For example, for the 2D object detection annotation of the dataset to be annotated, after performing object detection on the original image, its output can be used as the supervision signal of the model together with the image object detection annotation, and can be directly used for model training or for subsequent fusion and completion processes. The reliability of the automatic annotation can be obtained through the uncertainty prediction of the model or through manual screening of the input data scenarios.

[0035] 102. Fuse and complete the annotation data of the same frame in the first dataset to obtain the second dataset.

[0036] In this embodiment, the annotation data is fused and completed by fusing and completing semantic segmentation information, fusing and completing object detection information, and using temporal information to complete motion annotation.

[0037] Specifically, for the speed annotation of the target, for example, on the object detection information, by calculating the displacement distance of the target within a unit time and fusing it in two ways of point cloud and image, the speed annotation of the target in any frame can be obtained. Among them, in the fusion and completion of object detection information, due to different data sources (image and point cloud), the missing and incorrect data of either party can be corrected through mutual verification between the two to obtain accurate speed annotation. Similarly, in the process of fusing and completing semantic segmentation information, fusing and completing object detection information, and using temporal information to complete, a large amount of redundant data will be generated for cross-verification, so as to obtain more accurate annotation data.

[0038] 103. Determine the annotation data with redundant annotations in the second dataset, and cross-verify and integrate the second dataset according to the annotation data with redundant annotations to obtain the completed dataset.

[0039] In this embodiment, there is redundant information between the annotation data obtained from the second dataset obtained by the automatic annotation model, the annotation data after fusion and completion, and the original annotation data in the dataset to be annotated, which can be used for cross-verification between consecutive frames and different annotations. Among them, redundant annotations are the duplicate or incorrect annotation data generated during the fusion and completion process. Among them, the error can be a random error value based on the original dataset. For example, the speed information set collected by the speed sensor obtains a data of 100 km / h in a certain frame, but considering the information of the upper and lower frames such as 60 km / h, the incorrect data can be cross-verified and completed.

[0040] Specifically, in terms of object detection information, in the annotations of the same object in different frames within the second dataset, there are both category information of point cloud annotations and category information of image annotations. By weighting the category information of point cloud annotations and the category information of image annotations, the annotations between different frames are smoothed, and possible incorrect annotations or missing annotations are corrected and complemented. For example, in a sequence of consecutive frames, an object has an image annotation in a certain frame, but in some frames, due to occlusion or other reasons, detailed semantic information annotations are lost, or due to blurriness, incorrect annotations occur. The situation of missing or incorrect annotations in some frames can be corrected through cross-validation integration.

[0041] Among them, in terms of object detection information, the number of objects is not limited. Here, only a single object is described for illustration. Similarly, it can be applied to one or more objects for the above correction process.

[0042] In this embodiment, the dataset to be annotated is input into a preset automatic annotation model to obtain a first dataset. Among them, the first dataset includes at least one frame of annotation data, and the annotation data contains one or more annotations; the annotation data of the same frame in the first dataset is fused and complemented between the annotations to obtain a second dataset; the annotation data with redundant annotations in the second dataset is determined, and the second dataset is cross-validated and integrated according to the annotation data with redundant annotations to obtain a complemented dataset. In this method, for the way that the annotation data is correlated, point cloud and image data are acquired through a collector, multi-modal annotation cross-validation is used to enrich semantic information, temporal information such as object speed and acceleration is generated, and multi-modal annotations are fused, enriching the annotation semantic information while smoothing the annotations and increasing the annotation accuracy.

[0043] Please refer to Figure 2 , the second embodiment of the data annotation method in the embodiment of the present invention includes:

[0044] 201. Input the dataset to be annotated into a preset automatic annotation model to obtain a first dataset;

[0045] 202. Determine whether there are both point cloud object detection annotations and image object detection annotations in the annotation data of the same frame;

[0046] 203. If so, obtain one or more 3D detection frames corresponding to the camera pose information, lidar pose information, and point cloud object detection annotations in each frame of annotation data in the first dataset;

[0047] 204. According to the camera pose information and lidar pose information, construct a stereo-plane mapping relationship between the three-dimensional coordinate system and the plane coordinate system in the same frame;

[0048] In this embodiment, by extracting the three-dimensional coordinates of the 3D detection box for the target in the point cloud target detection annotation in the point cloud target detection information, the 2D detection box of the target in the image plane is obtained through the mapping formula between coordinate points.

[0049] Specifically, the mapping formula between coordinates is as follows:

[0050] x max = max(p 1x , p 2x ,... p 8x )

[0051] x min = min(p 1x , p 2x ,... p 8x )

[0052] y max = max(p 1y , p 2y ,... p 8y )

[0053] y min = min(p 1y , p 2y ,... p 8y )

[0054] The 2D detection box coordinates mapped from the 3D detection box constructed for the target based on the point cloud are obtained through the above mapping formula.

[0055] 205. Convert the 3D detection box into the corresponding 2D detection box according to the stereo plane mapping relationship;

[0056] 206. Determine whether the overlap rate between the 2D detection box and the image target detection annotation is higher than the preset threshold;

[0057] In this embodiment, the calculation of the overlap rate can be by calculating the coordinates of the mapped coordinates and the corresponding target in the image target detection annotation, calculating the difference between the corresponding coordinate points, and dividing the preset overlap rate threshold that is allowed to exist; or by calculating the ratio of the overlapping area between the 2D detection box and the image target detection annotation and the ratio of the overlapping area between the image target detection annotation and the 2D detection box, and verifying the ratio between the two to prevent the special situation where one party is included in the other party, and any one party fails to frame the complete target or the framed area is much larger than the target area.

[0058] 207. If it is higher, mark the object detected by the point cloud target detection annotation and the target detected by the image target detection annotation as the same target to obtain the second data set;

[0059] In this embodiment, when the calculated overlap rate is higher than the preset threshold, it can be determined that the target marked by the point cloud object detection annotation and the target marked by the image object detection annotation are the same target. In subsequent steps, the point cloud object detection annotation and the image object detection annotation will be shared for this same target.

[0060] 208. Determine the annotation data with redundant annotations in the second dataset, and cross-validate and integrate the second dataset based on the annotation data with redundant annotations to obtain a complemented dataset.

[0061] Based on the previous embodiment, this embodiment details the process of determining whether there are both point cloud object detection annotations and image object detection annotations in the annotation data of the same frame; if so, convert the point cloud object detection annotations included in the annotation data of the same frame into 2D detection frames based on a preset coordinate system; determine whether the overlap rate between the 2D detection frame and the image object detection annotation is higher than the preset threshold; if it is higher, mark the object detected by the point cloud object detection annotation and the target detected by the image object detection annotation as the same target to obtain the second dataset. Compared with traditional methods through this embodiment, by calculating the differences between corresponding coordinate points, a preset overlap rate threshold that allows for existence is divided; or by calculating the ratio of the overlapping area between the 2D detection frame and the image object detection annotation and the ratio of the overlapping area between the image object detection annotation and the 2D detection frame, and by verifying the ratio between the two, special situations where one party is included in the other party and either party fails to enclose the complete target or the enclosed area is much larger than the target area are prevented.

[0062] Please refer to Figure 3 , the third embodiment of the data annotation method in the embodiments of the present invention includes:

[0063] 301. Input the dataset to be annotated into a preset automatic annotation model to obtain a first dataset;

[0064] 302. Determine whether there are at least two of the point cloud semantic segmentation annotation, the point cloud semantic map annotation, or the image semantic segmentation annotation in the annotation data of the same frame;

[0065] In this embodiment, when any two of the point cloud semantic segmentation annotation, the point cloud semantic map annotation, or the image semantic segmentation annotation exist simultaneously in any one frame of the first dataset, the projection relationship between the coordinate system of the point cloud and the coordinate system of the image can be obtained by adopting the camera pose and the lidar pose corresponding to the current frame. The specific projection relationship formula is as follows:

[0066] pt(x,y,z)=img(u,v)

[0067] 303. If there are both point cloud semantic segmentation annotations and image semantic segmentation annotations, a mapping relationship between the point cloud laser points corresponding to the point cloud semantic segmentation annotations and the camera pixels corresponding to the image semantic segmentation annotations is constructed according to the stereo plane mapping relationship;

[0068] Specifically, the image pixel coordinates (u, v) closest to the projection on the camera plane can be calculated through the position (x, y, z) of each laser point in the 3D space. Through this corresponding relationship, the laser points and pixels can be matched, and the semantic segmentation annotation of the corresponding image pixel point (u, v) is assigned to the corresponding laser point. Similarly, the point cloud semantic segmentation annotation can be optimized through semantic map annotation. For example, the point cloud segmentation result of the lane / sidewalk is optimized through the curb annotation.

[0069] After fusion, starting from the ego-vehicle in the top view, rays are emitted at certain angular intervals for scanning. When a moving object bounding box is encountered, it is set as an occluded area, such as pedestrians / vehicles, etc. In other cases, when an un-drivable area or a static object annotation in the semantic segmentation annotation is encountered, it is set as an un-drivable area, such as curbs, traffic cones, etc. Among them, the semantic segmentation annotation includes the point cloud semantic segmentation annotation and the image semantic segmentation annotation.

[0070] 304. According to the mapping relationship, the annotation between the point cloud laser points and the camera pixels is complemented to obtain a second data set;

[0071] Specifically, if any target is jointly confirmed in the point cloud object detection annotation and the image object detection annotation, the point cloud object detection annotation and the image object detection annotation corresponding to the confirmed target can be used to complement the annotation of the confirmed target with the confirmed target as the connection relationship.

[0072] 305. If there are both point cloud semantic segmentation annotations and point cloud semantic map annotations, the semantic map annotation is optimized according to the point cloud semantic segmentation annotation to obtain a second data set;

[0073] 306. The annotations in the first data set are cross-validated based on the connection relationship and annotation specifications between the upper and lower frames to obtain point cloud class verification information and image class verification information;

[0074] Specifically, using frame-by-frame annotation and timestamps, the kinematic information (speed, acceleration, etc.) of the object is calculated as supervised data. Through the connection relationship between frames and the supplement of timestamps, the obtained point cloud class verification information and image class verification information can be used to complement the missing annotation data.

[0075] 307. The annotation data in the continuous frames are smoothed by adopting a weighted algorithm through the point cloud class verification information and the image class verification information to obtain processed annotations;

[0076] 308. By traversing the processed annotations after smoothing frame by frame to complete the missing annotation data in the second dataset, a completed dataset is obtained.

[0077] Specifically, for example, there are 10 consecutive frames, and the annotations of the moving object's speed in each frame are {60, 61, 59, 0, 62, 0, 0, 60, 58, 60} respectively. Here, 0 can mean that the speed annotation corresponding to the moving object in the corresponding frame is missing. By smoothing this set of data, the processed annotation obtained can be {60, 60, 60, 60, 60, 60, 60, 60, 60, 60}.

[0078] Based on the previous embodiment, this embodiment details the process of cross - validating the annotations in the first dataset based on the connection relationship and annotation specifications between adjacent frames to obtain point - cloud - type verification information and image - type verification information; using the weighted algorithm with the point - cloud - type verification information and the image - type verification information to smooth the annotation data in consecutive frames to obtain processed annotations; and traversing frame by frame with the processed annotations after smoothing to complete the missing annotation data in the second dataset to obtain a completed dataset. Compared with the traditional method through this embodiment, the category information of point - cloud annotations and image annotations existing in the annotations of the same target in different frames is refined, and weighting can be used to smooth and unify the annotations, and at the same time, the specific method for correcting possible incorrect annotations is provided.

[0079] Please refer to Figure 4 , the fourth embodiment of the data annotation method in the embodiment of the present invention includes:

[0080] 401. Input the dataset to be annotated into a preset automatic annotation model to obtain a first dataset;

[0081] 402. Perform fusion and completion between the annotation data of the same frame in the first dataset to obtain a second dataset;

[0082] 403. Frame by frame, construct the temporal information corresponding to the target for the first dataset that at least includes point - cloud object - detection annotations and image object - detection annotations;

[0083] In this embodiment, when there are multiple frames of temporally related annotations, the annotation of the object's motion state is completed through modeling. Using a matching algorithm or existing target - tracking annotation data, the temporal information (x, y, z, t) of the same object in each frame, that is, the spatial coordinates of the object at time t, can be obtained. The multi - frame information can be fused by means of linear interpolation, etc. to obtain the discrete trajectory of the object. By performing differential calculation on the trajectory, the motion information such as the speed, angular velocity, and acceleration of the object can be obtained.

[0084] 404. Combine the frame-by-frame timing information to fuse and obtain the discrete trajectory of the target;

[0085] In this embodiment, through the timing information between frames, the target data within each frame can be linked together to obtain a discrete data trajectory including time, that is, the discrete trajectory for any target.

[0086] On the other hand, to complete the target motion trajectory, not only linear interpolation can be used, but also a vehicle model can be added for calculation and fusion to obtain the discrete trajectory of the target.

[0087] Specifically, the linear interpolation method means that if there is a linear relationship between two quantities, if A(X1, Y1) and B(X2, Y2) are two points on this straight line, and the Y0 value of another point P is known, then the corresponding value X0 of point P can be obtained using their linear relationship.

[0088] 405. Through differential calculation of the discrete trajectory, obtain the motion annotation of the target;

[0089] In this embodiment, due to factors such as errors and interferences on the device, the obtained frame-by-frame annotation data shows a discrete state of the target trajectory information. By performing differential operations on the obtained discrete trajectory, and without affecting the effectiveness of the results, it is converted into a calculation method in an ideal state to obtain approximate and smooth values.

[0090] 406. Determine the annotation data with redundant annotations in the second dataset, and perform cross-validation integration on the second dataset according to the annotation data with redundant annotations to obtain a completed dataset;

[0091] 407. Collect the preprocessed sample dataset, randomly mask some sample semantic segmentation annotations and / or sample object detection annotations and / or sample object tracking annotations of the sample dataset and input them into the automatic annotation model to be trained to obtain the predicted annotation data of the screening model to be trained;

[0092] 408. Calculate the preset semantic segmentation cross-entropy loss function based on the predicted semantic segmentation annotation and the sample semantic segmentation annotation to obtain the semantic segmentation cross-entropy loss function value;

[0093] In this embodiment, the calculation formula of the object detection cross-entropy loss function is:

[0094]

[0095] 409. Calculate the preset object detection cross-entropy loss function based on the predicted object detection annotation and / or predicted object tracking annotation and the sample object detection annotation and / or sample object tracking annotation to obtain the object detection cross-entropy loss function value;

[0096] In this embodiment, the calculation formula of the semantic segmentation cross-entropy loss function is as follows:

[0097]

[0098] 410. Determine whether the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value is less than a preset threshold;

[0099] 411. If not, adjust the model parameters of the to-be-trained automatic annotation model based on the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value, and re-train the model with a sample data set in which part of the sample semantic segmentation annotation and / or sample object detection annotation and / or sample object tracking annotation are randomly masked until the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value is less than the preset threshold;

[0100] 412. If so, update the to-be-trained screening model to an automatic annotation model.

[0101] Based on the previous embodiment, this embodiment details the process of collecting the preprocessed sample data set, inputting the sample data set in which part of the sample semantic segmentation annotation and / or sample object detection annotation and / or sample object tracking annotation are randomly masked into the to-be-trained automatic annotation model to obtain the predicted annotation data of the to-be-trained screening model, where the predicted annotation data includes predicted semantic segmentation annotation and / or predicted object detection annotation and / or predicted object tracking annotation, and the sample data set also includes the completion data set after the completion method is executed for the current time; the to-be-trained automatic annotation model also includes the automatic annotation model after the completion method is executed for the current time; calculate the preset semantic segmentation cross-entropy loss function based on the predicted semantic segmentation annotation and the sample semantic segmentation annotation to obtain the semantic segmentation cross-entropy loss function value; calculate the preset object detection cross-entropy loss function based on the predicted object detection annotation and / or the predicted object tracking annotation and the sample object detection annotation and / or the sample object tracking annotation to obtain the object detection cross-entropy loss function value; determine whether the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value is less than the preset threshold; if not, adjust the model parameters of the to-be-trained automatic annotation model based on the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value, and re-train the model with the sample data set in which part of the sample semantic segmentation annotation and / or sample object detection annotation and / or sample object tracking annotation are randomly masked until the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value is less than the preset threshold; if so, update the to-be-trained screening model to an automatic annotation model. By clarifying the training process of the automatic annotation model, the specific details of adjusting the loss function during training are clarified, and the accuracy of the model is improved.

[0102] The data annotation method in the embodiments of the present invention has been described above. Next, the data annotation device in the embodiments of the present invention will be described. Please refer to Figure 5 , an embodiment of the data annotation device in the embodiments of the present invention includes:

[0103] An automatic annotation module 501, configured to input a dataset to be annotated into a preset automatic annotation model to obtain a first dataset, where the first dataset includes at least one frame of annotated data, and the annotated data includes one or more annotations;

[0104] An annotation fusion and completion module 502, configured to perform fusion and completion between the annotations of the annotated data in the same frame in the first dataset to obtain a second dataset;

[0105] An annotation verification and integration module 503, configured to determine the annotated data with redundant annotations in the second dataset, and perform cross-verification and integration on the second dataset according to the annotated data with redundant annotations to obtain a completed dataset.

[0106] In the embodiments of the present invention, the data annotation device runs the above data annotation method, including obtaining a graph structure containing nodes, where the nodes include nodes to be predicted and known nodes; constructing a corresponding graph convolutional network model based on the graph structure, and inputting the nodes into the graph convolutional network model to calculate the node attention of each node in the graph structure; constructing an attenuation factor corresponding to each node based on the node attention; aggregating messages of all the nodes based on the attenuation factor to obtain a first label of the node to be predicted. This method obtains a graph structure containing nodes to be predicted, inputs it into a pre-trained graph convolutional network model to obtain node attention, calculates the attenuation factor based on the node attention, and aggregates messages of nodes by determining the value of the attenuation factor to obtain the first label of the node to be predicted. The whole process has better inference performance and more general pre-trained embedding representations, and can be effectively used for various downstream tasks, transforming the original process that greatly relied on manual hyperparameter tuning into a large-scale and replicable industrial production mode.

[0107] Please refer to Figure 6 , a second embodiment of the data annotation device in the embodiments of the present invention includes:

[0108] An automatic annotation module 501, configured to input a dataset to be annotated into a preset automatic annotation model to obtain a first dataset, where the first dataset includes at least one frame of annotated data, and the annotated data includes one or more annotations;

[0109] The annotation fusion and completion module 502 is used to fuse and complete the annotations between the annotation data of the same frame in the first dataset to obtain a second dataset;

[0110] The annotation verification and integration module 503 is used to determine the annotation data with redundant annotations in the second dataset, and perform cross-verification and integration on the second dataset according to the annotation data with redundant annotations to obtain a completed dataset.

[0111] In this embodiment, the annotation fusion and completion module 502 is specifically used for:

[0112] The object detection judgment unit 5021 judges whether there are both point cloud object detection annotations and image object detection annotations in the annotation data of the same frame; the plane detection box conversion unit 5022, if so, converts the point cloud object detection annotations included in the annotation data of the same frame into 2D detection boxes based on a preset coordinate system; the overlap rate judgment unit 5023 judges whether the overlap rate between the 2D detection box and the image object detection annotation is higher than a preset threshold; the second dataset acquisition unit 5024, if higher, marks the objects detected by the point cloud object detection annotations and the targets detected by the image object detection annotations as the same target to obtain the second dataset.

[0113] In this embodiment, the plane detection box conversion unit 5022 is specifically used for:

[0114] Obtain one or more 3D detection boxes corresponding to the camera pose information, lidar pose information, and the point cloud object detection annotations in each frame of annotation data in the first dataset; according to the camera pose information and the lidar pose information, construct a stereo-plane mapping relationship between the three-dimensional coordinate system and the plane coordinate system in the same frame; convert the 3D detection box into the corresponding 2D detection box according to the stereo-plane mapping relationship.

[0115] In this embodiment, the annotation fusion and completion module 502 is specifically further used for:

[0116] Judge whether there are at least two of the point cloud semantic segmentation annotation, the point cloud semantic map annotation, or the image semantic segmentation annotation in the annotation data of the same frame; if there are both the point cloud semantic segmentation annotation and the image semantic segmentation annotation, construct a mapping relationship between the point cloud laser points corresponding to the point cloud semantic segmentation annotation and the camera pixels corresponding to the image semantic segmentation annotation according to the stereo-plane mapping relationship; perform annotation completion between the point cloud laser points and the camera pixels according to the mapping relationship to obtain the second dataset; if there are both the point cloud semantic segmentation annotation and the point cloud semantic map annotation, optimize the semantic map annotation according to the point cloud semantic segmentation annotation to obtain the second dataset.

[0117] In this embodiment, the data annotation device further includes a motion annotation acquisition module 504, and the motion annotation acquisition module 504 is specifically configured to:

[0118] Construct the temporal information corresponding to the target frame by frame for the first data set including at least the point cloud object detection annotation and the image object detection annotation; combine the temporal information frame by frame to fuse and obtain the discrete trajectory of the target; obtain the motion annotation of the target by performing differential calculation on the discrete trajectory.

[0119] In this embodiment, the data annotation device further includes a model training module 505, and the model training module 505 is specifically configured to:

[0120] Collect the preprocessed sample data set, input the sample data set with some sample semantic segmentation annotations and / or sample object detection annotations and / or sample object tracking annotations randomly masked into the automatic annotation model to be trained, and obtain the predicted annotation data of the screening model to be trained, where the predicted annotation data includes predicted semantic segmentation annotations and / or predicted object detection annotations and / or predicted object tracking annotations, and the sample data set further includes the completion data set after the completion method is executed this time; the automatic annotation model to be trained further includes the automatic annotation model after the completion method is executed this time; calculate the preset semantic segmentation cross-entropy loss function based on the predicted semantic segmentation annotation and the sample semantic segmentation annotation to obtain the semantic segmentation cross-entropy loss function value; calculate the preset object detection cross-entropy loss function based on the predicted object detection annotation and / or the predicted object tracking annotation and the sample object detection annotation and / or the sample object tracking annotation to obtain the object detection cross-entropy loss function value; determine whether the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value is less than a preset threshold; if not, adjust the model parameters of the automatic annotation model to be trained based on the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value, and re-train the model with the sample data set with some sample semantic segmentation annotations and / or sample object detection annotations and / or sample object tracking annotations randomly masked until the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value is less than the preset threshold; if so, update the screening model to be trained to an automatic annotation model.

[0121] Based on the previous embodiment, this embodiment details the specific functions of each module and the unit composition of some modules. Through the above modules, the specific functions of the original modules are refined, the operation of the data annotation device is improved, the reliability during its operation is enhanced, and the actual logic between each step is clarified, thereby improving the practicality of the device.

[0122] AboveFigure 5 and Figure 6 The data annotation device in the embodiments of the present invention will be described in detail from the perspective of modular functional entities. Next, the data annotation device in the embodiments of the present invention will be described in detail from the perspective of hardware processing.

[0123] Figure 7 FIG. is a schematic structural diagram of a data annotation device provided by an embodiment of the present invention. The data annotation device 700 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 710 (for example, one or more processors) and a memory 720, and one or more storage media 730 (for example, one or more mass storage devices) for storing application programs 733 or data 732. Among them, the memory 720 and the storage media 730 may be transient storage or persistent storage. The program stored in the storage media 730 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the data annotation device 700. Further, the processor 710 may be configured to communicate with the storage media 730 and execute a series of instruction operations in the storage media 730 on the data annotation device 700 to implement the steps of the above data annotation method.

[0124] The data annotation device 700 may further include one or more power supplies 740, one or more wired or wireless network interfaces 750, one or more input / output interfaces 760, and / or one or more operating systems 731, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 7 the shown structural diagram of the data annotation device does not limit the data annotation device provided in the present application, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0125] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is made to execute the steps of the above data annotation method.

[0126] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, or units can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0127] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0128] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A data annotation method, characterized in that, The data annotation method includes: Inputting the dataset to be annotated into a preset automatic annotation model to obtain a first dataset, where the first dataset includes at least one frame of annotated data, and the annotated data contains one or more annotations; Performing fusion and complementation between the annotations of the annotated data in the same frame in the first dataset to obtain a second dataset; Determining the annotated data with redundant annotations in the second dataset, and performing cross-validation integration on the second dataset according to the annotated data with redundant annotations to obtain a complemented dataset; The annotated data further includes motion annotations. Before determining the annotated data with redundant annotations in the second dataset and performing cross-validation integration on the second dataset according to the annotated data with redundant annotations to obtain a complemented dataset, it includes: constructing temporal information corresponding to the target frame by frame for the first dataset including at least point cloud object detection annotations and image object detection annotations; combining the temporal information frame by frame to fuse and obtain the discrete trajectory of the target; obtaining the motion annotation of the target by performing differential calculation on the discrete trajectory; Determining the annotated data with redundant annotations in the second dataset and performing cross-validation integration on the second dataset according to the annotated data with redundant annotations to obtain a complemented dataset, including: performing cross-validation on the annotations in the first dataset based on the connection relationship and annotation specifications between adjacent frames to obtain point cloud type verification information and image type verification information; performing smoothing processing on the annotated data in consecutive frames by using a weighted algorithm through the point cloud type verification information and the image type verification information to obtain processed annotations; traversing frame by frame with the processed annotations after smoothing processing to complement the missing annotated data in the second dataset to obtain a complemented dataset.

2. The data annotation method according to claim 1, wherein The annotation includes object detection annotations, and the object detection annotations include point cloud object detection annotations and image object detection annotations. Performing fusion and complementation between the annotations of the annotated data in the same frame in the first dataset to obtain a second dataset includes: Judging whether both point cloud object detection annotations and image object detection annotations exist in the annotated data of the same frame; If so, converting the point cloud object detection annotations included in the annotated data of the same frame into 2D detection boxes based on a preset coordinate system; Judging whether the overlap rate between the 2D detection box and the image object detection annotation is higher than a preset threshold; If it is higher, marking the object detected by the point cloud object detection annotation and the target detected by the image object detection annotation as the same target to obtain the second dataset.

3. The data annotation method according to claim 2, wherein The coordinate system includes a three-dimensional coordinate system and a planar coordinate system; Converting the point cloud object detection annotations included in the annotated data of the same frame into 2D detection boxes based on a preset coordinate system includes: Obtaining one or more 3D detection boxes corresponding to the camera pose information, lidar pose information, and the point cloud object detection annotations in each frame of annotated data in the first dataset; Constructing a stereo-plane mapping relationship between the three-dimensional coordinate system and the planar coordinate system in the same frame according to the camera pose information and the lidar pose information; Convert the 3D detection box into the corresponding 2D detection box according to the three-dimensional plane mapping relationship.

4. The data annotation method according to claim 3, wherein The annotation further includes semantic segmentation annotation, and the semantic segmentation annotation at least includes point cloud semantic segmentation annotation, point cloud semantic map annotation, and image semantic segmentation annotation. The step of fusing and complementing the annotation data of the same frame in the first data set to obtain the second data set further includes: Determine whether at least two of the point cloud semantic segmentation annotation, the point cloud semantic map annotation, and the image semantic segmentation annotation exist simultaneously in the annotation data of the same frame; If the point cloud semantic segmentation annotation and the image semantic segmentation annotation exist simultaneously, construct a mapping relationship between the point cloud laser points corresponding to the point cloud semantic segmentation annotation and the camera pixels corresponding to the image semantic segmentation annotation according to the three-dimensional plane mapping relationship; Perform annotation complementation between the point cloud laser points and the camera pixels according to the mapping relationship to obtain the second data set; If the point cloud semantic segmentation annotation and the point cloud semantic map annotation exist simultaneously, optimize the semantic map annotation according to the point cloud semantic segmentation annotation to obtain the second data set.

5. The data annotation method according to claim 1, wherein After determining the annotation data with redundant annotations in the second data set and cross-validating and integrating the second data set according to the annotation data with redundant annotations to obtain the complemented data set, it includes: Collect the preprocessed sample data set, and input the sample data set with some sample semantic segmentation annotations and / or sample object detection annotations and / or sample object tracking annotations randomly masked into the to-be-trained automatic annotation model to obtain the predicted annotation data of the to-be-trained screening model, where the predicted annotation data includes predicted semantic segmentation annotations and / or predicted object detection annotations and / or predicted object tracking annotations, and the sample data set further includes the complemented data set obtained after the completion of the complementation method for the current time; the to-be-trained automatic annotation model further includes the automatic annotation model obtained after the completion of the complementation method for the current time; Calculate the preset semantic segmentation cross-entropy loss function based on the predicted semantic segmentation annotation and the sample semantic segmentation annotation to obtain the semantic segmentation cross-entropy loss function value; Calculate the preset object detection cross-entropy loss function based on the predicted object detection annotation and / or the predicted object tracking annotation and the sample object detection annotation and / or the sample object tracking annotation to obtain the object detection cross-entropy loss function value; Determine whether the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value is less than a preset threshold; If not, adjust the model parameters of the to-be-trained automatic annotation model based on the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value, and re-train the model with the sample data set with some sample semantic segmentation annotations and / or sample object detection annotations and / or sample object tracking annotations randomly masked until the semantic segmentation cross-entropy loss function value and / or the object detection cross-entropy loss function value is less than the preset threshold; If so, update the to-be-trained screening model to an automatic annotation model.

6. A data annotation device, characterized in that, The data annotation device includes: An automatic annotation module for inputting a dataset to be annotated into a preset automatic annotation model to obtain a first dataset, where the first dataset includes at least one frame of annotated data, and the annotated data contains one or more annotations; An annotation fusion and completion module for fusing and completing the annotations between the annotated data of the same frame in the first dataset to obtain a second dataset; An annotation verification and integration module for determining the annotated data with redundant annotations in the second dataset, and cross-verifying and integrating the second dataset according to the annotated data with redundant annotations to obtain a completed dataset; The data annotation device further includes a motion annotation acquisition module, and the motion annotation acquisition module is specifically configured to: frame by frame construct the temporal information corresponding to the target for the first dataset including at least point cloud object detection annotations and image object detection annotations; combine the temporal information frame by frame to fuse and obtain the discrete trajectory of the target; obtain the motion annotation of the target by performing differential calculation on the discrete trajectory; The annotation verification and integration module is specifically configured to: cross-verify the annotations in the first dataset based on the connection relationship and annotation specifications between adjacent frames to obtain point cloud type verification information and image type verification information; perform smoothing processing on the annotated data in consecutive frames by adopting a weighted algorithm through the point cloud type verification information and the image type verification information to obtain processed annotations; traverse the missing annotated data in the second dataset frame by frame by adopting the processed annotations after smoothing processing to obtain a completed dataset.

7. A data annotation device, characterized in that, The data annotation device includes: a memory and at least one processor, where instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line; The at least one processor invokes the instructions in the memory so that the data annotation device executes each step of the data annotation method according to any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it realizes each step of the data annotation method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Vehicle model prediction method and device based on image recognition, equipment and medium

    CN113850263A

  • KR20200092715A