Data labeling method and device and storage medium

By pre-labeling the images of autonomous driving environment using high-precision and high speed perception models, the problems of low efficiency and high cost of autonomous driving data labeling are solved, and efficient and accurate data labeling is achieved.

CN120148031APending Publication Date: 2025-06-13BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311717104.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In autonomous driving technology, the training of perception algorithms requires a large amount of labeling data, and the landing scenarios of autonomous driving are diverse, resulting in low data labeling efficiency and high cost.

Method used

Through the pre-trained first perception model and the second perception model, the data set to be marked image is determined from the continuous multi-frame environment images, including a multi-frame environment image to be marked and a target label corresponding to each frame environment image to be marked. The detection accuracy of the first perception model is higher than that of the second perception model, and the detection speed of the second perception model is higher than that of the first perception model.

Benefits of technology

Pre-labeling of environmental images is realized, improving the efficiency of data labeling, reducing manual labeling costs, and improving the accuracy of data labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148031A_ABST
    Figure CN120148031A_ABST
Patent Text Reader

Abstract

The invention relates to a data annotation method and device and a storage medium. Continuous multi-frame environment images can be obtained; determining a to-be-labeled image data set through a first perception model and a second perception model obtained through pre-training, wherein the to-be-labeled image data set comprises multiple frames of to-be-labeled environment images determined from the continuous multiple frames of environment images and a target label corresponding to each frame of to-be-labeled environment image; the detection precision of the first sensing model is higher than that of the second sensing model, and the detection speed of the second sensing model is higher than that of the first sensing model; and for each frame of the to-be-labeled environment image in the to-be-labeled image data set, performing data labeling on the frame of the to-be-labeled environment image according to the target label corresponding to the frame of the to-be-labeled environment image to obtain a target labeled image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, and in particular, to a data annotation method, apparatus, and storage medium. Background Art

[0002] Perception algorithms play an important role in autonomous driving technology. The training of perception algorithms requires a large amount of annotated data. The data collection in autonomous driving is real-scene data collected by on-vehicle cameras. That is, before applying the perception algorithm to the autonomous driving technology of a vehicle, it is necessary to perform model training on the perception model corresponding to the perception algorithm based on the real-scene data collected by the on-vehicle cameras.

[0003] In addition, there are many landing scenarios for autonomous driving. In order to improve the applicability of the perception model to various scenarios corresponding to autonomous driving, it is necessary to separately collect real-scene data for various scenarios to perform model training on the perception model. Summary of the Invention

[0004] To overcome the problems in the related art, the present disclosure provides a data annotation method, apparatus, and storage medium.

[0005] According to a first aspect of an embodiment of the present disclosure, a data annotation method is provided, including:

[0006] Obtain a plurality of consecutive frames of environmental images;

[0007] Determine a dataset of images to be annotated through a pre-trained first perception model and a second perception model. The dataset of images to be annotated includes a plurality of frames of environmental images to be annotated determined from the plurality of consecutive frames of environmental images and a target label corresponding to each frame of the environmental image to be annotated. The detection accuracy of the first perception model is higher than that of the second perception model, and the detection speed of the second perception model is higher than that of the first perception model;

[0008] For each frame of the environmental image to be annotated in the dataset of images to be annotated, perform data annotation on the frame of the environmental image to be annotated according to the target label corresponding to the frame of the environmental image to be annotated, to obtain a target annotated image.

[0009] Optionally, the determining the dataset of images to be annotated through the pre-trained first perception model and the second perception model includes:

[0010] Perform frame extraction processing on the plurality of consecutive frames of environmental images according to a preset frame extraction ratio to obtain a plurality of frames of first environmental images;

[0011] Determine the dataset of images to be annotated through the first perception model and the second perception model according to the plurality of frames of first environmental images and the plurality of consecutive frames of environmental images.

[0012] Optionally, determining the image dataset to be labeled according to the multi-frame first environmental images and the continuous multi-frame environmental images through the first perception model and the second perception model includes:

[0013] For each frame of the first environmental images, input the first environmental images into the first perception model and the second perception model respectively, and obtain a first detection result output by the first perception model and a second detection result output by the second perception model;

[0014] For each frame of the continuous multi-frame environmental images, input the environmental images into the first perception model, and obtain a third detection result output by the first perception model;

[0015] Determine the image dataset to be labeled according to the first detection result, the second detection result, and the third detection result.

[0016] Optionally, the detection result includes a detection area corresponding to the target object in the environmental image and a confidence level that the detection area is an image area corresponding to the target object;

[0017] Determining the image dataset to be labeled according to the first detection result, the second detection result, and the third detection result includes:

[0018] For each frame of the first environmental images, according to the first detection result and the second detection result corresponding to the first environmental image, determine the overlapping degree between the first detection area of the target object in the first detection result and the second detection area of the target object in the second detection result;

[0019] Determine the image dataset to be labeled according to the overlapping degree, the confidence level in the first detection result, and the third detection result.

[0020] Optionally, determining the overlapping degree between the first detection area of the target object in the first detection result and the second detection area of the target object in the second detection result according to the first detection result and the second detection result corresponding to the first environmental image includes:

[0021] Calculate the intersection over union (IoU) of the first detection area and the second detection area;

[0022] Take the IoU as the overlapping degree.

[0023] Optionally, the image dataset to be labeled includes at least one of a first subset, a second subset, a third subset, and a fourth subset. Determining the image dataset to be labeled according to the degree of overlap, the confidence in the first detection result, and the third detection result includes:

[0024] Determining the first subset and the second subset according to the degree of overlap;

[0025] Determining the third subset according to the confidence in the first detection result;

[0026] Determining the fourth subset according to the third detection result.

[0027] Optionally, determining the first subset and the second subset according to the degree of overlap includes:

[0028] Taking, among the multiple frames of first environmental images, the first image with the degree of overlap less than or equal to a first preset overlap threshold as the environmental image to be labeled in the first subset, and recording the first detection result corresponding to the first image as the target label of the first image in the first subset;

[0029] Taking, among the multiple frames of first environmental images, the second image with the degree of overlap less than or equal to a second preset overlap threshold as the environmental image to be labeled in the second subset, and recording the first detection result corresponding to the second image as the target label of the second image in the second subset;

[0030] Wherein, the first preset overlap threshold is greater than the second preset overlap threshold.

[0031] Optionally, determining the third subset according to the confidence in the first detection result includes:

[0032] Taking, among the multiple frames of first environmental images, the third image with the confidence greater than or equal to a preset confidence threshold as the environmental image to be labeled in the third subset, and recording the first detection result corresponding to the third image as the target label of the third image in the third subset.

[0033] Optionally, determining the fourth subset according to the third detection result includes:

[0034] After performing target tracking on the target object on each frame of environmental image according to the third detection result by using a preset optical flow tracking algorithm, determining the tracking result;

[0035] According to the tracking result, the fourth image in the consecutive multiple frames of environmental images is used as the environmental image to be labeled in the fourth subset, and the fourth image is the environmental image in which the target object is not detected according to the third detection result in every preset number of consecutive multiple frames of environmental images;

[0036] Use the third detection results corresponding to the preset number of consecutive multiple frames of environmental images as the target label corresponding to the fourth image;

[0037] Record the target label corresponding to the fourth image into the fourth subset.

[0038] Optionally, after performing target tracking on the target object on each frame of environmental image according to the third detection result using a preset optical flow tracking algorithm, the determined tracking result includes:

[0039] For every preset number of consecutive multiple frames of environmental images in the consecutive multiple frames of environmental images, according to the third detection result, perform target tracking on the target object of each adjacent two frames of environmental images in the preset number of consecutive multiple frames of environmental images using the preset optical flow tracking algorithm, and then determine the tracking result;

[0040] Wherein, the tracking result indicates whether there is at least one frame of environmental image missing the target object in the preset number of consecutive multiple frames of environmental images in the third detection result.

[0041] Optionally, the method further includes:

[0042] Perform a deduplication operation on the dataset of images to be labeled to obtain a target image dataset;

[0043] The data annotation of each frame of the environmental image to be labeled according to the target label to obtain the target annotated image includes:

[0044] Perform data annotation on each frame of the environmental image to be labeled in the target image dataset according to the target label.

[0045] Optionally, the method further includes:

[0046] Train the second perception model according to the target annotated image to obtain a target perception model.

[0047] According to the second aspect of the embodiments of the present disclosure, a data annotation device is provided, including:

[0048] An acquisition module, configured to acquire consecutive multiple frames of environmental images;

[0049] A determination module, configured to determine a dataset of images to be labeled through a first perception model and a second perception model obtained by pre-training. The dataset of images to be labeled includes multiple frames of environmental images to be labeled determined from the continuous multi-frame environmental images and target labels corresponding to each frame of the environmental images to be labeled. The detection accuracy of the first perception model is higher than that of the second perception model, and the detection speed of the second perception model is higher than that of the first perception model.

[0050] A labeling module, configured to, for each frame of the environmental images to be labeled in the dataset of images to be labeled, perform data labeling on the frame of the environmental images to be labeled according to the target label corresponding to the frame of the environmental images to be labeled, to obtain target labeled images.

[0051] According to a third aspect of the embodiments of the present disclosure, there is provided an apparatus for data labeling, including:

[0052] A processor;

[0053] A memory for storing instructions executable by the processor;

[0054] Wherein, the processor is configured to: execute the data labeling method described in the first aspect of the present disclosure.

[0055] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the data labeling method provided in the first aspect of the present disclosure are implemented.

[0056] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: Based on the first perception model and the second perception model, multiple frames of environmental images to be labeled and target labels corresponding to each frame of the environmental images to be labeled are determined from continuous multi-frame environmental images, thereby completing the pre-labeling work of the environmental images. In this way, multiple frames of environmental images to be labeled can be screened from the continuous multi-frame environmental images through the first perception model and the second perception model for data labeling, and the target labels obtained by pre-labeling can be used as the basis for image labeling, thereby improving the efficiency of data labeling and reducing the manual labeling cost. In addition, since the detection accuracy of the first perception model is higher than that of the second perception model, using the target labels corresponding to each frame of the environmental images to be labeled comprehensively determined by the first perception model and the second perception model as the basis for data labeling can also improve the accuracy of data labeling.

[0057] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings

[0058] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used in conjunction with the specification to explain the principles of the present disclosure.

[0059] Figure 1 is a flowchart of a data annotation method shown according to an exemplary embodiment.

[0060] Figure 2 is according to Figure 1 a flowchart of a data annotation method shown according to the illustrated embodiment.

[0061] Figure 3 is according to Figure 2 a flowchart of a data annotation method shown according to the illustrated embodiment.

[0062] Figure 4 is according to Figure 1 a flowchart of a data annotation method shown according to the illustrated embodiment.

[0063] Figure 5 is according to Figure 1 a flowchart of a data annotation method shown according to the illustrated embodiment.

[0064] Figure 6 is a block diagram of a data annotation device shown according to an exemplary embodiment.

[0065] Figure 7 is according to Figure 6 a block diagram of a data annotation device shown according to the illustrated embodiment.

[0066] Figure 8 is according to Figure 6 a block diagram of a data annotation device shown according to the illustrated embodiment.

[0067] Figure 9 is a block diagram of a device for data annotation shown according to an exemplary embodiment. Detailed Description of the Embodiments

[0068] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0069] It should be noted that all actions of obtaining signals, information, or data in this application are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining the authorization given by the owner of the corresponding device.

[0070] First, the application scenarios of the present disclosure will be described. The present disclosure is mainly applied to the scenario of data annotation for the sample data required for model training. For example, the model can be a perception model applied to the perception algorithm in autonomous driving technology.

[0071] The data collection in autonomous driving is the real-scene data collected by in-vehicle cameras. Most of this data will have duplicate situations. Therefore, the collected data first needs to be screened to filter out a large amount of duplicate data, and then the data is annotated. In addition, there are many landing scenarios for autonomous driving, so there are also many data scenarios that need to be annotated, which will result in an overly large amount of annotated data. Relying entirely on manual annotation will incur a large amount of time cost. And currently, the selection of in-vehicle models for autonomous driving is mainly based on real-time performance. Therefore, it is difficult to deploy a model with accurate detection accuracy but long time consumption to the vehicle side, while a model with less time consumption and performance meeting the requirements requires a large amount of annotated data to complete training.

[0072] Therefore, in order to improve the annotation efficiency of the training data for in-vehicle real-time models and reduce the manual annotation cost, the present disclosure provides a data annotation method, apparatus, and storage medium. The following will describe the specific embodiments of the present disclosure in detail with reference to the accompanying drawings.

[0073] Figure 1 is a flowchart of a data annotation method shown according to an exemplary embodiment. As Figure 1 shown, the method includes the following steps:

[0074] In step S101, a plurality of consecutive frames of environmental images are acquired.

[0075] Among them, the plurality of consecutive frames of environmental images can be a plurality of consecutive frames of environmental images around the vehicle collected in real time by an in-vehicle image acquisition device.

[0076] In step S102, a dataset of images to be annotated is determined by a pre-trained first perception model and a second perception model. The dataset of images to be annotated includes a plurality of frames of environmental images to be annotated determined from the plurality of consecutive frames of environmental images and the target label corresponding to each frame of the environmental images to be annotated. The detection accuracy of the first perception model is higher than that of the second perception model, and the detection speed of the second perception model is higher than that of the first perception model.

[0077] Among them, the first perception model can be understood as a large model with a relatively complex structure and high recognition accuracy obtained through offline training. The second perception model can be understood as an in-vehicle real-time model that needs to be deployed on the vehicle for real-time object detection. The complexity of the second perception model is less than that of the first perception model, and the detection accuracy of the second perception model is less than that of the first perception model. However, the detection speed of the second perception model is higher than that of the first perception model, so as to realize the real-time perception of the surrounding environment during vehicle autonomous driving.

[0078] It should be noted that there can also be multiple first perception models in the present disclosure. The difference between each first perception model lies in that the detection object may be different, or for different first perception models, their detection accuracies are also different. The present disclosure does not limit this. In this way, for each first perception model, the dataset of images to be labeled can be determined through the first perception model and the second perception model. After data fusion of the datasets of images to be labeled determined by each first perception model, the final dataset of images to be labeled is obtained.

[0079] It should also be noted that in order to improve the detection accuracy of the in-vehicle real-time model that needs to be deployed on the vehicle, the present disclosure can use the target labeled images after data labeling to train the second perception model, and deploy the trained target perception model to the vehicle.

[0080] In addition, the multiple frames of environment images to be labeled included in the dataset of images to be labeled can be understood as the environment images obtained by performing a duplicate removal operation on the continuously acquired multiple frames of original environment images, selecting from the continuously acquired multiple frames of environment images, which is helpful for improving the detection accuracy of the in-vehicle real-time model, and at the same time deleting some data with low value for the model training of the in-vehicle real-time model. Therefore, after data labeling the multiple frames of environment images to be labeled to obtain target labeled images, and using the target labeled images to train the second perception model, the detection accuracy of the second perception model can be significantly improved.

[0081] In step S103, for each frame of the environment image to be labeled in the dataset of images to be labeled, data labeling is performed on the frame of the environment image to be labeled according to the target label corresponding to the frame of the environment image to be labeled, to obtain a target labeled image.

[0082] The target label is an image annotation label determined by comprehensively considering the model detection results of the first perception model and the second perception model. The model detection results can include the detection results of the model for the image regions corresponding to the target objects (such as pedestrians, vehicles, buildings, etc.) included in the environment image.

[0083] In one possible implementation of this step, for each frame of the to-be-annotated environmental image in the to-be-annotated image dataset, the corresponding target label can be directly used as the annotation label for this frame of the to-be-annotated environmental image to obtain the target annotated image corresponding to this frame of the to-be-annotated environmental image.

[0084] In addition, in order to improve the accuracy of data annotation, the to-be-annotated image dataset can also be handed over to humans for annotation. During the process of manually annotating each frame of the to-be-annotated environmental image, the target label given by the first perception model and the second perception model can be used as the pre-annotation label for annotation, which can improve the accuracy of data annotation. At the same time, when manually annotating, there is no need to sequentially screen and annotate the original continuous multi-frame environmental images, which can also improve the efficiency of data annotation.

[0085] Using the above method, based on the first perception model and the second perception model, multiple frames of to-be-annotated environmental images and the target label corresponding to each frame of the to-be-annotated environmental image are determined from the continuous multi-frame environmental images, thus completing the pre-annotation work of the environmental images. In this way, multiple frames of to-be-annotated environmental images can be screened from the continuous multi-frame environmental images through the first perception model and the second perception model for data annotation, and the target label obtained based on the pre-annotation can be used as the basis for image annotation, thereby improving the efficiency of data annotation and reducing the manual annotation cost. In addition, since the detection accuracy of the first perception model is higher than that of the second perception model, using the target label corresponding to each frame of the to-be-annotated environmental image comprehensively determined by the first perception model and the second perception model as the basis for data annotation can also improve the accuracy of data annotation.

[0086] Figure 2 is according to Figure 1 The flowchart of a data annotation method shown in the illustrated embodiment is as Figure 2 shown, and step S102 includes the following sub-steps:

[0087] In step S1021, the continuous multi-frame environmental images are frame-sampled according to a preset frame-sampling ratio to obtain multiple frames of first environmental images.

[0088] As described above, the data collection in autonomous driving is the real-scene data collected by on-vehicle cameras, and most of the data will have duplicate situations. Therefore, in order to minimize the duplicate environmental images in the continuous multi-frame environmental images as much as possible, the present disclosure can perform frame-sampling processing on the continuous multi-frame environmental images according to a preset frame-sampling ratio to obtain multiple frames of first environmental images.

[0089] Among them, the preset frame extraction ratio can be, for example, 1 / 3 or 1 / 4. Taking the preset frame extraction ratio of 1 / 3 as an example, one frame of image can be randomly selected and retained as one frame of the first environmental image from every three consecutive frames of the continuous multi-frame environmental images, and the other two frames of environmental images are deleted. In this way, multiple frames of the first environmental images after frame extraction processing on the continuous multi-frame environmental images can be obtained. This is only an example, and the present disclosure does not limit this.

[0090] In step S1022, according to the multiple frames of the first environmental images and the continuous multi-frame environmental images, the dataset of the images to be labeled is determined through the first perception model and the second perception model.

[0091] Exemplarily, Figure 3 is a flowchart of a data annotation method shown in the illustrated embodiment. As Figure 2 shown, step S1022 includes the following sub-steps: Figure 3 In step S10221, for each frame of the first environmental image, the first environmental image is respectively input into the first perception model and the second perception model to obtain a first detection result output by the first perception model and a second detection result output by the second perception model.

[0092] Among them, the detection result (including the first detection result, the second detection result, and the third detection result below) includes the detection area corresponding to the target object in the environmental image and the confidence level that the detection area is the image area corresponding to the target object. The target object can be, for example, any object such as a pedestrian, a vehicle, and a building in the vehicle surrounding environment.

[0093] In step S10222, for each frame of the continuous multi-frame environmental images, the environmental image is input into the first perception model to obtain a third detection result output by the first perception model.

[0094] In step S10223, the dataset of the images to be labeled is determined according to the first detection result, the second detection result, and the third detection result.

[0095] In this step, for each frame of the first environmental image, according to the first detection result and the second detection result corresponding to the first environmental image, the overlapping degree between the first detection area of the target object in the first detection result and the second detection area of the target object in the second detection result can be determined; the dataset of the images to be labeled is determined according to the overlapping degree, the confidence level in the first detection result, and the third detection result.

[0096]

[0097] ​In one implementation, the overlapping degree can be represented by the intersection over union of the first detection region and the second detection region. Therefore, the overlapping degree of the first detection region and the second detection region can be calculated in the following manner:

[0098] Calculate the intersection over union of the first detection region and the second detection region; use the intersection over union as the overlapping degree.

[0099] In addition, in the present disclosure, the image dataset to be labeled may include at least one of a first subset, a second subset, a third subset, and a fourth subset. For example, the image dataset to be labeled may be the union of the first subset, the second subset, the third subset, and the fourth subset.

[0100] In the process of the present disclosure determining the image dataset to be labeled according to the overlapping degree, the confidence level in the first detection result, and the third detection result, the first subset and the second subset may be determined according to the overlapping degree; the third subset may be determined according to the confidence level in the first detection result; and the fourth subset may be determined according to the third detection result.

[0101] In one implementation, the present disclosure may determine the first subset according to the overlapping degree in the following manner:

[0102] Among the multiple frames of first environmental images, use the first images with the overlapping degree less than or equal to a first preset overlapping threshold as the environmental images to be labeled in the first subset, and record the corresponding first detection results of the first images as the target labels of the first images in the first subset.

[0103] Among them, the first detection region of the target object refers to the image region of the target object in the corresponding first environmental image obtained after target detection by the first perception model with higher detection accuracy. The second detection region of the target object refers to the image region of the target object in the corresponding first environmental image obtained after target detection by the second perception model with lower detection accuracy. It can be understood that the higher the overlap degree between the first detection region of the target object and the second detection region of the target object, the more accurate the detection of the target object in the first environmental image by the second perception model. In other words, the second perception model can already have a good recognition result for the first environmental image, and the value of the first environmental image for further training to obtain a second perception model with higher detection accuracy is generally average. In this case, the first environmental image can be not used as the to-be-annotated environmental image for training the second perception model. On the contrary, the lower the overlap degree between the first detection region of the target object and the second detection region of the target object, the higher the value of the corresponding first environmental image for further training to obtain a second perception model with higher detection accuracy. In this case, the first environmental image can be screened as the to-be-annotated environmental image for training the second perception model.

[0104] Therefore, in an implementation manner of the present disclosure, among multiple frames of first environmental images, the first images with the overlap degree between the first detection region of the target object and the second detection region of the target object being less than or equal to the first preset overlap threshold can be used as the to-be-annotated environmental images in the first subset. Moreover, the first detection result of the first perception model with higher detection accuracy for the first image can be used as the target label of the first image, improving the accuracy of annotation. In this way, after data annotation of the to-be-annotated environmental images in the first subset, it can be used to train the second perception model, thereby improving the detection accuracy of the second perception model.

[0105] In an implementation manner, the present disclosure can determine the second subset according to the overlap degree in the following way:

[0106] Among the multiple frames of first environmental images, the second images with the overlap degree being less than or equal to the second preset overlap threshold are used as the to-be-annotated environmental images in the second subset, and the corresponding first detection results of the second images are recorded as the target labels of the second images in the second subset; where the first preset overlap threshold is greater than the second preset overlap threshold. For example, the second preset overlap threshold can be 0 or a relatively small number close to 0.

[0107] Taking the case where the second preset overlap threshold is equal to 0 as an example, if the overlap degree between the first detection area and the second detection area of the target object is 0, it means that the image area corresponding to the target object is detected in the first perception model, but the image area of the target object is not detected in the second perception model. In this case, the corresponding first environmental image should also be used as the environmental image to be labeled for training the second perception model. Therefore, among the multiple frames of the first environmental images, the second image with the overlap degree less than or equal to the second preset overlap threshold can be used as the environmental image to be labeled in the second subset, and the first detection result corresponding to the second image can be recorded as the target label of the second image in the second subset.

[0108] In order to further improve the accuracy of data annotation, in one implementation, the present disclosure can determine the third subset according to the confidence level in the first detection result in the following manner:

[0109] Among the multiple frames of the first environmental images, the third image with the confidence level in the first detection result greater than or equal to the preset confidence threshold can be used as the environmental image to be labeled in the third subset, and the first detection result corresponding to the third image can be recorded as the target label of the third image in the third subset.

[0110] Among them, the confidence level in the first detection result represents the credibility of the detection area obtained after the first perception model performs target detection on the input first environmental image. The higher the confidence level, the more accurate the detection result of the target object in the input image by the first perception model. Therefore, in order to improve the detection accuracy of the model after training the second perception model, it is also necessary to use the third image with the confidence level in the first detection result greater than or equal to the preset confidence threshold as the environmental image to be labeled in the third subset for subsequent model training of the second perception model.

[0111] In one implementation, the present disclosure can determine the fourth subset according to the third detection result in the following manner:

[0112] After performing target tracking on the target object on each frame of the environmental image according to the third detection result by using the preset optical flow tracking algorithm, the tracking result is determined; according to the tracking result, the fourth image in the continuous multiple frames of environmental images is used as the environmental image to be labeled in the fourth subset. The fourth image is the environmental image in which the target object is not detected according to the third detection result among every preset number of continuous multiple frames of environmental images; the third detection results corresponding to the preset number of continuous multiple frames of environmental images are used as the target label corresponding to the fourth image; and the target label corresponding to the fourth image is recorded in the fourth subset.

[0113] Among them, for every preset number of consecutive multi-frame environmental images in the consecutive multi-frame environmental images, according to the third detection result, after performing target tracking on the target object in every two adjacent environmental images in the preset number of consecutive multi-frame environmental images through the preset optical flow tracking algorithm, the tracking result is determined; wherein, the tracking result indicates whether there is at least one frame of environmental image missing the target object in the preset number of consecutive multi-frame environmental images in the third detection result.

[0114] Among them, the third detection result is the detection result obtained after inputting each frame of environmental image in the original consecutive multi-frame environmental images without frame extraction into the first perception model.

[0115] In the present disclosure, the preset optical flow tracking algorithm can also be used to perform tracking processing on the third detection result, so as to determine the environmental images in which the first perception model does not detect the target object in every preset number of consecutive multi-frame environmental images. This frame of environmental image can also be used as the environmental image to be labeled for data annotation and then used for subsequent training of the second perception model. Therefore, in the present disclosure, the environmental images in which the first perception model does not detect the target object in every preset number of consecutive multi-frame environmental images can be stored as the environmental images to be labeled in the fourth subset.

[0116] For example, assume that the consecutive multi-frame environmental images include x 1 , x 2 ,......x n These N frames of environmental images. Among them, the preset number can be 5 for example. Then the N frames of environmental images can be divided into N / 5 groups of environmental images, and each group includes 5 consecutive environmental images. For each group of environmental images, the detection results of the target object in the third detection result can be tracked for every two adjacent environmental images in the group through the preset optical flow tracking algorithm, so as to determine whether there is an environmental image in the group that is not detected by the third detection result for the target object, and the frame of environmental image that does not detect the target object is used as the fourth image and stored in the fourth subset. This is only an example here, and the present disclosure does not make any limitations in this regard.

[0117] That is to say, the dataset of images to be labeled can include any one of the first subset, the second subset, the third subset, and the fourth subset, or can also include the union of any subsets of the first subset, the second subset, the third subset, and the fourth subset. The present disclosure does not make any limitations in this regard.

[0118] In summary, the first subset includes the environment images to be labeled as the first images, where the first images are the images in the multiple frames of the first environment images with the overlapping degree between the first detection region and the second detection region of the target object being less than or equal to the first preset overlapping threshold. The second subset includes the environment images to be labeled as the second images, where the second images are the images in the multiple frames of the first environment images with the overlapping degree between the first detection region and the second detection region of the target object being less than or equal to the second preset overlapping threshold, and the first preset overlapping threshold is greater than the second preset overlapping threshold, and the second preset overlapping threshold can be 0 for example. The third subset includes the environment images to be labeled as the third images, where the third images are the images in the first detection results with the confidence level being greater than or equal to the preset confidence threshold. The fourth subset includes the environment images to be labeled as the fourth images, where the fourth images are the environment images in which the target object is not detected in every preset number of consecutive frames of the consecutive multiple frames of environment images.

[0119] In order to further improve the efficiency of data annotation and reduce the duplicate images in the images to be labeled, the present disclosure can further perform a deduplication operation on the dataset of the images to be labeled.

[0120] Figure 4 is according to Figure 1 shown in the flowchart of a data annotation method according to the illustrated embodiment, as Figure 4 shown, the method further includes the following steps:

[0121] In step S104, perform a deduplication operation on the dataset of the images to be labeled to obtain a target image dataset.

[0122] Exemplarily, a preset deduplication algorithm (such as a hash algorithm) can be used for the deduplication operation, and only one frame of the images with relatively large image similarity in the dataset of the images to be labeled is retained, and the redundant images with relatively high similarity are deleted.

[0123] In this way, in the process of executing step S103, the data annotation can be performed on each frame of the environment images to be labeled in the target image dataset according to the target labels.

[0124] Figure 5 is according to Figure 1 shown in the flowchart of a data annotation method according to the illustrated embodiment, as Figure 5 shown, the method includes the following steps:

[0125] In step S105, perform model training on the second perception model according to the target annotated images to obtain a target perception model.

[0126] For the model training of the second perception model, since the target labeled image is a labeled image that is of high value for improving the detection accuracy of the trained model, therefore, the target perception model obtained by training the second perception model based on this labeled image will have a significantly improved model detection accuracy compared to the second perception model. Moreover, compared to the first perception model, the second perception model is lighter and has better real-time performance in target detection. Therefore, the target perception model can be deployed as an in-vehicle real-time model to the vehicle for real-time and high-precision target detection of the surrounding driving environment, improving the safety of vehicle driving.

[0127] By using the above method, useful data for training the in-vehicle real-time model (i.e., the second perception model) can be determined through the offline high-precision model (i.e., the first perception model), which greatly utilizes the collected data set. At the same time, for the model training of the second perception model, it also avoids adding too much data with low value to the annotation. Here, the data with low value refers to the image data that the second perception model can already identify or detect well. For example, an environmental image with a high degree of overlap between the first detection area and the second detection area can be regarded as an image that the second perception model can already identify or detect well, and this type of environmental image does not need to be used as the to-be-labeled environmental image in the to-be-labeled image data set.

[0128] In addition, using the model inference result of the first perception model with high detection accuracy as the pre-labeling label can save most of the manual labeling costs, improve the manual labeling efficiency, and shorten the data labeling cycle.

[0129] Figure 6 It is a block diagram of a data labeling device shown according to an exemplary embodiment. As Figure 6 shown, the device includes:

[0130] An acquisition module 601, configured to acquire a continuous multi-frame environmental image;

[0131] A determination module 602, configured to determine a to-be-labeled image data set through a pre-trained first perception model and a second perception model. The to-be-labeled image data set includes multiple to-be-labeled environmental images determined from the continuous multi-frame environmental images and the target label corresponding to each to-be-labeled environmental image. The detection accuracy of the first perception model is higher than that of the second perception model, and the detection speed of the second perception model is higher than that of the first perception model;

[0132] A labeling module 603, configured to perform data labeling on each to-be-labeled environmental image in the to-be-labeled image data set according to the target label corresponding to the to-be-labeled environmental image to obtain a target labeled image.

[0133] Optionally, the determining module 602 is configured to perform frame extraction processing on the continuous multi-frame environmental images according to a preset frame extraction ratio to obtain multiple first environmental images; and determine the image dataset to be annotated based on the multiple first environmental images and the continuous multi-frame environmental images through the first perception model and the second perception model.

[0134] Optionally, for each frame of the first environmental image, the determining module 602 is configured to input the first environmental image into the first perception model and the second perception model respectively to obtain a first detection result output by the first perception model and a second detection result output by the second perception model; for each frame of the environmental image in the continuous multi-frame environmental images, input the environmental image into the first perception model to obtain a third detection result output by the first perception model; and determine the image dataset to be annotated based on the first detection result, the second detection result, and the third detection result.

[0135] Optionally, the detection result includes a detection area corresponding to the target object in the environmental image and a confidence level that the detection area is the image area corresponding to the target object.

[0136] For each frame of the first environmental image, the determining module 602 is configured to determine the overlapping degree between a first detection area of the target object in the first detection result and a second detection area of the target object in the second detection result according to the first detection result and the second detection result corresponding to the first environmental image; and determine the image dataset to be annotated based on the overlapping degree, the confidence level in the first detection result, and the third detection result.

[0137] Optionally, the determining module 602 is configured to calculate the intersection over union (IoU) of the first detection area and the second detection area; and use the IoU as the overlapping degree.

[0138] Optionally, the image dataset to be annotated includes at least one of a first subset, a second subset, a third subset, and a fourth subset. The determining module 602 is configured to determine the first subset and the second subset according to the overlapping degree; determine the third subset according to the confidence level in the first detection result; and determine the fourth subset according to the third detection result.

[0139] Optionally, the determining module 602 is configured to use, as the environment image to be labeled in the first subset, the first image in the multiple-frame first environment images whose overlapping degree is less than or equal to a first preset overlapping threshold, and record, as the target label of the first image, the first detection result corresponding to the first image in the first subset; use, as the environment image to be labeled in the second subset, the second image in the multiple-frame first environment images whose overlapping degree is less than or equal to a second preset overlapping threshold, and record, as the target label of the second image, the first detection result corresponding to the second image in the second subset; wherein the first preset overlapping threshold is greater than the second preset overlapping threshold.

[0140] Optionally, the determining module 602 is configured to use, as the environment image to be labeled in the third subset, the third image in the multiple-frame first environment images whose confidence level is greater than or equal to a preset confidence threshold, and record, as the target label of the third image, the first detection result corresponding to the third image in the third subset.

[0141] Optionally, the determining module 602 is configured to, after performing target tracking on the target object on each frame of the environment image according to the third detection result by using a preset optical flow tracking algorithm, determine a tracking result; according to the tracking result, use, as the environment image to be labeled in the fourth subset, the fourth image in the consecutive multiple-frame environment images, where the fourth image is the environment image in which the target object is not detected according to the third detection result in every preset number of consecutive multiple-frame environment images; use the third detection results corresponding to the preset number of consecutive multiple-frame environment images as the target label corresponding to the fourth image; and record the target label corresponding to the fourth image in the fourth subset.

[0142] Optionally, the determining module 602 is configured to, for every preset number of consecutive multiple-frame environment images in the consecutive multiple-frame environment images, perform target tracking on the target object of every two adjacent frames of the environment images in the preset number of consecutive multiple-frame environment images according to the third detection result by using the preset optical flow tracking algorithm, and then determine the tracking result; wherein the tracking result indicates whether there is at least one frame of environment image in which the target object is lost in the third detection result for the preset number of consecutive multiple-frame environment images.

[0143] Optionally, Figure 7 is a block diagram of a data annotation device shown in the Figure 6 illustrated embodiment, as Figure 7 shown, the device further includes:

[0144] The deduplication module 604 is configured to perform deduplication operations on the image dataset to be labeled, and obtain a target image dataset;

[0145] The annotation module 603 is configured to perform data annotation on each frame of the environment image to be labeled in the target image dataset according to the target label.

[0146] Optionally, Figure 8 is a block diagram of a data annotation device shown in the Figure 6 illustrated embodiment. As shown in Figure 8 shown, the device further includes:

[0147] The model training module 605 is configured to perform model training on the second perception model according to the target annotated image, and obtain a target perception model.

[0148] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.

[0149] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the data annotation method provided by the present disclosure are implemented.

[0150] Figure 9 is a block diagram of a device for data annotation shown in an exemplary embodiment. For example, the device 900 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0151] Referring to Figure 9 , the device 900 may include one or more of the following components: a processing component 902, a memory 904, a power component 906, a multimedia component 908, an audio component 910, an input / output interface 912, a sensor component 914, and a communication component 916.

[0152] The processing component 902 generally controls the overall operation of the device 900, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 902 may include one or more processors 920 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 902 may include one or more modules to facilitate the interaction between the processing component 902 and other components. For example, the processing component 902 may include a multimedia module to facilitate the interaction between the multimedia component 908 and the processing component 902.

[0153] The memory 904 is configured to store various types of data to support the operation of the device 900. Examples of such data include instructions for any application or method operating on the device 900, contact data, phone book data, messages, pictures, videos, etc. The memory 904 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0154] The power supply component 906 provides power to various components of the device 900. The power supply component 906 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 900.

[0155] The multimedia component 908 includes a screen that provides an output interface between the device 900 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 908 includes a front camera and / or a rear camera. When the device 900 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0156] The audio component 910 is configured to output and / or input audio signals. For example, the audio component 910 includes a microphone (MIC) that is configured to receive external audio signals when the device 900 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 904 or transmitted via the communication component 916. In some embodiments, the audio component 910 further includes a speaker for outputting audio signals.

[0157] The input / output interface 912 provides an interface between the processing component 902 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.

[0158] The sensor assembly 914 includes one or more sensors for providing a status assessment of various aspects of the device 900. For example, the sensor assembly 914 can detect the on / off state of the device 900, the relative positioning of components, such as the display and keypad of the device 900. The sensor assembly 914 can also detect a change in the position of the device 900 or a component of the device 900, the presence or absence of user contact with the device 900, the orientation or acceleration / deceleration of the device 900, and the temperature change of the device 900. The sensor assembly 914 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 914 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 914 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0159] The communication component 916 is configured to facilitate communication between the device 900 and other devices in a wired or wireless manner. The device 900 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 916 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 916 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0160] In an exemplary embodiment, the device 900 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above data annotation method.

[0161] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 904 including instructions, and the above instructions can be executed by a processor 920 of the device 900 to complete the above data annotation method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0162] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program executable by a programmable device, and the computer program has a code portion for performing the above-described data annotation method when executed by the programmable device.

[0163] Other embodiments of the present disclosure will be readily apparent to those skilled in the art after considering the specification and practicing the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0164] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A data annotation method, characterized in that, it includes: Obtain consecutive multiple frames of environmental images; Determine a dataset of images to be annotated through a pre-trained first perception model and a second perception model, where the dataset of images to be annotated includes multiple frames of environmental images to be annotated determined from the consecutive multiple frames of environmental images and target labels corresponding to each frame of the environmental images to be annotated; the detection accuracy of the first perception model is higher than that of the second perception model, and the detection speed of the second perception model is higher than that of the first perception model; For each frame of the environmental images to be annotated in the dataset of images to be annotated, perform data annotation on the frame of the environmental image to be annotated according to the target label corresponding to the frame of the environmental image to be annotated, to obtain a target annotated image.

2. The method according to claim 1, characterized in that, The determination of the dataset of images to be annotated through the pre-trained first perception model and second perception model includes: Perform frame extraction on the consecutive multiple frames of environmental images according to a preset frame extraction ratio to obtain multiple frames of first environmental images; Determine the dataset of images to be annotated through the first perception model and the second perception model according to the multiple frames of first environmental images and the consecutive multiple frames of environmental images.

3. The method according to claim 2, characterized in that, The determination of the dataset of images to be annotated through the first perception model and the second perception model according to the multiple frames of first environmental images and the consecutive multiple frames of environmental images includes: For each frame of the first environmental images, input the first environmental image into the first perception model and the second perception model respectively to obtain a first detection result output by the first perception model and a second detection result output by the second perception model; For each frame of environmental images in the consecutive multiple frames of environmental images, input the environmental image into the first perception model to obtain a third detection result output by the first perception model; Determine the dataset of images to be annotated according to the first detection result, the second detection result, and the third detection result.

4. The method according to claim 3, characterized in that, The detection result includes a detection area corresponding to a target object in the environmental image and a confidence level that the detection area is an image area corresponding to the target object; The determination of the dataset of images to be annotated according to the first detection result, the second detection result, and the third detection result includes: For each frame of the first environmental images, determine the degree of overlap between a first detection area of the target object in the first detection result and a second detection area of the target object in the second detection result according to the first detection result and the second detection result corresponding to the first environmental image; Determine the dataset of images to be annotated according to the degree of overlap, the confidence level in the first detection result, and the third detection result.

5. The method according to claim 4, characterized in that, Determining the degree of overlap between the first detection region of the target object in the first detection result and the second detection region of the target object in the second detection result according to the first detection result and the second detection result corresponding to the first environmental image includes: Calculating the intersection over union (IoU) of the first detection region and the second detection region; Using the IoU as the degree of overlap.

6. The method according to claim 4, wherein, the image dataset to be labeled includes at least one of a first subset, a second subset, a third subset, and a fourth subset, and determining the image dataset to be labeled according to the degree of overlap, the confidence level in the first detection result, and the third detection result includes: Determining the first subset and the second subset according to the degree of overlap; Determining the third subset according to the confidence level in the first detection result; Determining the fourth subset according to the third detection result.

7. The method according to claim 6, wherein, determining the first subset and the second subset according to the degree of overlap includes: Regarding a first image in the multiple frames of the first environmental image, where the degree of overlap is less than or equal to a first preset overlap threshold, as the environmental image to be labeled in the first subset, and recording the first detection result corresponding to the first image as the target label of the first image in the first subset; Regarding a second image in the multiple frames of the first environmental image, where the degree of overlap is less than or equal to a second preset overlap threshold, as the environmental image to be labeled in the second subset, and recording the first detection result corresponding to the second image as the target label of the second image in the second subset; wherein the first preset overlap threshold is greater than the second preset overlap threshold.

8. The method according to claim 6, wherein, determining the third subset according to the confidence level in the first detection result includes: Regarding a third image in the multiple frames of the first environmental image, where the confidence level is greater than or equal to a preset confidence threshold, as the environmental image to be labeled in the third subset, and recording the first detection result corresponding to the third image as the target label of the third image in the third subset.

9. The method according to claim 6, wherein, determining the fourth subset according to the third detection result includes: After performing target tracking on the target object on each frame of the environmental image according to the third detection result using a preset optical flow tracking algorithm, determining the tracking result; According to the tracking result, regarding a fourth image in the consecutive multiple frames of the environmental image, where the fourth image is an environmental image in which the target object is not detected according to the third detection result among every preset number of consecutive multiple frames of the environmental image, as the environmental image to be labeled in the fourth subset; Regarding the third detection results corresponding to the preset number of consecutive multiple frames of the environmental image as the target label corresponding to the fourth image; Recording the target label corresponding to the fourth image in the fourth subset.

10. The method according to claim 9, wherein, after performing target tracking on the target object in each frame of the environmental image according to the preset optical flow tracking algorithm based on the third detection result, the determined tracking result includes: for each preset number of consecutive frames of environmental images in the continuous multi-frame environmental images, according to the third detection result, performing target tracking on the target object in each adjacent two frames of environmental images in the preset number of consecutive frames of environmental images through the preset optical flow tracking algorithm, and then determining the tracking result; wherein, the tracking result indicates whether there is at least one frame of environmental image in the preset number of consecutive frames of environmental images that loses the target object in the third detection result.

11. The method according to claim 1, wherein, the method further includes: performing a duplicate removal operation on the to-be-annotated image dataset to obtain a target image dataset; the data annotation of each frame of the to-be-annotated environmental image according to the target label to obtain the target annotated image includes: performing data annotation on each frame of the to-be-annotated environmental image in the target image dataset according to the target label.

12. The method according to any one of claims 1-11, wherein, the method further includes: performing model training on the second perception model according to the target annotated image to obtain a target perception model.

13. A data annotation device, wherein, it includes: an acquisition module configured to acquire continuous multi-frame environmental images; a determination module configured to determine a to-be-annotated image dataset through a pre-trained first perception model and a second perception model, the to-be-annotated image dataset includes multiple frames of to-be-annotated environmental images determined from the continuous multi-frame environmental images and the target label corresponding to each frame of the to-be-annotated environmental image; the detection accuracy of the first perception model is higher than that of the second perception model, and the detection speed of the second perception model is higher than that of the first perception model; a annotation module configured to, for each frame of the to-be-annotated environmental image in the to-be-annotated image dataset, perform data annotation on the frame of the to-be-annotated environmental image according to the target label corresponding to the frame of the to-be-annotated environmental image to obtain a target annotated image.

14. A data annotation device, wherein, it includes: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to: execute the steps of the data annotation method according to any one of claims 1-12.

15. A computer-readable storage medium, on which computer program instructions are stored, wherein, when the program instructions are executed by a processor, the steps of the method according to any one of claims 1-12 are implemented.