Automatic driving perception data annotation quality evaluation and optimization method
By using Bayesian neural network and prior knowledge database in autonomous driving perception data annotation, quantifying label uncertainty and combining search enhancement technology, the problems of labeling accuracy and cost are solved, and efficient and reliable data annotation are achieved.
Patent Information
- Application Number
- CN202411918133.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-30
AI Technical Summary
The existing technology has insufficient sample analysis and screening in the labeling of autonomous driving perception data, resulting in reduced labeling accuracy, lack of scientific basis for labeling quality evaluation, high cost of data labeling, prominent problems in scarce data, and insufficient utilization of domain knowledge.
The Bayesian neural network is combined with a priori knowledge database, and the uncertainty is quantified through probabilistic modeling, and the samples are hierarchical screened to accurately distinguish low-quality and high-quality annotated data, and combined with retrieval enhancement technology to alleviate the limitations of scarce data on model performance.
It improves the accuracy and efficiency of data labeling, reduces the need for manual intervention, ensures the reliability of labeling results, optimizes resource allocation, and reduces labeling costs.
Smart Images

Figure CN120070934A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data annotation, and particularly to a method for evaluating and optimizing the annotation quality of autonomous driving perception data. Background Art
[0002] As a core component of autonomous driving, the perception system realizes a closed-loop from perception to decision-making in traffic scenarios. However, the accuracy and reliability of perception models highly depend on high-quality annotated data. The sources of autonomous driving perception data include multiple sensors, such as cameras, lidars, and millimeter-wave radars. These data integrate visual, depth, and motion information to support key tasks such as object detection, semantic segmentation, and scene understanding. How to efficiently and accurately annotate these data has become one of the main bottlenecks restricting the development of autonomous driving.
[0003] In response to the requirements of autonomous driving perception data annotation, scholars at home and abroad have proposed various methods, which can generally be classified into the following categories. The rule- and template-based method is an early mainstream technology that realizes the annotation of simple objects through preset logical rules and templates. For example, geometric features are used to annotate pedestrians or vehicles. This type of method has high efficiency and interpretability, but its flexibility is poor, and its performance significantly degrades in dynamic or complex scenarios (such as occlusion and multi-object overlap). In recent years, automatic annotation methods based on deep learning have made rapid progress, using technologies such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs) to perform object detection and semantic segmentation on sensor data. These methods can generate annotation results in diverse scenarios through data-driven feature learning. However, their annotation accuracy highly depends on the quality of training data, and the effect is poor in extreme scenarios (such as bad weather or strong light reflection). Some scholars have explored semi-automatic annotation methods that combine domain knowledge, and try to improve the annotation quality by introducing prior information (such as traffic rules or traffic sign distribution characteristics). For example, annotations are generated through preliminary model prediction and then fine-tuned by manual review. This type of method is applicable to specific scenarios and can improve accuracy, but its degree of manual participation is high and the cost is still large.
[0004] Although these methods have promoted the progress of automatic annotation, there are still significant limitations. First, the complexity of dynamic scenarios limits the annotation accuracy. In situations such as crowded traffic or multi-target interference, it is difficult for existing technologies to identify the boundaries of targets. For example, there are problems with the overlap between fast-moving vehicles and stationary objects. Second, the mechanism for evaluating the quality of annotations lacks a scientific basis. Existing methods mostly rely on the confidence of the model as the criterion for judging the quality of data. However, confidence fails to effectively characterize the uncertainty of annotations, resulting in a significant negative impact of low-quality samples on model training. Third, the cost of data annotation is high, and the problem of scarce data is prominent. Especially in special scenarios of driverless vehicles (such as tunnels or rare traffic signs), the acquisition of high-quality annotated data is extremely costly. Finally, the utilization of domain knowledge is insufficient. Many existing methods fail to effectively combine specific prior knowledge in the field, such as regional differences in road signs or vehicle characteristics, resulting in limited generalization ability.
[0005] After retrieval, the invention patent with Chinese patent number CN201910239565.8 discloses a method for evaluating the quality of image classification data annotation, including: providing an image data set, where the image data set includes images and classification data obtained after manually annotating each image; image feature extraction, extracting multiple feature vectors describing the color of each image based on the HSV channel of the image, and extracting multiple feature vectors describing the appearance of each image based on the local features of the image; measuring the degree of feature dispersion, using statistical analysis to model the degree of dispersion of the color and / or appearance feature vectors; automatic scoring, scoring and sorting the images based on the dispersion model obtained by modeling, and evaluating the classification data based on the sorting results.
[0006] Compared with the prior art, when the invention patent with Chinese patent number CN201910239565.8 is used, by modeling the degree of dispersion of the color features and appearance features of the annotated data, quantifying and scoring the quality of this pair of data annotations, and then achieving automatic evaluation of the quality of data annotation, providing a quantitative basis to assist manual evaluation and reducing the time cost.
[0007] However, in the actual use process of the above method, the analysis and screening of samples are weak, resulting in a decrease in the accuracy of data annotation. The present invention combines the probability modeling ability of the Bayesian neural network to quantify the uncertainty of annotations, classifies and screens the samples, accurately distinguishes low-quality and high-quality annotated data, and at the same time combines retrieval enhancement technology to alleviate the limitation of scarce data on the performance of the model, improving the annotation efficiency while ensuring the reliability of the annotation results. Summary of the Invention
[0008] The object of the present invention is to solve the problem in the prior art that the analysis and screening of samples are weak, resulting in a decrease in the accuracy of data annotation, and to propose an evaluation and optimization method for the quality of autonomous driving perception data annotation.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] An evaluation and optimization method for the quality of autonomous driving perception data annotation, comprising the following steps:
[0011] Step 1: Construct a prior knowledge database to provide more accurate prior information for the input data through image preprocessing and feature vector conversion;
[0012] Step 2: Use a Bayesian neural network to predict data annotation, generate the mean and variance of the prediction results through probability distribution prediction, use variance and entropy for scoring and measurement of the results generated by the Bayesian neural network, standardize the two and perform weighted averaging as the quality score to measure the model uncertainty;
[0013] Step 3: Evaluate the annotation quality according to the generated quality score.
[0014] The above technical solution further includes:
[0015] Perform image preprocessing operations on the original data in the database and construct the database. Convert the input data into the form of feature vectors through a pre-trained transformer module, use the Euclidean distance as the similarity measurement method, find the data in the database that is most similar to the input data and input them into the model together, and perform a series of preprocessing operations on the input data, including denoising, data segmentation, and standardization, to reduce the impact of noise on subsequent feature extraction;
[0016] Extract features from the preprocessed data. At this time, the model will automatically extract high-level features required for data annotation from the input perception data;
[0017] Use a pre-trained deep learning model to extract features. The deep learning model is a convolutional neural network and a Transformer model. Construct a pre-trained model and then use a Bayesian neural network for uncertainty modeling. Perform uncertainty modeling on each layer of parameters through the Bayesian neural network, so that the model can generate probability-based prediction results and give scores.
[0018] An optimization method for evaluating the quality of autonomous driving perception data annotation, comprising the following steps:
[0019] S1: Calculate the variance and entropy of the deviation between the annotation data and the standard data;
[0020] S2: Standardize the variance and entropy, and perform weighted fusion as the final quality score.
[0021] Classify the labeled samples according to the quality score, divide the samples into high-score samples and low-score samples, directly upload the high-score samples to the dataset, and the low-score samples enter the manual review process.
[0022] By sorting the low-score samples according to the priority, the samples with the lowest scores are processed first and manually reviewed;
[0023] Among them, on this basis, a Bayesian neural network is used to predict the data annotation. The Bayesian neural network generates multiple prediction results by dealing with the uncertainty of the input data, so that the annotation is not only deterministic, but a result based on probability.
[0024] The preprocessed data will be fed into a deep neural network that has been pre-trained on a large-scale dataset for feature extraction. At this time, the model will automatically extract useful high-level features from the input perceptual data, and then provide necessary information for data annotation.
[0025] To introduce the handling of model uncertainty, the present invention transforms these pre-trained models into Bayesian neural networks. In a Bayesian neural network, the parameters of the model are no longer fixed values, but random variables drawn from a certain probability distribution. By introducing Bayesian regularization layers in each layer of the neural network, these layers learn the posterior distribution of each parameter, enabling the network to estimate the parameter uncertainty of each layer. The Bayesian neural network can generate a probability distribution, rather than just a deterministic prediction result, so that the model can provide multiple possible prediction results for different input samples. When generating data annotation, the Bayesian neural network generates multiple different prediction results by sampling the model parameters multiple times. Through this sampling method, we can provide multiple prediction results for each input sample, rather than just a single label. These multiple prediction results will be statistically analyzed according to the mean and variance, and finally a mean annotation representing the annotation result will be generated.
[0026] Quality scoring and quantification. For the generation of multiple prediction results, the uncertainty is measured by calculating the variance and entropy of these results. The variance reflects the volatility of the prediction results, while the entropy represents the confidence of the model in classification. The smaller the entropy value, the stronger the confidence of the model in the annotation. To unify the scoring range, the variance and entropy of all samples are first standardized, and then they are weighted and calculated according to a certain weight to obtain a comprehensive quality score. Finally, this score can assign a quantified quality index to each sample, where samples with low uncertainty indicate high annotation quality, while samples with high uncertainty may have annotation problems.
[0027] Sample Classification and Resource Optimization Processing Based on Quality Scoring. Quality scores are assigned to each sample, and based on the scoring results, the samples are divided into two categories: high-score samples and low-score samples. High-score samples indicate higher annotation quality and lower prediction uncertainty, so they can be directly added to the dataset, eliminating the need for manual intervention and significantly reducing the annotation time and labor costs. Low-score samples, due to the high uncertainty in annotation, are marked as "pending review" and enter the manual review process. The system sorts the low-score samples according to the scores and gives priority to processing the samples with the lowest scores. The manual annotators review these samples and correct inaccurate annotation information or supplement missing annotation content.
[0028] The present invention has the following beneficial effects:
[0029] 1. In the present invention, by uniformly preprocessing the original data in the database and combining with the pre-trained transformer module to generate feature vectors, similarity measurement is carried out through Euclidean distance, which effectively improves the matching degree between the input data and prior knowledge and enhances the accuracy and efficiency in the data annotation task.
[0030] 2. In the present invention, the present invention innovatively considers uncertainty based on the Bayesian neural network to improve the accuracy and efficiency of data annotation. By introducing the Bayesian neural network, the model can model the uncertainty in the data annotation process, generate multiple prediction results, and accurately measure the quality of each annotation through the quantification of variance and entropy, thus effectively improving the reliability of annotation.
[0031] 3. In the present invention, through the resource optimization scheme that combines automation and manual review, samples are classified based on quality scores. High-score samples are directly added to the dataset, and low-score samples enter the manual review process and are processed in order of priority. This scheme not only reduces the need for manual intervention but also ensures that low-quality data is corrected in a timely manner, optimizes resource allocation, and improves the data annotation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a flowchart of a method for evaluating and optimizing the quality of autonomous driving perception data annotation proposed by the present invention;
[0033] Figure 2 is a schematic diagram of a retrieval enhancement structure based on a knowledge base;
[0034] Figure 3 is a schematic diagram of a Bayesian neural network model structure;
[0035] Figure 4 is a schematic diagram of the overall process of this method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0037] Embodiment 1:
[0038] For a large number of special scenarios in autonomous driving, a database is constructed for retrieval enhancement to overcome the model's understanding of special scenarios, including the following steps:
[0039] 1. A large amount of data on special scenarios is collected and stored in the database.
[0040] 2. All data is converted into high-dimensional feature vectors using a pre-trained Transformer module.
[0041] 3. An efficient vector index is created using FAISS, enabling the rapid retrieval of data most similar to a given query from a large amount of data.
[0042] 4. When given data is used as a query input, high-dimensional feature vectors are also extracted through the Transformer module, and the most similar data in the database is quickly found using the FAISS index.
[0043] Embodiment 2:
[0044] To solve the uncertainty problem in autonomous driving data annotation, the present invention designs a method based on a Bayesian neural network, which can automatically generate high-quality annotations according to perception data. This method can quantify uncertainty during the automatic annotation process, thereby improving the accuracy and efficiency of data annotation.
[0045] It includes the following steps:
[0046] 101: Data preprocessing and feature extraction.
[0047] Among them, a series of preprocessing operations need to be performed on the original perceptual data. These steps include data cleaning, denoising, segmentation, and normalization, etc., aiming to ensure the quality of the input data so that the subsequent deep learning model can efficiently extract useful features and perform annotation. In the preprocessing stage, first, useless or abnormal image data are removed through data cleaning, and then filtering techniques are used to remove noise in the images. Next, different regions in the images (such as roads, pedestrians, vehicles) are separated through image segmentation technology, so as to extract more representative parts. Finally, through normalization processing, the pixel values of the images are unified to the same scale, so that the subsequent deep learning model can efficiently process and analyze these image data.
[0048] After these preprocessing operations are completed, feature extraction will be carried out. In this stage, convolutional neural networks and Transformer models are often used to extract high-dimensional features from the preprocessed data. Convolutional neural networks perform local feature extraction on image data through convolution operations, and can effectively capture important features such as edges, textures, and shapes in the images, which is crucial for object recognition and object detection. The Transformer model captures global features in the data through the self-attention mechanism and can identify the relationships between objects in the images. These feature extraction methods combine local details and global information, and we splice them and then output the annotation results through a fully connected layer.
[0049] 102: Bayesian neural network conversion.
[0050] Among them, in the process of autonomous driving data annotation, in order to handle the uncertainty problem of the model, the present invention converts a pre-trained deep learning model into a Bayesian neural network.
[0051] Traditional deep learning models usually output a deterministic prediction result, while Bayesian neural networks model the uncertainty of the parameters of each layer of the network, enabling the model to output prediction results with a probability distribution rather than a single deterministic value.
[0052] The core advantage of this conversion is that Bayesian neural networks can quantitatively handle the prediction uncertainty when processing different data, so that the model can adjust the credibility of its predictions according to the degree of uncertainty of the data.
[0053] Bayesian neural networks model the parameters of each layer by introducing a normal distribution, so that each parameter of the network is no longer a fixed value, but a random variable sampled from a probability distribution. Suppose the parameter W of the l-th layer (l) obeys a normal distribution with a mean of u (l) and a standard deviation of σ (l) :
[0054] W (l)~N(u (l) ,σ (l) ),
[0055] In the above formula, u (l) and σ (l) represent the mean and standard deviation of the parameters of the l-th layer respectively, representing the posterior distribution of the weights of that layer. By sampling this distribution, we obtain a set of parameters These different sampled values represent different possible states of the weights of that layer.
[0056] By sampling multiple times from the posterior distribution, we can obtain multiple prediction results. These prediction results are not just single values, but a distribution, which enables the model to estimate the uncertainty of its predictions. For example, for an input sample x, the prediction results obtained by sampling multiple times are The mean of these prediction results and variance can be used to quantify the uncertainty of the prediction.
[0057] The formula for calculating the mean is:
[0058] The formula for calculating the variance is:
[0059] where is the mean of all sampling results, representing the final prediction output, while is the variance, representing the uncertainty of the prediction result. If the variance is large, it indicates that the model has a high uncertainty in predicting this sample, usually meaning that this sample requires further manual review or more complex automated correction methods. Through this uncertainty quantification, the Bayesian neural network can not only generate prediction results, but also provide a measure of the uncertainty of the predictions. For those samples with high uncertainty, the system will mark them as samples that need to be reviewed, while for samples with low uncertainty, the final annotation results will be directly generated.
[0060] In summary, this process improves the accuracy of annotation, reduces the workload of manual annotation, and ensures the reliability of the annotated data.
[0061] Example 3
[0062] To quantify the quality of each annotated sample, the present invention assigns a quality score Q i to each sample by calculating the uncertainty measure of the sample. Lower uncertainty usually indicates higher annotation quality, while higher uncertainty may reflect potential problems in the annotation results. The construction process of the quality score includes the following steps:
[0063] Step 1: Selection and calculation of uncertainty quantification indicators. For each sample xi , generate M prediction results through multiple samplings of the Bayesian neural network Based on these prediction results, select the variance and entropy H as quantitative indicators for measuring uncertainty. The formula for calculating entropy is:
[0064] H = ∑P(y k |x k ) log P(y k |x k )
[0065] where P(y k |x k ) represents the prediction probability of the model for k category k. The smaller the entropy, the more confident the model is in classifying the sample.
[0066] Step 2: Calculate and screen the comprehensive index weights. To ensure the consistency of the scoring range, standardize the variances and entropies of all samples. The standardization formula is:
[0067]
[0068]
[0069] Then, perform a weighted calculation of the standardized variances and entropies to obtain the comprehensive score. The calculation formula is as follows:
[0070]
[0071] where w 1 and w 2 are the weights of the variance and entropy respectively, satisfying w 1 + W 2 = 1. This method can integrate various uncertainty information to generate a more reliable quality score.
[0072] Example 4
[0073] Set the threshold Q th resh old according to the application requirements, and divide all samples into the following two categories.
[0074] High-score samples: The score is higher than Q th resh old , indicating higher annotation quality and lower prediction uncertainty.
[0075] Low-score samples: The score is lower than Q th resh old , indicating lower annotation quality and higher prediction uncertainty. The two types of samples enter different processing flows respectively to achieve efficient resource allocation.
[0076] For high-scoring samples, the system considers that their annotation quality meets the training requirements, so these samples are directly added to the dataset. The advantage of this process is that it requires no manual intervention, significantly reducing the annotation time and labor costs. For low-scoring samples, due to the high uncertainty in their annotations, the system marks them as "pending review" and enters the manual review process to further improve their quality.
[0077] According to the quality score Q i , sort the low-scoring samples from low to high, and give priority to processing the samples with the lowest scores. And use manual annotators to review these samples, correcting inaccurate annotation information or supplementing missing annotation content.
[0078] The following combines specific datasets to verify the solutions in Examples 1-4. See the following description for details:
[0079] The task of recognizing and classifying traffic signs is selected as a representative task. In an autonomous driving system, accurately recognizing traffic signs is an important part of ensuring driving safety. Traffic signs in different countries and regions have different colors, shapes, and symbols. Annotating these traffic sign data is the key to training an autonomous driving classification model. However, there are a wide variety of traffic signs, which often overlap with other objects (such as buildings, trees), resulting in occlusion and blur problems. Some special traffic signs appear very rarely, and the annotated data is relatively scarce. We constructed a dataset including 3,402 samples through collection. A total of 3,589 samples, including this self-built dataset, were constructed from multiple channels on the network, and 792 samples were used as the annotated test set.
[0080] To fairly evaluate the performance of all methods, we conducted two aspects of evaluation. On the one hand, we used the classification accuracy (Acc) and mean average precision (mAP) measurement criteria to generate accuracy. On the other hand, we used the percentage (Per) of the number of samples that require manual review to measure the consideration of this invention for saving labor costs in automatic annotation. The results are shown in Table 1:
[0081] Table 1 Comparison of quantitative indicators of related methods
[0082]
[0083] The present invention selects the following three comparison methods for verification:
[0084] Traditional deep learning method: a standard method based on convolutional neural network without uncertainty modeling. Ensemble learning method: improving robustness through the integrated prediction of multiple independent models, but without introducing Bayesian modeling. Prior-free enhancement method: only using Bayesian neural network without an enhancement module for prior knowledge retrieval. Table 1 shows the evaluation results of relevant quantitative indicators. This table comprehensively demonstrates the performance comparison of the methods in terms of classification accuracy, annotation quality, and labor cost savings, highlighting the comprehensive advantages of the method of the present invention.
[0085] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for evaluating the quality of autonomous driving perception data annotation, characterized in that: The following steps are involved: Step 1: Build a prior knowledge database to provide more accurate prior information for input data through image preprocessing and feature vector conversion; Step 2: Use the Bayesian neural network to predict the data annotation, and use the probability distribution to predict the mean and variance of the generated results. The results generated by the Bayesian neural network are scored using variance and entropy. The two are standardized and weighted averaged as a quality score to measure model uncertainty. Step 3: Evaluate the annotation quality based on the generated quality score.
2. The method for evaluating the quality of autonomous driving perception data annotation according to claim 1, characterized in that: The original data in the database is subjected to image preprocessing and the database is constructed. The input data is converted into feature vector form through the pre-trained transformer module, and the Euclidean distance is used as the similarity measurement method to find the data in the database that is most similar to the input data and input them into the model together.
3. The method for evaluating the quality of autonomous driving perception data annotation according to claim 2, characterized in that: A series of preprocessing operations are performed on the input data, including denoising, data segmentation and normalization, to reduce the impact of noise on subsequent feature extraction.
4. The method for evaluating the quality of autonomous driving perception data annotation according to claim 3, characterized in that: Feature extraction is performed on the preprocessed data. At this time, the model will automatically extract the high-level features required for data labeling based on the input perception data.
5. The method for evaluating and optimizing the quality of autonomous driving perception data annotation according to claim 4, characterized in that: Use pre-trained deep learning models to extract features. The deep learning models are convolutional neural networks and Transformer models. After building the pre-trained model, Bayesian neural networks are used for uncertainty modeling. The uncertainty modeling of each layer parameter is performed through the Bayesian neural network, so that the model can generate probability-based prediction results and give scores.
6. The optimization method for evaluating the quality of autonomous driving perception data annotation according to claim 5, characterized in that: The following steps are involved: S1: Calculate the variance and entropy of the deviation between the labeled data and the standard data; S2: The variance and entropy are normalized and weighted fusion is performed as the final quality score.
7. The optimization method for evaluating the quality of autonomous driving perception data annotation according to claim 6, characterized in that: The labeled samples are classified according to the quality scores, and the samples are divided into high-scoring samples and low-scoring samples. The high-scoring samples are directly uploaded to the dataset, and the low-scoring samples enter the manual review process.
8. The optimization method for evaluating the quality of autonomous driving perception data annotation according to claim 7, characterized in that: By prioritizing the low-scoring samples, the lowest-scoring samples are processed first and manually reviewed.
Citation Information
Patent Citations
A method for evaluating the quality of image classification data annotation
CN111652258B
Cited By
A method and system for verifying the consistency of velocity direction in 3D obstacle annotation data
CN122574829A