System and method for embedding uncertainty estimation into deep neural network-based automatic driving perception framework
By embedding uncertainty estimation in the autonomous driving perception model, the problem of the perception model predicting uncertainty in complex environments is solved, and the safety and decision-making accuracy of the autonomous driving system are improved.
Patent Information
- Application Number
- CN202510286309.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-11
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-11
AI Technical Summary
Existing autonomous driving perception models fail to effectively estimate and handle uncertainty in predictions, resulting in incorrect or unsafe decisions that may be made in complex environments.
Embed uncertainty estimation in the perceptual model, by adding confidence branch to each prediction head, combining the confidence score for correction of the prediction output and calculation of the loss function, the regularized loss term is dynamically adjusted to punish low confidence predictions.
It improves the safety and reliability of the perceptual model, can more accurately identify uncertain areas in the environment, and improves the decision-making quality of the autonomous driving system.
Smart Images

Figure CN120298988A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure are generally related to autonomous driving. More specifically, the systems of the present disclosure relate to embedding uncertainty estimation into a deep neural network-based autonomous driving perception framework. Background Art
[0002] Autonomous driving perception is a key aspect of an autonomous driving system as it enables an autonomous vehicle to make informed decisions based on real-time data of the environment. Autonomous driving perception can include a large number of individual perception tasks, each of which makes an important contribution to the overall efficiency of the autonomous driving system. For example, environmental perception is a task that provides key information about the driving environment for the vehicle, including free drivable areas and the positions, speeds, and even predictions of the future states of surrounding obstacles.
[0003] There has been a large amount of research work in the field of machine learning-based autonomous driving perception. However, most of this work has ignored the impact of uncertainty in predictions. Given safety and reliability considerations, estimating uncertainty in each perception task is crucial for the practical application of autonomous driving. Summary of the Invention
[0004] Embodiments of the present disclosure may provide a system and method for training a perception model that performs autonomous driving tasks. In an implementation, the system may obtain training data containing annotations of images captured by multiple cameras installed at different positions on a vehicle, and the perception model may generate task-related prediction outputs and confidence scores in parallel based on the annotated training data. The confidence scores may characterize the level of uncertainty associated with the prediction outputs. The system may generate an uncertainty-weighted prediction based on the ground truth, prediction outputs, and confidence scores characterized by the annotated training data; calculate a loss function based on the uncertainty-weighted prediction; and update the perception model based on the loss function.
[0005] In a variation of this embodiment, the perception model may include a Bird’s Eye View (BEV)-based perception model.
[0006] In a variation of this embodiment, the prediction outputs may include classification predictions and / or regression predictions.
[0007] In a variation of this embodiment, generating an uncertainty weighted prediction may include calculating a linear combination between the road ground truth and the prediction output, where the prediction output is weighted according to the confidence score.
[0008] In a variation of this embodiment, calculating the loss function may further include adding a regularization loss term that can be determined according to the confidence score and hyperparameters.
[0009] In another variation, the hyperparameters may be dynamically adjusted during the training process according to the upper limit value of the regularization loss term.
[0010] In another variation, the system may decrease the hyperparameters when the regularization loss term is greater than or equal to the upper limit value, and increase the hyperparameters when the regularization loss term is less than the upper limit value.
[0011] In another variation, the system may select a subset of the labeled training data and associate the prediction output generated based on the selected subset of the labeled training data with a static confidence score. Description of the Drawings
[0012] Figure 1 Shows an example scenario where the perception model according to the prior art does not consider the uncertainty in the prediction.
[0013] Figure 2 Shows an example architecture of a perception framework with embedded uncertainty estimation according to an embodiment of the present application.
[0014] Figure 3 Shows an example pseudocode for embedding uncertainty estimation into a bird's-eye view (BEV) perception framework according to an embodiment of the present application.
[0015] Figures 4A - 4C Shows an example scenario of map vectorization using a perception model enhanced with uncertainty according to an embodiment of the present application.
[0016] Figures 5A - 5B Shows an example scenario of object detection using a perception model enhanced with uncertainty according to an embodiment of the present application.
[0017] Figure 6 Shows an example scenario of semantic segmentation using a perception model enhanced with uncertainty according to an embodiment of the present application.
[0018] Figure 7 Shows the performance metrics of the perception models with and without uncertainty considerations according to an embodiment of the present application.
[0019] Figure 8 FIG. 2 shows an exemplary block diagram of an uncertainty-enhanced perception system for autonomous driving according to an embodiment of the present application.
[0020] Figure 9 FIG. 6 presents a flowchart showing an exemplary training process of an uncertainty-enhanced machine learning model according to an embodiment of the present application.
[0021] Figure 10 FIG. 10 shows an exemplary computer system for an uncertainty-enhanced perception system according to an embodiment of the present application.
[0022] In the figures, the same reference numerals denote the same components. DETAILED DESCRIPTION
[0023] The following description is intended to enable those skilled in the art to make and use the disclosed embodiments and is provided in the context of one or more specific applications and their requirements. Various modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the scope of the disclosed content. Therefore, the present invention or various aspects thereof are not intended to be limited to the embodiments shown, but should be accorded the widest scope consistent with the disclosed content.
[0024] Overview
[0025] Embodiments of the present disclosure provide a system and method for embedding uncertainty estimation into a deep neural network-based autonomous driving perception framework. Different from traditional perception models that ignore the degree of uncertainty in their predictions, the proposed perception network can embed uncertainty estimation from the training stage. More specifically, each prediction head in the perception network can add an additional confidence dimension to characterize the network's confidence in its corresponding prediction. During the training process, the predictions used to generate the loss can be corrected according to their corresponding confidence scores to make them closer to the Ground Truth (GT) targets. In some embodiments, uncertainty-weighted predictions can be generated by interpolating between the GT targets and the original predictions according to the confidence scores of the predictions. The perception model can also include a regularization loss for penalizing low-confidence predictions. Additionally, the regularization loss can be dynamically adjusted according to an upper bound value.
[0026] Uncertainty - Augmented Perception Models
[0027] In the field of autonomous driving, models based on deep neural networks are used to predict environmental information around autonomous vehicles (ego vehicles, e.g., vehicles equipped with various sensors), including features such as lanes, curbs, sidewalks, intersections, road paths, vehicles, pedestrians, etc., as well as various dynamic parameters such as speed estimation and motion. These prediction results can be utilized by downstream planning and control tasks. Traditional perception models often ignore the inherent uncertainty in perception tasks. However, autonomous vehicles typically operate in complex and dynamic environments, and significant uncertainty may arise in perception due to factors such as obstacles, limited perspectives, and the unpredictability of the behavior of other road users. When considering downstream tasks such as planning and control that rely on the perception output as the basic input in autonomous driving, the uncertainty in perception becomes particularly important.
[0028] Figure 1 An example scenario according to the prior art is shown, where the perception model does not consider the uncertainty in prediction. In this example, the perception model can be based on a Bird’s-Eye-View (BEV) framework that can extract an overall representation of the environment from multi-camera images. In Figure 1 the example shown, six images (e.g., Image 102 - Image 112) captured by cameras installed at different positions (e.g., the front side, the rear side, and the sides) of an autonomous vehicle can be used to generate a map vectorization prediction 114 that shows a road map. It is worth noting that Image 102 - Image 112 correspond to the road surface view or perspective view of the environment observed from different perspectives, while the map vectorization prediction 114 corresponds to the BEV of the environment. The lines of different colors in the prediction 114 correspond to different road objects (e.g., lane lines, curbs, crosswalks, etc.).
[0029] Figure 1 Occlusion regions 116 and 118 (located in Image 104 and Image 106 respectively) are also shown. In other words, due to the occlusion of the camera's field of view, regions 116 and 118 cannot be clearly seen, which may lead to an increase in uncertainty in the corresponding region 120 in the prediction 114. However, existing perception models usually treat all regions in the prediction 114 equally, ignoring the fact that region 120 has a high degree of uncertainty due to occlusion. This lack of discrimination in perception uncertainty may lead to incorrect or unsafe decisions in subsequent planning and control tasks.
[0030] Just as a human driver slows down due to uncertainty about the situation ahead when driving in fog, an autonomous driving system needs to measure the confidence of its perception results. Embedding uncertainty estimation into a deep neural network-based autonomous driving perception framework can not only improve safety by providing the system with a measure of its own prediction confidence, but also serve as an important input to the upper-level planning and control system. The improved autonomous driving system can thus make more informed and context-appropriate decisions, similar to how a human driver adjusts their driving style based on the degree of certainty about the surrounding environment.
[0031] To embed prediction uncertainty into the perception model, some embodiments of the present invention can modify each prediction head in the deep learning neural network to include an additional confidence branch or sub-task for predicting the confidence score of the current prediction. Additionally, when calculating the loss, regression and classification predictions can be combined with the confidence score (e.g., by interpolating with the ground truth).
[0032] Figure 2 An example architecture of a perception framework with embedded uncertainty estimation according to an embodiment of the present application is shown. In Figure 2 this, the multi-view image 202 can be sent to the BEV feature extraction unit 204 for feature extraction. The multi-view image 202 can include images captured by cameras mounted in front of, behind, and on both sides of the vehicle itself (i.e., the autonomous vehicle). In one example, six images (similar to images 102 - 112) can be sent to the BEV feature extraction unit 204. The BEV feature extraction unit 204 can extract BEV features 206 from these six images. In one embodiment, the BEV feature extraction unit 204 can use multiple mechanisms to convert the two-dimensional image features in the multi-view image 202 into BEV features. For example, the inverse perspective mapping (IPM) algorithm can be used to convert the perspective view of the traffic scene into BEV.
[0033] The BEV features 206 (which can exist in the form of a BEV feature map) can then be sent to a prediction head 208 that can perform specific tasks in autonomous driving, such as 3D object detection or semantic segmentation. It is important to note that only one prediction head is shown in the Figure 2 example shown. However, in practical applications, the perception framework can include multiple prediction heads, each performing a specific perception task and making predictions related to that task.
[0034] In one embodiment, the prediction head 208 can include a map decoder that can query the BEV feature map to output instance latent features, Figure 2(not shown). Then, these instance latent features can be sent to different prediction branches, each of which performs a subtask. Conventional prediction heads for 3D object detection typically include a classification branch for predicting object classification and a regression branch for predicting object location and pose. In Figure 2 In the example shown, each prediction head can embed uncertainty estimation in its predictions by including an additional confidence branch for predicting the confidence scores of its classification and regression predictions. In Figure 2 In the example shown, the prediction head 208 includes a classification branch 210, a confidence branch 212, and a regression branch 214. These three branches can operate in parallel based on the same BEV feature input. In some embodiments, the confidence score can be represented as a percentage value, where 0% indicates no confidence and 100% indicates full confidence in the prediction result.
[0035] The independent prediction outputs of the classification branch, the confidence branch, and the regression branch can be represented as , and . In some embodiments, to distinguish high-confidence predictions from low-confidence predictions, during training of the classification and regression models (e.g., deep learning neural networks), in each training iteration, the outputs of the classification branch 210 and the regression branch 214 can be combined with the confidence scores generated by the confidence branch 212. More specifically, during the backpropagation process, predictions with high confidence scores can play a more important role in generating the loss function than predictions with low confidence scores. In some embodiments, the prediction results can be combined with the ground truth according to the confidence scores. In one embodiment, a linear combination (e.g., convex combination) or interpolation technique can be used to interpolate the prediction results with the ground truth using the confidence scores as weight factors to generate uncertainty-weighted predictions.
[0036] In one embodiment, the uncertainty-weighted classification prediction can be based on generated, where represents the uncertainty-weighted prediction, and represents the ground truth of the classification. In the extreme case where the confidence score is 0, the prediction output of the classification branch 210 (e.g., ) will not be included in the loss calculation. Instead, the loss calculation will only use the ground truth of the classification. On the other hand, when the confidence score is 100%, the loss will be calculated only based on . In most cases, when the confidence score is between 0 and 100%, the loss calculation can be based on , which is a weighted combination of the prediction and the ground truth. Similarly, the uncertainty-weighted prediction for regression can be based on Generation
[0037] Experiments show that the perception model produces high uncertainty (or low confidence scores) for regions that are occluded due to limited viewing angles, too far away, or not visible, while producing low uncertainty (or high confidence scores) for clearly visible regions. These confidence scores enable the perception model to identify invisible, occluded, or blurred regions. Therefore, during the training process, the model can dynamically reduce the weights of these regions (e.g., low confidence scores mean lower weights), thus preventing the model from overfitting to road ground truths that are invisible or too blurred to learn.
[0038] Figure 3 An example pseudocode for embedding uncertainty estimation into a bird's-eye view (BEV) perception framework according to an embodiment of the present application is shown. In Figure 3 the example shown, the bird's-eye view perception task 300 can be an object detection task that outputs classification and regression predictions. More specifically, each iteration in the perception task 300 can include a prediction stage 302, an uncertainty embedding stage 304, and a backpropagation stage 306.
[0039] The prediction stage 302 can include parallel operations for predicting regression ( ), classification ( ), and confidence scores ( ). The uncertainty embedding stage 304 can include applying the confidence scores to the regression and classification predictions. In this example, the uncertainty-weighted predictions for regression and classification can be generated as a linear combination (or interpolation) between the predicted values and the road ground truth values, where the weight coefficients are determined based on the confidence scores. Other combination or interpolation techniques (e.g., non-linear interpolation) can also be used to generate the uncertainty-weighted predictions.
[0040] In the backpropagation stage 306, the loss function can be calculated based on the uncertainty-weighted predictions. The loss function can be of any type, including but not limited to the mean squared error (MSE) loss function, cross-entropy loss function, focal loss function, logarithmic loss function, etc. The scope of the present disclosure is not limited by the type of loss function used during the training process. The deep learning neural network is trained to minimize the loss function. Various optimization algorithms can be employed when training the neural network. In some embodiments, the gradient descent method (e.g., Adam algorithm) and the backpropagation algorithm can be used to train the network.
[0041] If the confidence scores are directly used in the training process (e.g., in loss calculation), it may cause the model to overly favor low-confidence data because the low-confidence data will be replaced by the ground truth data, thus minimizing the training loss. To address this issue, a structured training strategy can be adopted. More specifically, a regularization term can be added to the original loss function to enhance confidence. In some embodiments, the regularization term in the loss function (e.g., ) can be calculated based on the negative logarithm of the confidence score. For example, , where is a hyperparameter and can be dynamically adjusted during training. For example, if the original loss function is based on MSE, the uncertainty-enhanced loss function can be . In this way, very low confidence scores (e.g., C close to 0) can result in very high losses. The regularization loss term is designed to penalize low-confidence predictions, thus encouraging the perception model to generate high-confidence predictions as much as possible.
[0042] In some embodiments, to encourage learning in an uncertain environment, a random selection mechanism can be used to select a subset of the data to embed uncertainty estimation, while the other data is considered trustworthy (e.g., assigned a static confidence score of 100%), rather than applying uncertainty estimation (e.g., confidence scores) to all training data. In some embodiments, for each batch of data, the system can randomly select a portion of the data batch and embed uncertainty for it during training (e.g., predicting confidence scores in each training iteration). The proportion of the randomly selected data (denoted as ) can be an adjustable hyperparameter between 0 and 1 / 2. In one example, during training, approximately 1 / 4 of the data batch can be selected to embed uncertainty estimation. In an alternative embodiment, for each data batch, the system can randomly select a subset of the data and assign a static confidence score (e.g., 100%) to the predictions made based on the selected subset of data. The network will predict confidence scores for the prediction results made based on the other data.
[0043] The uncertainty in the prediction may change during the training process. A trained model usually has a lower uncertainty level than a model that has not been trained yet. In some embodiments, the training of the perception model can implement an upper bound on the loss related to uncertainty ( ) to limit the maximum loss caused by skewed confidence values. More specifically, the loss term related to uncertainty or the weight of the regularization loss can be adjusted according to the set upper bound. If the regularization loss is equal to or greater than the set upper bound , which can be achieved by reducing the hyperparameters used to calculate the regularization loss To slightly reduce the weight of the confidence score. In some embodiments, during training, the confidence score can be slightly reduced according to the additional hyperparameter To dynamically adjust hyperparameters ,in .if ,but (For example, ). On the contrary, if ,but (For example, ). In this way, no matter the hyperparameter Regardless of the initial value of is maintained within the predetermined range.
[0044] Once a perception model is trained, it can be used to perform a variety of perception tasks. In one example, a perception model can be trained to perform map vectorization, which refers to the process of building a map based on an image and converting each map element into a vector. In addition to predicting vectorized map elements (such as lanes, curbs, stop lines, crosswalks, etc.), the perception model can also output confidence scores for these predicted map elements. In one example, each predicted map element can be associated with a predicted confidence score. Therefore, downstream planning and control tasks that rely on these predictions can take into account the level of uncertainty associated with each prediction.
[0045] Figures 4A - 4C An example scenario of performing map vectorization using an uncertainty-enhanced perception model according to an embodiment of the present application is shown. Figure 4A shows an image of a BEV captured by the autonomous vehicle’s camera, Figure 4B shows the real data of the vectorized map, and Figure 4C The vectorized map predicted by the uncertainty-enhanced perception model is shown. Figure 4B and Figure 4C In , different colored lines represent different types of map elements. For example, green lines represent curbs, while red lines represent lane lines. Figure 4C , the confidence scores associated with predicted map elements are represented by the brightness of the lines. More specifically, brighter line segments (e.g., line segment 402) are predictions with higher confidence scores (or lower uncertainty levels), while darker line segments (e.g., line segment 404) are predictions with lower confidence scores (or higher uncertainty levels). Figure 4A and Figure 4CAs shown, the nearer and visible regions (e.g., the regions near line segment 402) generally have lower levels of uncertainty, while the farther and occluded regions (e.g., the regions near line segment 404) have higher levels of uncertainty.
[0046] Figures 5A - 5B Fig. shows an example scenario of object detection using an uncertainty-enhanced perception model according to an embodiment of the present application. Figure 5A Fig. shows a BEV image captured by a camera of an autonomous vehicle, while Figure 5B shows the results of object detection. As can be seen from Figure 5B it, the regions closer to the autonomous vehicle or within the red box (e.g., region 502) have lower levels of uncertainty, and their bounding boxes are represented by brighter lines, while the regions farther from the autonomous vehicle (e.g., region 504) have higher levels of uncertainty, and their bounding boxes are represented by darker lines.
[0047] Figure 6 Fig. shows an example scenario of semantic segmentation using an uncertainty-enhanced perception model according to an embodiment of the present application. In this example, the left figure shows a confidence map, where the confidence scores are mapped to grayscale pixel values, and the higher the confidence score, the higher the grayscale pixel value. More specifically, white can represent 100% confidence, and black can represent 0% confidence. The right figure shows the results of semantic segmentation, where the green lines represent the segmented curbs and the blue lines represent the segmented lane lines. Similar to Figs. 4C and 5B, the uncertainty associated with the predicted output can be reflected by the brightness of the lines. In Figure 6 the example shown, the darker lines in region 602 of the right figure correspond to the low-confidence region 604 in the left figure.
[0048] The uncertainty-enhanced perception model can be widely applied to various autonomous driving tasks. In addition to map vectorization, object detection, and semantic segmentation, the uncertainty-enhanced perception model can also be used for other tasks such as path prediction, motion detection, speed estimation, etc. Furthermore, the present disclosure takes the BEV framework as an example, where the BEV image is used as the input. In fact, other types of perception frameworks or architectures can also adopt similar methods to embed the uncertainty estimation in the prediction into the learning and training process of the model. For example, machine learning models for robot control can also utilize the uncertainty-enhanced model to improve the perception ability of the robot. The scope of the present disclosure is not limited by the architecture and application of the perception model.
[0049] Experimental results show that embedding uncertainty estimation into the perception model can effectively improve the performance of the model. Figure 7 Fig. shows the performance metrics of the perception model with and without uncertainty consideration according to an embodiment of the present application. InFigure 7 In [Table 700], each number can represent the average accuracy value of a specific type of map element (e.g., lane lines, curbs, stop lines, and crosswalks). More specifically, the numbers in row 702 represent the average accuracy values of the predictions generated by the baseline model (i.e., the perception model that does not consider prediction uncertainty). On the other hand, the numbers in row 704 represent the average accuracy values of the predictions generated by the uncertainty-enhanced perception model. As can be seen from Figure 7 it, including uncertainty estimation in the model training process can improve the performance of various tasks of the perception model. For example, after considering uncertainty, the average accuracy of lane line prediction can be improved from 51.1 to 52.6. The most significant improvement is in the prediction of crosswalks, where the average accuracy is increased from 15.7 to 18.2.
[0050] Figure 8 FIG. shows an exemplary block diagram of an uncertainty-enhanced perception system for autonomous driving according to an embodiment of the present application. The uncertainty-enhanced perception system 800 may include a plurality of sensors 802, a feature extraction unit 804, a deep learning neural network 806, an uncertainty embedding unit 808, a loss calculation unit 810, a sample selection unit 812, a model training unit 814, and a model execution unit 816.
[0051] The sensors 802 may include various types of sensors installed on the autonomous vehicle for collecting traffic data. In one embodiment, the sensors 802 may include image sensors (e.g., cameras), radars, lidars, Global Positioning System (GPS) sensors, Inertial Measurement Unit (IMU) modules, sound sensors, etc.
[0052] The feature extraction unit 804 may be responsible for extracting features from the raw sensor data. In one example, the sensor data may include multi-camera images of the environment around the autonomous vehicle, and the feature extraction unit 804 may extract BEV features from these images.
[0053] The deep learning neural network 806 may predict the environment around the autonomous vehicle based on the sensor data. The deep learning neural network 806 may include a plurality of prediction heads for performing different tasks, and each prediction head may include a plurality of branches for performing different subtasks. In some embodiments, the prediction heads in the deep learning neural network 806 may include: a classification branch for predicting the category of an object in an image; a regression branch for predicting the position / pose of each object; and a confidence branch for predicting the confidence scores of the classification and regression predictions.
[0054] The uncertainty embedding unit 808 may be responsible for embedding uncertainty estimates (e.g., confidence scores) into the perception model. In some embodiments, the uncertainty embedding unit 808 may generate uncertainty-weighted predictions based on the model predictions, ground truth values, and confidence scores associated with the predictions for each training iteration. More specifically, the uncertainty-weighted prediction for a particular task or subtask may be calculated as a linear combination of the model output and the ground truth, where the confidence score serves as the weight factor for the model output. For example, the uncertainty-weighted classification prediction may be generated according to while the uncertainty-weighted regression prediction may be generated according to .
[0055] The loss calculation unit 810 may be responsible for calculating the loss for each iteration. In some embodiments, when calculating the loss, the loss calculation unit 810 does not use the original model predictions (e.g., and ), but instead uses the uncertainty-weighted predictions (e.g., and ). Additionally, the loss may be increased by a regularization term, referred to as the regularization loss. In some embodiments, the regularization loss may be calculated according to . In another embodiment, the hyperparameter may be dynamically adjusted based on the upper bound value of the regularization loss. More specifically, when , may be adjusted downward; while when , may be adjusted upward.
[0056] The sample selection unit 812 may be responsible for selecting a subset of samples from each batch of training samples for embedding uncertainty. In some embodiments, confidence scores are predicted for only a portion (e.g., ) of the sample batch, while the remaining samples are assigned a static confidence score of 100%. In one example, approximately 1 / 4 of a batch of samples may be selected for embedding uncertainty, meaning that approximately 3 / 4 of the samples are assigned a static confidence score, such as 100%. In other embodiments, the sample selection unit 812 may also randomly select a subset and assign it a static, relatively high confidence score (e.g., 90%).
[0057] The model training unit 814 may be responsible for training the perception model using the labeled samples. In some embodiments, the model training unit 814 may use the loss calculated by the loss calculation unit 810 to optimize the model parameters. In another embodiment, gradient descent and backpropagation methods may be employed for model optimization.
[0058] The model execution unit 816 can be responsible for executing the trained uncertainty-enhanced perception model to generate a perception output. For example, new data (such as an image) acquired by the sensor 802 can be sent to the model execution unit 816 that can provide the new data as input to the trained perception model. The trained perception model can output a prediction and a confidence score for the prediction. If the model is used to detect objects in an image, the trained model can output bounding boxes, each associated with a confidence score; if the model is used for image segmentation, the trained model can output a segmented image, each segmented part associated with a confidence score. For an autonomous driving application, the output of the model (including the prediction and the corresponding confidence score) can be sent to downstream path planning or control tasks.
[0059] Since the uncertainty-enhanced perception model incorporates uncertainty estimation from the training phase, it can perceive the surrounding environment in an additional dimension (i.e., confidence), enabling downstream tasks to understand which areas are clearly visible and which are not. For example, if the confidence score of a prediction result is low, indicating poor visibility, the downstream control task may reduce the vehicle speed.
[0060] Figure 9 A flowchart is presented, showing an exemplary training process of an uncertainty-enhanced machine learning model according to an embodiment of the present application. In one or more embodiments, Figure 9 one or more of the steps in can be repeated and / or performed in a different order. Therefore, Figure 9 the specific arrangement of the steps shown in should not be construed as limiting the scope of the technology.
[0061] During operation, multiple annotated training samples can be acquired (operation 902). In one example, the machine learning model can include a perception model for autonomous driving, and the annotated training samples can include BEV images and corresponding ground truth vectorized maps, bounding boxes of detected objects, segmented images, paths, etc. The BEV images can be captured by multiple cameras installed at different positions on the autonomous vehicle. The ground truth can be provided by manually annotating the captured images.
[0062] Useful or relevant features can be extracted from the training samples (operation 904). In some embodiments, the machine learning model can include a BEV-based perception model, and BEV features can be extracted from BEV images captured by multiple cameras installed on the autonomous vehicle.
[0063] The extracted features can be sent as input to a machine learning model (operation 906). In some embodiments, the machine learning model can include a deep learning neural network with one or more prediction heads to perform one or more tasks. In one embodiment, the machine learning model can be used for autonomous driving applications, and the tasks that the prediction heads can perform include, but are not limited to, map vectorization, object detection, semantic segmentation, and path prediction.
[0064] In another embodiment, each prediction head can also include multiple branches for performing multiple subtasks. The multiple branches of the prediction head can operate in parallel to generate prediction outputs (operation 908). More specifically, the prediction outputs can include normal predictions (e.g., the outputs when the model does not consider the uncertainty in its predictions) and uncertainty predictions representing the confidence of the model in its predictions. In one example, a particular prediction head of the machine learning model can output classification and regression predictions, as well as confidence score predictions related to the classification and regression predictions. The confidence score can be between 0 and 1 and is negatively correlated with the uncertainty level of the classification and regression predictions.
[0065] The model training process can include the step of combining the normal predictions with the confidence scores to generate uncertainty-weighted predictions (operation 910). In some embodiments, the uncertainty-weighted predictions (e.g., classification or regression predictions) can be a linear combination of the normal predictions and the ground truth, where the confidence score C is the weight coefficient of the normal prediction and 1 - C is the weight coefficient of the ground truth. A higher confidence score means the prediction has a greater weight, while a lower confidence score means the ground truth has a greater weight.
[0066] The model training process can include the step of calculating a loss based on the uncertainty-weighted predictions (operation 912). In some embodiments, a regularization term can also be employed to increase the loss to prevent the model from being too biased towards low confidence. The regularization loss can be calculated as the negative logarithm of the confidence score multiplied by an adjustable hyperparameter . In another embodiment, the hyperparameter can be adjusted according to an uncertainty upper bound that limits the maximum value of the regularization loss. When , the hyperparameter can be adjusted downward; when , the hyperparameter can be adjusted upward.
[0067] The model training process can further include calculating the gradient of the loss (operation 914) and determining whether the loss has been minimized based on the gradient (operation 916). In some embodiments, the backpropagation algorithm can be used to calculate the gradient of the loss. If the loss has been minimized, the training process can terminate. Otherwise, the model parameters can be updated (operation 918), and further predictions can be made using the updated model (operation 908).
[0068] In Figure 9 the example shown, uncertainty estimation is applied to all training samples. In practice, uncertainty estimation during model training may only be applied to a subset of the training samples (e.g., a portion of the samples in each batch). In one example, a subset of the training samples can be randomly selected, and a static 100% confidence score can be assigned to the predictions made based on the selected training samples.
[0069] Figure 10 FIG. 1000 shows an example computer system 1000 for an uncertainty enhanced perception system according to an embodiment of the present application. The computer system 1000 includes a processor 1002, a memory 1004, and a storage device 1006. In addition, the computer system 1000 can be connected to a peripheral input / output (I / O) user device 1010, such as a display device 1012, a keyboard 1014, a pointing device 1016, and a camera 1018. The storage device 1006 can store an operating system 1020, an uncertainty enhanced perception system 1022, and data 1040. In some embodiments, the computer system 1000 can be implemented as part of an advanced driver-assistance system (ADAS) or an automated driving system (ADS) installed in a vehicle.
[0070] The uncertainty enhanced perception system 1022 contains a series of instructions that, when executed on the computer system 1000, cause the computer system 1000 or the processor 1002 to perform the methods and / or processes described in the present disclosure. Specifically, the uncertainty enhanced perception system 1022 can include: instructions for receiving training data (data receiving instructions 1024), instructions for extracting useful features from the training data (feature extraction instructions 1026), instructions for implementing a machine learning-based perception model (model implementation instructions 1028), instructions for generating uncertainty weighted predictions (uncertainty weighted prediction generation instructions 1030), instructions for determining a loss function based on the uncertainty weighted predictions (loss determination instructions 1032), instructions for selecting a subset of samples to apply confidence scores (sample selection instructions 1034), instructions for training the machine learning-based perception model (model training instructions 1036), and instructions for executing the trained model (model execution instructions 1038). The data 1040 can include labeled training samples 1042.
[0071] Generally speaking, the present disclosure proposes a solution to address the issue of handling uncertainty in prediction in an autonomous driving perception system. More specifically, the present disclosure can provide a direct and general method to implement uncertainty estimation in an autonomous driving perception model, which is applicable to various autonomous driving perception tasks, such as map vectorization, 3D object detection, path prediction, semantic segmentation, motion prediction, speed estimation, etc. The perception model can embed uncertainty estimation during the model training phase, where each prediction head can, in addition to making normal classification and regression predictions, also predict a confidence score representing the level of uncertainty associated with each prediction. The confidence score can be used to calculate uncertainty-weighted predictions for generating the loss. A regularization loss can be introduced to penalize low-confidence predictions. The regularization loss can also be limited by an upper bound on the uncertainty.
[0072] The proposed solution addresses key deficiencies in current uncertainty estimation in the field of autonomous driving perception, enhancing the overall safety, reliability, and performance of the autonomous driving system. Although the BEV perception framework is used as an example throughout the disclosure, the proposed solution makes very few assumptions about each individual perception task and requires only minimal modifications to existing neural network architectures in autonomous driving. Additionally, in addition to the perception model, the solution can be applied to other types of machine learning models to improve their prediction accuracy.
[0073] The data structures and program codes described in this detailed description are typically stored on a non-transitory computer-readable storage medium, which can be any device or medium capable of storing code and / or data for use by a computer system. Non-transitory computer-readable storage media include, but are not limited to, volatile memory; non-volatile memory; electrical, magnetic, and optical storage devices, solid-state drives, and / or other non-transitory computer-readable media known now or developed in the future.
[0074] The methods and processes described in the detailed description can be embodied as code and / or data, which can be stored in the above-mentioned non-transitory computer-readable storage medium. When a processor or computer system reads and executes the code stored on this medium and operates on the data stored on this medium, the processor or computer system will execute the methods and processes embodied in the form of code and data structures and stored on this medium.
[0075] In addition, the optimized parameters obtained from the above methods and processes can be programmed into hardware modules, including but not limited to application-specific integrated circuit (ASIC) chips, field-programmable gate arrays (FPGAs), and other programmable logic devices known now or developed in the future. When such a hardware module is activated, it executes the methods and processes contained within the module.
[0076] The above embodiments are for illustrative and descriptive purposes only and are not intended to be exhaustive or to limit the scope of the disclosure to the forms disclosed. Thus, many modifications and variations will be apparent to those skilled in the art. The scope of the disclosure is defined by the appended claims rather than the foregoing disclosure.
Claims
1. A method for training a perception model to perform an autonomous driving task, characterized in that, The method includes: obtaining training data containing annotations of images captured by a plurality of cameras installed at different positions on a vehicle; a perception model generating, based on the annotated training data, a task-related prediction output and a confidence score in parallel, the confidence score characterizing a level of uncertainty associated with the prediction output; generating an uncertainty-weighted prediction based on the ground truth indicated by the annotated training data, the prediction output, and the confidence score; calculating a loss function based on the uncertainty-weighted prediction; and updating the perception model based on the loss function.
2. The method according to claim 1, wherein The autonomous driving task includes: a map vectorization task; an object detection task; a semantic segmentation task; or a path prediction task.
3. The method according to claim 1, wherein The perception model includes a bird's-eye view (BEV)-based perception model.
4. The method according to claim 1, wherein The prediction output includes one or more of a classification prediction and a regression prediction.
5. The method according to claim 1, characterized in that Generating the uncertainty-weighted prediction includes calculating a linear combination of the ground truth and the prediction output weighted according to the confidence score.
6. The method according to claim 1, wherein Calculating the loss function further includes adding a regularization loss term, and wherein the regularization loss term is determined according to the confidence score and a hyperparameter.
7. The method according to claim 6, characterized in that, The hyperparameter is dynamically adjusted according to an upper limit value of the regularization loss term.
8. The method according to claim 7, wherein It further includes: responding to the regularization loss term being greater than or equal to the upper limit value by decreasing the hyperparameter; and responding to the regularization loss term being less than the upper limit value by increasing the hyperparameter.
9. The method according to claim 1, characterized in that It further includes: selecting a subset of the annotated training data; and associating a prediction output generated based on the selected subset of the annotated training data with a static confidence score.
10. A non-transitory computer-readable storage medium stores instructions that, when executed by a computer, cause the computer to perform a method for training a perception model to perform an autonomous driving task, characterized in that, The method includes: obtaining training data containing annotations of images captured by a plurality of cameras installed at different positions on a vehicle; a perception model generating, based on the annotated training data, a task-related prediction output and a confidence score in parallel, the confidence score characterizing a level of uncertainty associated with the prediction output; generating an uncertainty-weighted prediction based on the ground truth indicated by the annotated training data, the prediction output, and the confidence score; calculating a loss function based on the uncertainty-weighted prediction; and updating the perception model based on the loss function.
11. The non-transitory computer-readable storage medium according to claim 10, wherein The perception model includes a bird's-eye view (BEV)-based perception model.
12. The non-transitory computer-readable storage medium according to claim 10, wherein The prediction output includes a classification prediction and / or a regression prediction.
13. The non-transitory computer-readable storage medium according to claim 10, wherein Generating the uncertainty-weighted prediction includes calculating a linear combination of the ground truth and the prediction output weighted according to the confidence score.
14. The non-transitory computer-readable storage medium according to claim 10, wherein Calculating the loss function further includes adding a regularization loss term, and wherein the regularization loss term is determined according to the confidence score and a hyperparameter.
15. The non-transitory computer-readable storage medium according to claim 14, wherein The hyperparameter is dynamically adjusted according to an upper limit value of the regularization loss term, and the method further includes: responding to the regularization loss term being greater than or equal to the upper limit value by decreasing the hyperparameter; and responding to the regularization loss term being less than the upper limit value by increasing the hyperparameter.
16. The non-transitory computer-readable storage medium according to claim 10, wherein The method further includes: selecting a subset of the annotated training data; and associating a prediction output generated based on the selected subset of the annotated training data with a static confidence score.
17. A computing system, characterized in that, It includes: a processor; and A memory coupled to the processor, the memory storing instructions that, when executed by the processor, cause the processor to perform a method for training a perception model to perform an autonomous driving task, the method comprising: Obtaining training data containing annotations of images captured by a plurality of cameras installed at different positions on a vehicle; The perception model generates, in parallel based on the annotated training data, a prediction output related to the task and a confidence score, the confidence score characterizing the level of uncertainty associated with the prediction output; Generating an uncertainty-weighted prediction based on the ground truth indicated by the annotated training data, the prediction output, and the confidence score; Calculating a loss function based on the uncertainty-weighted prediction; and Updating the perception model based on the loss function.
18. The computing system according to claim 17, wherein Generating the uncertainty-weighted prediction includes calculating a linear combination of the ground truth and the prediction output weighted according to the confidence score.
19. The computing system according to claim 17, wherein Calculating the loss function further includes adding a regularization loss term, and wherein the regularization loss term is determined according to the confidence score and a hyperparameter.
20. The computing system according to claim 19, wherein The hyperparameter is dynamically adjusted according to an upper limit value of the regularization loss term, and the method further includes: In response to the regularization loss term being greater than or equal to the upper limit value, decreasing the hyperparameter; and In response to the regularization loss term being less than the upper limit value, increasing the hyperparameter.