Predicted uncertainty correction method and correction device

By acquiring additional features and adjusting the number of variations in the prediction model using the correction scaling factor, the problem of trajectory prediction uncertainty in the autonomous driving system is solved, the accuracy of trajectory prediction and the effectiveness of motion planning are improved, and driving safety is ensured.

CN120340244APending Publication Date: 2025-07-18HON HAI PRECISION INDUSTRY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510079115.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-18
Filing Date
2025-01-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The uncertainty of trajectory prediction in autonomous driving systems leads to insufficient planning safety, which may affect driving safety and even lead to accidents.

Method used

By obtaining additional features, the scaling factor is corrected using the decoder output and the number of variations in the trained predicted model is adjusted to correct the range of variations predicted in the line.

Benefits of technology

It improves the accuracy of trajectory prediction and the efficiency of motion planning, and improves the safety and reliability of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340244A_ABST
    Figure CN120340244A_ABST
Patent Text Reader

Abstract

The invention provides a prediction uncertainty correction method and a prediction uncertainty correction device. Obtaining one or more additional features from the data set; outputting a corrected scaling factor by inputting the additional feature to the decoder; inputting to-be-evaluated data to the trained prediction model, and outputting online prediction; and adjusting the variation number of the online prediction through the correction scaling factor, wherein the variation number corresponds to the variation range of the online prediction. Therefore, the prediction accuracy and the correction efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a calibration technique, and more particularly to a method and apparatus for predicting uncertain calibration. Background Art

[0002] An autonomous driving system relies on accurate trajectory prediction to achieve safe and efficient motion planning. The trajectory prediction function of the autonomous driving system can predict the future positions of surrounding vehicles. The planning function of the autonomous driving system uses these predicted outputs (i.e., future positions) to derive a collision-free motion path. The mutual dependence of the above two functions has caused urgent concern that inaccurate prediction may affect planning safety and even lead to serious accidents. Due to data noise and incomplete observations, inherent uncertainties still exist. Therefore, resolving the uncertainties between these two functions is crucial for ensuring driving safety. Summary of the Invention

[0003] The present invention provides a method and apparatus for predicting uncertain calibration, which can improve the prediction accuracy in related fields.

[0004] The method for predicting uncertain calibration according to an embodiment of the present invention is implemented by a processor. The method for predicting uncertain calibration includes (but is not limited to) the following steps: obtaining one or more additional features from a data set; outputting a calibration scaling factor by inputting the additional features into a decoder; outputting an online prediction by inputting data to be evaluated into a trained prediction model; and adjusting the variance of the online prediction by the calibration scaling factor, where the variance corresponds to the range of variation of the online prediction.

[0005] The calibration apparatus according to an embodiment of the present invention includes (but is not limited to) a memory and a processor. The memory is used to store program codes. The processor is coupled to the memory. The processor is configured to load the program codes to execute: obtaining one or more additional features from a data set; outputting a calibration scaling factor by inputting the additional features into a decoder; outputting an online prediction by inputting data to be evaluated into a prediction model; and adjusting the variance of the online prediction by the calibration scaling factor, where the variance corresponds to the range of variation of the online prediction.

[0006] Based on the above, the method and apparatus for predicting uncertain calibration according to an embodiment of the present invention can obtain additional features from a data set, generate a calibration scaling factor according to the additional features, and use the calibration scaling factor to reduce or increase the variance of another prediction result of a prediction model. Thereby, the accuracy of trajectory prediction can be improved, and the performance of motion planning can be improved.

[0007] To make the above features and advantages of the present invention more obvious and understandable, the following specific embodiments are given and described in detail in conjunction with the accompanying drawings. Brief Description of the Drawings

[0008] Figure 1 is a block diagram of elements of a calibration device according to an embodiment of the present invention;

[0009] Figure 2 is a flowchart of regularization training according to an embodiment of the present invention;

[0010] Figure 3 is a schematic diagram of regularization training according to an embodiment of the present invention;

[0011] Figure 4 is a flowchart of a method for predicting and uncertain calibration according to an embodiment of the present invention;

[0012] Figure 5A is a schematic diagram of prediction and post-training according to an embodiment of the present invention;

[0013] Figure 5B is a schematic diagram of post-training according to an embodiment of the present invention;

[0014] Figure 6 is a flowchart of spatio-temporal feature extraction according to an embodiment of the present invention;

[0015] Figure 7 is a schematic diagram of calibration according to an embodiment of the present invention;

[0016] Figures 8A to 8C is a schematic diagram illustrating performance verification according to an embodiment of the present invention;

[0017] Figure 9 is a schematic diagram illustrating performance verification of multiple basic components according to an embodiment of the present invention;

[0018] Figure 10 is a graph showing the relationship between sample size and expected calibration error (ECE) according to an embodiment of the present invention.

[0019] Description of Reference Numerals

[0020] 100: Calibration device;

[0021] 110: Memory;

[0022] 120: Processor;

[0023] S210~S220, S410~S440, S610~S630: Steps;

[0024] 300: Regularization training;

[0025] 310: Sensing data;

[0026] 320: Regularizer;

[0027] 330: Prediction;

[0028] 340: Motion Planning;

[0029] 331: Initial Prediction;

[0030] 332: Online Prediction;

[0031] 333: Prediction for Post - training;

[0032] 500: Post - training;

[0033] 510: (Post - training) Dataset;

[0034] 511: Observation Image;

[0035] 520: Motion Feature Extractor;

[0036] 530: Social Feature Extractor;

[0037] 540: Spatiotemporal Feature Extractor;

[0038] 550: Combination;

[0039] 560: Decoder;

[0040] 570: Calibration Scaling Factor;

[0041] 700: Motion Planning;

[0042] 111: Temperature Scaling;

[0043] 112: Overall Temperature Scaling;

[0044] 113: Isotonic Regression;

[0045] 114: Embodiment of the present invention. Detailed implementation manners

[0046] Figure 1 is a block diagram of the components of the calibration device 100 according to an embodiment of the present invention. Please refer to Figure 1 , the calibration device 100 includes (but is not limited to) a memory 110 and a processor 120. The calibration device 100 can be a mobile phone, a tablet computer, a laptop computer, a desktop computer, a server, a voice assistant device, a smart home appliance, a wearable device, a vehicle system, or other electronic devices.

[0047] The memory 110 can be any type of fixed or removable random access memory (RAM), read only memory (ROM), flash memory, traditional hard disk drive (HDD), solid-state drive (SSD), or similar components. In one embodiment, the memory 110 is used to store program code, software modules, configuration settings, data (e.g., parameters of a model, data sets, samples, features, predictions, factors, or variances), or files, which will be described in detail in subsequent embodiments.

[0048] The processor 120 is coupled to the memory 110. The processor 120 can be a central processing unit (CPU), a graphics processing unit (GPU), or other programmable general-purpose or special-purpose microprocessors, digital signal processors (DSPs), programmable controllers, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), neural processing units (NPUs), tensor processing units (TPUs), artificial intelligence (AI) accelerators, neural engines, or other similar components or combinations of the above components. In one embodiment, the processor 120 is used to execute all or part of the operations of the calibration device 100, and can load and execute each program code, software module, file, and data stored in the memory 110.

[0049] In the following, the method described in the embodiments of the present invention will be described in conjunction with various devices, components, and modules in the calibration device 100. Each process of this method can be adjusted according to the implementation situation and is not limited thereto.

[0050] Figure 2 is a flowchart of regularization training according to an embodiment of the present invention. Please refer to Figure 2, during the training phase, the processor 120 outputs an initial prediction by inputting a training set into an untrained prediction model (step S210). Specifically, the training set includes one or more training samples. Depending on different application scenarios, the training samples can be images, videos, sounds, sensing intensities, distances, angles, amplitudes, positions, trajectories, quantities, or other forms / types / morphologies.

[0051] In one application scenario, the prediction model is used for predicting the trajectory of a target object. The target object can be any type of moving vehicle (e.g., a car, a motorcycle, or a truck), a person, or other animals. The target object can be adjacent to the main object. The definition of adjacent is related to the distance between the target object and the main object and can be adjusted according to actual needs. For example, a target object within the sensing range of a sensor. The main object can be any type of moving vehicle (e.g., a car, a motorcycle, or a truck), a person, or other animals. In one application scenario, an autonomous vehicle or an in-vehicle system uses the prediction model to predict the trajectory of an adjacent target object at a future time point. The trajectory includes or is related to the position, its sequence, and the turning direction at one or more time points. In some application scenarios, the trajectory can be recorded as one or more timestamps and their corresponding positions (e.g., coordinates in latitude and longitude, geographical, or other coordinate systems). The future time point is a time point after the reference time point (e.g., the current time point). For example, the reference time point is 12:12:11, and the future time point is 12:12:13.

[0052] In one embodiment, the training sample can include a historical trajectory and a future trajectory (as the ground truth corresponding to the historical trajectory). The historical trajectory includes or is related to the position, its sequence, and the turning direction at one or more past time points. The past time point is a time point before the reference time point (e.g., the current time point). For example, the reference time point is 12:12:11, and the past time point is 12:12:10. The future trajectory includes or is related to the position, its sequence, and the turning direction at one or more future time points. The definition and examples of the future time point are as described above and will not be elaborated here.

[0053] In one embodiment, the training sample can include sensing data and a future trajectory (as the ground truth corresponding to the sensing data). The main object can be equipped with or carry sensors and obtain sensing data related to images, sounds, distances, directions, etc. The sensing data can be used to generate the positional relationship between the main object and the target object. For example, the relative distance or direction. The sensing data can be used to generate the motion information of the main object and / or the target object. For example, speed, acceleration, or attitude.

[0054] In one embodiment, the training samples may further include environmental information. The main object may be equipped or carry sensors, and accordingly obtain sensor data related to, for example, images, sounds, distances, temperatures, and humidities (which may be used as training samples). The environmental information is, for example, the above-mentioned sensor data.

[0055] Figure 3 It is a schematic diagram of regularization training according to an embodiment of the present invention. Please refer to Figure 3 , the processor 120 can train a prediction model through a machine learning algorithm. The machine learning algorithm can analyze the labeled training samples (for example, historical trajectories and / or sensor data with corresponding true labels) to establish the association between the historical trajectories and / or sensor data (that is, the input of the model, the perception data 310 shown in the figure) and the future trajectory (that is, the output of the model, the prediction 330 shown in the figure). The prediction model can be learned and can be used to infer the data to be evaluated (for example, the perception data 310 to be evaluated) to output the prediction 330 corresponding to the perception data 310.

[0056] The type of the machine learning algorithm can be changed according to the application scenario. Taking trajectory prediction as an example, the machine learning algorithm can be Encoder-Decoder Long Short-Term Memory (ED-LSTM), Hierarchical Vector Transformer (HiVT), Transformer in Transformer (TNT), Learning LaneGraph Representations for Motion Forecasting (LaneGCN), or AutoBots, but not limited thereto. In other application scenarios, the machine learning algorithm can be Multiple Layer Perception (MLP), Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM) network, or Temporal Convolutional Network (TCN) (for example, Conv-TasNet), but not limited to this.

[0057] In one embodiment, the prediction 330 includes an initial prediction 331. The prediction 330 obtained by the prediction model inferring one or more training samples (for example, historical trajectories and / or sensor data) in the input training set is the initial prediction 331 (for example, the future trajectory).

[0058] In one embodiment, during the training phase of the prediction model, the parameters of the prediction model are recursively updated by minimizing a loss function (which is related to the error / loss between the output of the prediction model (i.e., the initial prediction 331) and the true label in the training sample (e.g., the future trajectory)). The parameters of the model are, for example, weights, the number of layers, the position or number of neurons, activation functions, or offsets, and are not limited thereto. The method of updating the parameters is, for example, by gradient descent, Adaptive Moment Estimation (Adam) optimizer, Momentum method, Adagrad, or conjugate gradient method, and is not limited thereto. That is, one of the multiple objectives of the training phase is to make the initial prediction output by the prediction model close to or the same as the corresponding true label.

[0059] Please refer to Figure 2 , the processor 120 updates the parameters of the prediction model according to the total loss function (step S220). Specifically, one of the multiple objectives of the training phase is to minimize the total loss function. In one embodiment, the total loss function is the sum of the Negative Log-Likelihood (NLL) and the calibration loss function. For example, the total loss function L CCTR The mathematical expression of is:

[0060] L CCTR = L NNL + L CAL …(1)

[0061] , L NNL is the negative log-likelihood, and L CAL is the calibration loss function.

[0062] The mathematical expression of the negative log-likelihood L_NNL is:

[0063]

[0064] is the distribution P under the initial prediction i (e.g., the i-th future trajectory) of the true label y (e.g., the uncertainty corresponding to the i-th training sample, or variance) made by the prediction model with variance (e.g., the i-th initial prediction). For example, the training set includes N training samples By inputting the training sample (x i, y i , h i ) to the initial prediction made by the prediction model accompanied by the variance (or corresponding to the initial prediction of the uncertainty) under the distribution P, x i is the i-th historical trajectory, and h i is the i-th environmental information. The distribution P is a predefined distribution, such as a Gaussian distribution, but not limited thereto. However, the training samples are not limited to historical trajectories and environmental information and can be changed according to application requirements. And the variance corresponds to the range of change of the initial prediction . Taking trajectory prediction as an example, the initial prediction is the future trajectory corresponding to the i-th historical trajectory and environmental information. Assuming that the future trajectory is the position at a certain future time, then the variance is the range / area of change of this position. That is, this position may occur within the specific range / area defined by the variance . For example, the prediction model predicts that the position at this future time point is within this range / area of change. The larger the variance , the larger the range / area of change, and the farther the maximum change in position is from the original predicted position; the smaller the variance , the smaller the range / area of change, and the closer the maximum change in position is to the original predicted position.

[0065] The calibration loss function L CAL is represented by the following mathematical formula:

[0066]

[0067] . The calibration loss function L CAL regularizes the difference between the first value and the second value. The first value (i.e., ) is the square of the difference between the initial prediction and the corresponding true label y i (for example, obtained by subtracting the true label y i and the initial prediction ), and the second value is the variance of the initial prediction The variance should match the actual difference between the initial prediction and the true label and the true label . For example,

[0068] As Figure 3 shown, in the regularization training 300, the regularizer 320 additionally assigns the calibration loss function L to the difference between the initial prediction 331 and the true labelCAL In one embodiment, the prediction 330 includes an online prediction 332. The processor 120 can output the online prediction 332 by inputting the data to be evaluated into the trained prediction model. The prediction 330 deduced by the trained prediction model from the input data to be evaluated (e.g., historical trajectories and / or sensed data) is the online prediction 332 (e.g., future trajectories, positions at future time points, or other deduced results). In one embodiment, a trained prediction model means that its (total) loss function has converged, the prediction accuracy has reached the corresponding threshold value, or the training has reached the standard for early stopping of training, but the training completion standard can still be adjusted according to other tasks or requirements.

[0069] In one embodiment, the online prediction 332 can be used to determine the motion planning data of the main object (i.e., Figure 3 the motion planning 340 shown). The motion planning data includes the motion parameters of the main object at one or more future time points. The motion parameters are, for example, speed, acceleration, direction, angular velocity, attitude, or a combination thereof. Algorithms related to motion planning are, for example, the Rapidly-exploring Random Tree (RRT) or Neural Message Passing (NMP) in full name, and are not limited thereto. These algorithms can calculate a series of actions or control instructions to guide a robot or a mobile vehicle (i.e., the main object) from the starting state to the target state safely and efficiently while meeting one or more constraints (e.g., avoiding obstacles, complying with traffic rules, etc.). The processor 120 can generate control instructions for the main object (e.g., a vehicle or a vehicle-mounted system) according to the motion parameters. The control instructions are, for example, used for accelerating, braking, steering, or reversing.

[0070] Figure 4 is a flowchart of an uncertain correction method for prediction according to an embodiment of the present invention. Refer to Figure 4 , the processor 120 obtains one or more additional features from the data set (step S410). Specifically, the additional features are features generated based on the data set. Depending on different application scenarios, the types of additional features may be different. The additional features may be related to motion, interaction, relative relationship, time, and / or space. The data set can be a validation set or other types of data sets. The data set includes one or more input samples. The introduction of the input samples can refer to the description of the training samples above and will not be repeated here. For example, the data set includes M input samples i.e., x j is the j-th historical trajectory, y j is the j-th true label, and h j is the j-th environmental information.

[0071] In one embodiment, the processor 120 outputs a post-training prediction by inputting a data set into a trained prediction model. Specifically, Figure 5A is a schematic diagram of prediction and post-training according to an embodiment of the present invention. Please refer to Figure 5A , Figure 3 The prediction 330 may include a post-training prediction 333. The post-training prediction 333 is used for the post-training 500 of the trained prediction model. For example, it is used to update the parameters (e.g., weights, number of layers, position or number of neurons, activation function, or bias) of the network or model used in the post-training 500. The post-training 500 is a series of techniques and strategies adopted after the prediction model has completed (initial) training (e.g., Figure 3 the regularization training 300), in order to further improve its performance or adapt to new data. In some application scenarios, the post-training 500 can be fine-tuned or optimized based on the trained prediction model. The prediction 330 inferred by the trained prediction model from the input data set (e.g., historical trajectories and / or sensed data) is the post-training prediction 333 (e.g., future trajectories).

[0072] Figure 5B is a schematic diagram of the post-training 500 according to an embodiment of the present invention. Please refer to Figure 5B , the positions at multiple past time points in the (post-processed) data set 500 (i.e., past trajectories) can be mapped to a geographical coordinate system or a map, and an observation image 511 can be generated accordingly. The (post-processed) data set 500 may include one or more observation images 511.

[0073] In one embodiment, the (post-processed) data set 500 includes the positions of one or more target objects at one or more past time points, and one or more additional features include motion features. The motion features can be speed and / or acceleration. The processor 120 can execute the motion feature extractor 520 and determine the motion features by comparing the positions corresponding to multiple past time points. For example, the distance is obtained by comparing the positions at two different future time points and is used to calculate the speed. Or, the speeds at two different future time points are used to calculate the acceleration.

[0074] In one embodiment, the (post-processed) data set 500 includes the positions of one or more target objects, the additional features include social features, and the social features are the distance between the target objects and the number of target objects. The processor 120 can execute the social feature extractor 530 and determine the social features from the input samples in the (post-processed) data set 500. For example, calculate the distance between the main object and the target objects in the observation image 511 corresponding to a certain past time point and the number of target objects.

[0075] In one embodiment, the (post-processed) dataset 500 includes the positions of one or more target objects at one or more past time points, and the additional features include spatio-temporal features. The processor 120 may execute the spatio-temporal feature extractor 540 and obtain the spatio-temporal features. Figure 6 is a flowchart of spatio-temporal feature extraction according to an embodiment of the present invention. Refer to Figure 6 , the processor 120 may generate a plurality of top views (step S610) that record the positions of one or more target objects at multiple past time points based on the map information. Specifically, the map information may be information about coordinates, routes, directions, road segments, buildings, or other objects in a Geographic Information System (GIS). The processor 120 may generate a map image of a certain area (e.g., covering multiple positions in the trajectory) according to the map information. The map image may be a view obtained by orthographically projecting downward from above the object. That is, a top view, a bird's-eye view, or an upper view. The map image may include a road area. That is, the image area of the road. Then, the processor 120 may map the positions of the target objects at multiple past time points or the past trajectory in the (post-processed) dataset 500 to the map image. For example, mark patterns or texts at the corresponding positions in the map image according to the latitude and longitude coordinates of one or more positions in the past trajectory. The mapped top view may be used as the observation image 511. In some embodiments, the mapping of the positions also refers to the environmental information.

[0076] The processor 120 may output spatial representations corresponding to multiple past time points by inputting one or more top views to the first machine learning network (step S620). Specifically, the first machine learning network is a network trained by a machine learning algorithm. The machine learning algorithm is, for example, a Convolutional Neural Network (CNN), AlexNet related to the convolutional neural network, VGGNet (Very Deep Convolutional Networks), ResNet (Residual Neural Network), or Inception. In the training phase of the first machine learning network, the parameters of the first machine learning network are recursively updated by minimizing a loss function (the error / loss related to the output of the first machine learning network and the true label in the input). The parameters of the first machine learning network are, for example, weights, the number of layers, the position or number of neurons, activation functions, or biases, and are not limited thereto. The method for updating the parameters is, for example, by gradient descent, Adaptive Moment Estimation (Adam) optimizer, momentum method, Adagrad, or conjugate gradient method, but is not limited thereto. The spatial representation is, for example, the distance and / or direction in space between the main object and the target object, and / or the position of the lane line in space.

[0077] The processor 120 may output spatio-temporal features by inputting the spatial representations corresponding to multiple past time points to the second machine learning network (step S630). Specifically, the second machine learning network is a network trained by a machine learning algorithm. The machine learning algorithm is, for example, a Gated Recurrent Unit (GRU), Long Short-Term Memory (LSTM), or Transformer. In the training phase of the second machine learning network, the parameters of the second machine learning network are recursively updated by minimizing a loss function (the error / loss related to the output of the second machine learning network and the true label in the input). The parameters of the second machine learning network are, for example, weights, the number of layers, the position or number of neurons, activation functions, or biases, and are not limited thereto. The method for updating the parameters is, for example, by gradient descent, Adaptive Moment Estimation (Adam) optimizer, momentum method, Adagrad, or conjugate gradient method, but is not limited thereto. The spatio-temporal features are, for example, the distance and / or direction in space between the main object and the target object at multiple (past) time points and their correlation, and the movement information of the lane line.

[0078] Please refer to Figure 5B, the processor 12 can combine the 550 motion features, social features, and spatio-temporal features. The combination of the motion features, social features, and spatio-temporal features is used as input to the decoder 560. The combination method is, for example, concatenate or arranged according to rules.

[0079] It should be noted that for other application fields, the types, contents, and acquisition methods of additional features may be different.

[0080] Please refer to Figure 4 and Figure 5B , the processor 120 outputs a calibration scaling factor by inputting additional features to the decoder (step S420). Specifically, the decoder is a network with a linear layer accompanied by an activation function. The activation function is, for example, softplus, softmax, swish, and is not limited thereto. One or more weight coefficients are assigned to the linear layer. The additional features are respectively operated with these weight coefficients, and the operation output passes through the activation function, that is, a calibration scaling factor is generated. The calibration scaling factor is a parameter used to shrink or magnify the variance and will be detailed in subsequent embodiments.

[0081] In the training stage of the encoder, the parameters of the encoder are recursively updated by minimizing the loss function (the error / loss related to the output of the encoder and the true label corresponding to the additional features). The parameters of the encoder are, for example, weights, number of layers, positions or numbers of neurons, activation functions, or biases, and are not limited thereto. The method for updating the parameters is, for example, through the gradient descent method, Adaptive Moment Estimation (Adam) optimizer, momentum method, Adagrad, or conjugate gradient method, but is not limited thereto.

[0082] In one embodiment, the processor 120 can update the feature acquirer (for example, Figure 5B the motion feature acquirer 520, social feature acquirer 530, and / or spatio-temporal feature acquirer 540) and / or the parameters of the decoder 560 used to obtain additional features according to the dataset, post-training prediction 333, and one or more additional features. For example, the loss function is defined by negative log-likelihood (NLL): where μ i is the position or future trajectory of the future time point in the post-training prediction 333, and γ i is the combination of the i-th additional feature or multiple additional features. The parameters of the feature acquirer (for example, Figure 5B the motion feature acquirer 520, social feature acquirer 530, and / or spatio-temporal feature acquirer 540) and / or the decoder 560 are recursively updated by minimizing the loss function using an update algorithm (for example, Adaptive Moment Estimation (Adam) optimizer, gradient descent method, or Adagrad).

[0083] Please refer to Figure 4 , the processor 120 outputs an online prediction 332 by inputting the information to be evaluated into the trained prediction model (step S430). Specifically, the generation of the online prediction 332 can refer to the foregoing description and will not be elaborated herein.

[0084] The processor 120 adjusts the variance of the online prediction 332 by correcting the scaling factor (step S440). Specifically, the variance corresponds to the change range of the online prediction 332. The reference to the variance can refer to the foregoing description and will not be elaborated herein. In one embodiment, in response to the correction scaling factor being greater than 1, the processor 120 amplifies the variance of the online prediction 332. For example, increasing the change range of the position. In another embodiment, in response to the correction scaling factor being less than 1, the processor 120 reduces the variance of the online prediction 332. For example, reducing the change range of the position.

[0085] Figure 7 is a schematic diagram of correction according to an embodiment of the present invention. Please refer to Figure 7 , the correction scaling factor generated by the post-training 500 can be used to correct the variance corresponding to the online prediction 332. In one embodiment, the processor 120 can determine the motion planning data of the main object according to the online prediction 332 with the adjusted variance (i.e., perform motion planning 700). The motion planning data includes the motion parameters of the main object at one or more future time points. The introduction of the motion planning data and the motion parameters can refer to the foregoing description and will not be elaborated herein. The online prediction 332 with the adjusted variance means that the variance of this online prediction 332 has been amplified or reduced by the correction scaling factor. Since the motion planning 700 may involve collision avoidance, a more reasonable driving path or motion planning parameters can be generated by correcting the variance.

[0086] Figures 8A to 8C is a schematic diagram for explaining the performance verification according to an embodiment of the present invention. Please refer to Figures 8A to 8C , the embodiments of the present invention can be applied to various machine learning models related to trajectory prediction. For example, ED-LSTM, HiVT, TNT, LaneGCN, and AutoBots. The performance verification is to perform simulations using the above models respectively. The baselines include:

[0087] Temperature Scaling (TS): Using the global temperature to scale the variance.

[0088] Isotonic Regression (IR): Training an auxiliary model based on isotonic regression.

[0089] Ensemble Temperature Scaling (ETS): Learn a mixture of uncalibrated, temperature-scaled calibrated, and uniform probability outputs.

[0090] Calibration metrics include: Expected Calibration Error (ECE), Mean Calibration Error (MCE), and Noise Calibration Error (NCE). The three calibration metrics of the embodiments of the present invention are all the lowest. The L2 error is for Rapidly-exploring Random Trees (RRT) and Neural Message Passing (NMP), and the errors of the embodiments of the present invention are all the lowest. For Average Displacement Error (ADE) and Final Displacement Error (FDE), the errors of the embodiments of the present invention are all the lowest. It can be seen that the embodiments of the present invention can improve the prediction accuracy of trajectory prediction and motion planning.

[0091] Figure 9 is a schematic diagram showing the performance verification of multiple basic components according to an embodiment of the present invention. Please refer to Figure 9 , the basic components may include Figure 3 regularizer 320 of Figure 5B motion feature extractor 520, social feature extractor 530, spatio-temporal feature extractor 540, and post-processing 500. The calibration metric for executing the calibration procedures of all basic components is the lowest. Although the calibration metric is higher when any one of the basic components is missing, it is still within an acceptable range.

[0092] Figure 10 is a graph showing the relationship between the sample size and the Expected Calibration Error (ECE) according to an embodiment of the present invention. Please refer to Figure 10 , the horizontal axis is the size of the calibration dataset. As the size of the calibration dataset increases, the expected calibration errors of temperature scaling 111, ensemble temperature scaling 112, and isotonic regression 113 are significantly higher than the expected calibration error of the embodiment 114 of the present invention. It can be seen that the embodiments of the present invention can provide higher calibration performance.

[0093] In summary, in the prediction uncertainty correction method and correction device according to the embodiments of the present invention, an additional feature extracted from a self-dataset is used to generate a correction scaling factor, and the correction scaling factor is used to adjust the variance corresponding to the prediction of a trained prediction model. In addition, the embodiments of the present invention provide a corresponding loss function for the variance, so that the variance more closely matches the error between the prediction and the true label. Thereby, the prediction efficiency and the correction data efficiency can be improved, and a better correction baseline can be obtained.

[0094] Although the present invention has been disclosed as above by way of embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the scope defined by the appended claims.

Claims

1. A method for correcting prediction uncertainty, implemented by a processor, the method for correcting prediction uncertainty comprising: Obtaining at least one additional feature from a data set; Outputting a correction scaling factor by inputting the at least one additional feature into a decoder; Outputting an online prediction by inputting data to be evaluated into a trained prediction model; And Adjusting the variance of the online prediction by the correction scaling factor, wherein the variance corresponds to the range of change of the online prediction.

2. The method for correcting prediction uncertainty according to claim 1, wherein the data set includes the positions of at least one target object at past time points, the at least one additional feature includes a motion feature, and the step of obtaining the at least one additional feature from the data set: Determining the motion feature by comparing the positions corresponding to a plurality of the past time points.

3. The method for correcting prediction uncertainty according to claim 1, wherein the data set includes the positions of at least one target object, the at least one additional feature includes a social feature, the social feature is the distance from the at least one target object and the number of the at least one target object, and the step of obtaining the at least one additional feature from the data set: Determining the social feature from the input samples in the data set.

4. The method for correcting prediction uncertainty according to claim 1, wherein the data set includes the positions of at least one target object at past time points, the at least one additional feature includes a spatio-temporal feature, and the step of obtaining the at least one additional feature from the data set: Generating a plurality of top views recording the positions of the at least one target object at a plurality of the past time points based on map information; Outputting a spatial representation corresponding to the past time point by inputting the top view into a first machine learning network; And Outputting the spatio-temporal feature by inputting the spatial representation corresponding to the past time point into a second machine learning network.

5. The method for correcting prediction uncertainty according to claim 4, wherein the first machine learning network is a convolutional neural network.

6. The method for correcting prediction uncertainty according to claim 4, wherein the second machine learning network is a gated recurrent unit.

7. The method for correcting prediction uncertainty according to claim 1, wherein the data set includes the positions of at least one target object at past time points, the at least one additional feature includes a motion feature, a social feature and a spatio-temporal feature, and the step of obtaining the at least one additional feature from the data set: Combining the motion feature, the social feature and the spatio-temporal feature, wherein the combination of the motion feature, the social feature and the spatio-temporal feature is used for input into the decoder.

8. The method for correcting prediction uncertainty according to claim 1, further comprising: Outputting a post-training prediction by inputting the data set into the trained prediction model; And Updating the parameters of the feature acquirer and / or the decoder used to obtain the at least one additional feature according to the data set, the post-training prediction and the at least one additional feature.

9. The method for correcting prediction uncertainty according to claim 1, further comprising: During the training phase, an initial prediction is output by inputting a training set into the untrained prediction model; and updating the parameters of the prediction model according to a total loss function, where the total loss function is the sum of a negative log-likelihood and a calibration loss function, the calibration loss function regularizes the difference between a first value and a second value, the first value is the square of the difference between the initial prediction and the corresponding true label, and the second value is the variance of the initial prediction.

10. The method for calibrating prediction uncertainty according to claim 1, wherein the online prediction includes the position of at least one target object at a future time point, the at least one target object is adjacent to a main object, and the method for calibrating prediction uncertainty further includes: Determining motion planning data of the main object according to the online prediction with the adjusted variance, where the motion planning data includes motion parameters of the main object at another future time point.

11. A calibration device, comprising: a memory for storing program code; and a processor coupled to the memory and configured to load the program code to execute: extracting at least one additional feature from a data set; outputting a calibration scaling factor by inputting the at least one additional feature into a decoder; outputting an online prediction by inputting data to be evaluated into a trained prediction model; and adjusting the variance of the online prediction by the calibration scaling factor, where the variance corresponds to the range of change of the online prediction.

12. The calibration device according to claim 11, wherein the data set includes the position of at least one target object at a past time point, the at least one additional feature includes a motion feature, and the processor is further configured to: determine the motion feature by comparing the positions corresponding to multiple past time points.

13. The calibration device according to claim 11, wherein the data set includes the position of at least one target object, the at least one additional feature includes a social feature, the social feature is the distance from the at least one target object, and the number of the at least one target object, and the processor is further configured to: determine the social feature from input samples in the data set.

14. The calibration device according to claim 11, wherein the data set includes the position of at least one target object at a past time point, the at least one additional feature includes a spatio-temporal feature, and the processor is further configured to: generate a plurality of top views recording the positions of the at least one target object at multiple past time points based on map information; outputting a spatial representation corresponding to the past time point by inputting the top view into a first machine learning network; and outputting the spatio-temporal feature by inputting the spatial representation corresponding to the past time point into a second machine learning network.

15. The calibration device according to claim 14, wherein the first machine learning network is a convolutional neural network.

16. The calibration device according to claim 14, wherein the second machine learning network is a gated recurrent unit.

17. The calibration device according to claim 11, wherein the data set includes the positions of at least one target object at past time points, the at least one additional feature includes motion features, social features and spatio-temporal features, and the processor is further configured to: Combine the motion features, the social features and the spatio-temporal features, wherein the combination of the motion features, the social features and the spatio-temporal features is used as an input to the decoder.

18. The calibration device according to claim 11, wherein the processor is further configured to: Output a post-training prediction by inputting the data set to the trained prediction model; and Update the parameters of the feature extractor and / or the decoder used to extract the at least one additional feature according to the data set, the post-training prediction and the at least one additional feature.

19. The calibration device according to claim 11, wherein the processor is further configured to: Output an initial prediction by inputting a training set to the untrained prediction model during a training phase; and Update the parameters of the prediction model according to a total loss function, wherein the total loss function is the sum of a negative log-likelihood and a calibration loss function, and the calibration loss function regularizes the difference between a first value and a second value, the first value being the square of the difference between the initial prediction and the corresponding true label, and the second value being the variance of the initial prediction.

20. The calibration device according to claim 11, wherein the online prediction includes the positions of at least one target object at future time points, the at least one target object is adjacent to the main object, and the processor is further configured to: Determine motion planning data of the main object according to the online prediction with an adjusted variance, wherein the motion planning data includes motion parameters of the main object at another future time point.