Post-fusion-mode-oriented uncertainty quantification method under vehicle-road cooperation perception scene
By performing uncertainty modeling and supervising learning on a single agent in the vehicle-road collaboration perception scenario, the uncertainty of vehicle-side and road-side perception results is quantified, and the problem of unknown credibility of the fusion result in vehicle-way collaboration perception is solved, and the reliability and safety of the system are improved.
Patent Information
- Application Number
- CN202510066398.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-27
AI Technical Summary
In vehicle-road collaboration perception, the credibility of the fusion results is often unknown, and existing methods are difficult to effectively evaluate and manage the credibility of different information sources, resulting in possible misjudgment and decision-making errors.
A method of uncertainty quantification of backward fusion method in vehicle-road collaboration perception scenarios is proposed. By modeling uncertainty on a single agent, supervised learning is performed in combination with real classification regression labels, classification and regression uncertainty of vehicle-side and road-side perception results is calculated, and post-fusion uncertainty after collaboration-perception is calculated based on matching results.
The reliability and security of perceived results in the collaborative perception process are quantified, the overall performance of the system is improved, and the risk of misjudgment and decision-making errors is reduced.
Smart Images

Figure CN120047728A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to an uncertainty quantification method and a computer storage medium for a backward fusion method in a cooperative perception scenario. Background Art
[0002] In recent years, the rapid development of autonomous driving technology has attracted extensive attention and research globally. Driven by intelligent transportation systems, autonomous driving is not only regarded as an important way of future travel, but also a key means to achieve road safety, traffic efficiency, and environmental protection. With the continuous progress of sensor technology, artificial intelligence algorithms, and communication networks, the perception, decision-making, and control capabilities of autonomous vehicles are increasingly enhanced. However, in the face of complex and changing traffic environments, relying solely on the vehicle's own perception ability often makes it difficult to comprehensively and accurately understand the surrounding situation. Therefore, vehicle-road cooperative perception, as an important way to improve the performance of autonomous driving, has gradually become a research hotspot.
[0003] Vehicle-road cooperative perception refers to the real-time information sharing and interaction between vehicles and road infrastructure (such as traffic lights, surveillance cameras, etc.), other vehicles, and cloud data, thereby forming a more comprehensive and dynamic environmental perception system. This cooperative mode can not only provide rich context information but also improve the early warning ability for potential dangers, effectively enhancing driving safety and comfort. For example, through real-time interaction with traffic lights, autonomous vehicles can better predict signal changes, optimize driving routes, and thus reduce traffic congestion and accident risks.
[0004] However, vehicle-road cooperative perception also faces a series of challenges in practical applications. First, the credibility of the fusion result is often unknown. In the process of fusing data from different sources, the reliability of each information source may vary due to various factors (such as sensor failures, data delays, and environmental interference, etc.). How to effectively evaluate and manage the credibility of these information sources has become the key to improving the overall performance of the system. Second, modeling the credibility of the fusion result is also complex. Existing methods often cannot fully capture the trust relationship between different information sources, resulting in possible misjudgments and decision-making errors during the information fusion process. Therefore, to address the above problems, an uncertainty quantification method for a backward fusion method in a vehicle-road cooperative perception scenario is needed to quantify the reliability and safety of the perception results during the cooperation process. Summary of the Invention
[0005] Aiming at the deficiencies in the prior art, the purpose of the present invention is to provide an uncertainty quantification method for a backward fusion method in a vehicle-road cooperative perception scenario.
[0006] To solve the above problems, the technical solution of the present invention is as follows:
[0007] An uncertainty quantification method for the backward fusion method in a vehicle-road cooperation perception scenario, comprising the following steps:
[0008] Perform uncertainty modeling on a single agent;
[0009] Combine real classification regression labels for supervised learning;
[0010] Calculate the classification uncertainty and regression uncertainty of the perception results at the vehicle end and the road end;
[0011] Match the perception results at the vehicle end and the road end; and
[0012] Calculate the fused uncertainty after cooperative perception according to the matching result.
[0013] Preferably, an agent is defined as a system that can perceive its environment and act based on these perceptions. The agent can change its strategy through reasoning and learning to adapt to changes in the environment. Specifically, the object of this patent is mainly a vehicle-mounted perception system or a roadside perception system equipped with data perception devices (such as lidar, cameras, etc.), a deep learning computing platform, and a deployed visual perception model.
[0014] Preferably, uncertainty quantification mainly focuses on the quantification of random uncertainty: there are two types of uncertainties faced by the agent during the perception process: epistemic uncertainty and random uncertainty. Epistemic uncertainty stems from insufficient information or knowledge and can be mitigated or resolved by obtaining more information; while random uncertainty stems from the randomness of the real world, such as the prediction of future possibilities. Random uncertainty cannot be completely eliminated, but can be coped with by evaluating the probabilities of different outcomes. For autonomous driving, understanding and effectively modeling these two types of uncertainties is crucial for safe decision-making, because imperfect perception data needs to be dealt with and future scenarios need to be accurately predicted during the driving process.
[0015] Preferably, the detection result of the agent in this patent can be a 2D or 3D detection result, and this patent mainly focuses on the more complex 3D detection.
[0016] 2D detection results generally refer to object detection performed on a two-dimensional plane (i.e., the image plane), such as detecting the position, bounding box, and classification label of an object in an image. In contrast, 3D detection results refer to the detection of target objects in three-dimensional space, including information such as object classification, position, and shape. For example, based on point cloud data, features such as the spatial position, size, and orientation of an object can be detected. Although this method is mainly applied to 3D detection scenarios, it is also applicable to 2D detection scenarios. The detection results in this study include object categories and 3D position and shape parameters, specifically including: x, y, z (the three-dimensional coordinates of the object's center position), l, w, h (the length, width, and height of the object), and θ (the rotation angle of the object), which together describe the geometric information of the 3D detection box.
[0017] Preferably, this method mainly aims at quantifying the uncertainty of cooperative perception in the post-fusion manner:
[0018] Vehicle-road cooperative perception mainly includes three methods, namely pre-fusion, mid-fusion, and post-fusion (abbreviated as post-fusion in this patent). Pre-fusion fuses raw sensor data such as raw point clouds; mid-fusion fuses the feature representations encoded by neural networks; post-fusion fuses the detection results from the vehicle side and the road side. The structure of this method mainly includes single-agent 3D detection and uncertainty modeling, and post-fusion of vehicle-road cooperative perception results.
[0019] The method model structure proposed in this patent mainly includes two parts: single-agent 3D detection and uncertainty modeling, and post-fusion of vehicle-road cooperative perception results.
[0020] Optionally, the model for the uncertainty of a single agent:
[0021] In the traditional PointPillars model, the main tasks are classification and bounding box regression, that is, predicting the category of an object and the bounding box of its location and size. However, the PointPillars model does not consider the uncertainty of detection results. To solve this problem, two additional regression terms are added to the detection head to predict the classification variance and (bounding box) regression variance, so as to provide uncertainty estimation for classification and regression tasks. See Figure 3 。
[0022] a) Classification variance:
[0023] In the original classification head (for predicting the category of an object), a new detection head is added to predict
[0024] the classification variance. The structure of this regression term is exactly the same as that of the original classification detection head, so its implementation method is the same as that of the traditional classification task, except that its goal becomes to estimate the variance of the category, indicating the uncertainty of classification.
[0025] b) Variance regression:
[0026] For bounding box regression (including the prediction of the center position, length, width, height, and orientation of the bounding box), a new detection head is added to predict the regression variance. Assuming that each 3D bounding box regression variable (including position, size, and orientation) is independent and each regression variable follows a univariate Gaussian distribution, through the regression
[0027] variance, the random uncertainty of each regression variable (i.e., the mentioned x, y, z, l, w, h, θ) can be quantified.
[0028] c) Modification of the model architecture.
[0029] In the implementation process, these two additional regression terms (classification variance and regression covariance) are both implemented through a
[0030] convolutional head. The convolutional head enables the model to predict these uncertainty values at each position while maintaining the simplicity of the original structure. The classification and regression variance heads are exactly the same as the original convolutional head (in terms of convolutional kernel size, stride, etc.) and do not require re-design, except that their functions are to estimate variance and covariance respectively.
[0031] Optionally, the principle that this method can learn uncertainty without real uncertainty labels is:
[0032] The uncertainty quantification of this method is based on loss attenuation. This method is mainly achieved by adding additional regression terms and modifying the loss function. This method can implicitly learn the variance in the loss function with real classification and regression labels, and this variance is the uncertainty. For the specific detailed derivation, please refer to the paper.
[0033] Optionally, the input-output encoding method described in this method is:
[0034] The input sample x is the original sensor data (such as LiDAR point cloud data), which is the basic information for object detection. The 3D proposal z is generated by the LiDAR network and is used to initially estimate the category and position of the object.
[0035] This proposal includes:
[0036] c z : The category label of the object. For example, the category labels for person, car, and bicycle are 0, 1, and 2 respectively.
[0037] s z : The softmax score of the category, indicating the confidence of the category prediction.
[0038] b z: The proposed 3D position, which is a 7D vector, represents the position and size of the target in 3D space. Specifically, it includes:
[0039] d x ,d y : The position offset on the horizontal plane (usually referring to the x and y coordinate offsets in a 2D plane).
[0040] d z : The height offset at the bottom of the target (i.e., the displacement along the z-axis).
[0041] log(l), log(w), log(h): Represent the logarithmic scales of the length, width, and height of the target respectively, aiming to better handle scale changes.
[0042] θ: The rotation angle of the target
[0043] The task of the fusion network is to further optimize and predict the final category and position of the target based on the 3D proposal z generated by LiDAR, combined with other information (such as context or prior information). Its output is y, which contains:
[0044] c y : The final target category label (e.g., 0, 1, 2), similar to c z and is used to indicate what kind of target it is.
[0045] s y : The softmax score of the target category, representing the confidence of the fusion network in the category prediction.
[0046] b y : The predicted 3D bounding box position, which is a 7D vector representing the position offset relative to the original 3D proposal b z That is to say, b y represents the correction or optimization of b z .
[0047] Optionally, the probability modeling process of this method is as follows:
[0048] Assume there are pre-trained LiDAR and camera networks, which can generate 3D proposals z and RoI (Region of Interest) features for fusion. From the perspective of maximum likelihood estimation, a set of network weights w is learned to maximize the observed likelihood of the training data. The negative log-likelihood is minimized by setting the following loss function:
[0049]
[0050] In the context of classification problems, p(y∣x,z) usually refers to the multinomial probability mass function, which is widely known as the cross-entropy loss. The Focal loss further adapts to the problem of positive and negative sample imbalance by introducing a modulating factor.
[0051] For deterministic regression problems, assuming that p(y∣x,z) is a Gaussian density function with a fixed variance, the corresponding loss function is the L2 loss. The calculation formula is:
[0052]
[0053] In this work, the proposed distribution p(z∣x) is introduced into the loss function, and the formula is as follows:
[0054]
[0055] where x, y, and z represent the input samples generated from the original sensor data (such as LiDAR point cloud data), the final output result of the model, and the 3D proposals generated by the LiDAR network (for preliminary estimation of the category and location of the target), respectively.
[0056] Since there is no analytical solution, this loss function is approximately calculated by sampling:
[0057] Sample z′~p(z∣x),
[0058]
[0059] Optionally, the process of designing the loss function for supervised learning by combining the true classification labels in this method is as follows:
[0060] Assume that the distribution of softmax logits is a Gaussian distribution, that is: where, represents the mean, corresponding to the standard output of the network (the predicted value of the logits). represents the variance, representing the classification noise, which is regressed by adding an additional output layer at the head of the LiDAR network. However, directly learning this variance variable is difficult. For this reason, the reparameterization trick is adopted to sample the logits, and the specific process is as follows:
[0061] a) Sample a standard normal distribution variable
[0062] b) Generate logits using the formula:
[0063]
[0064] The generated logits z are then converted to softmax scores, and then the Focal loss is calculated (see Figure 3 ).
[0065] Optionally, the process of designing the loss function for supervised learning by combining with the true regression label in this method is as follows:
[0066] For simplicity, this patent uses scalar notation instead of vector notation to introduce the proposed method. For example, b z represents a single regression variable in the regression vector b z . Assume that each regression variable follows a Gaussian distribution, that is where represents the mean of the distribution, corresponding to the standard output of the network (i.e., the regression prediction value); represents the regression variance.
[0067]
[0068] The bounding box regression is explicitly supervised in the loss function, and at the same time, the corresponding uncertainty variance representation is learned. The specific derivation process and principle of the loss function are prior art.
[0069] The logits in this method refer to the output result of the linear regression obtained by mapping the input data to the output space after the deep learning model undergoes forward propagation. In a multi-classification problem, the logits represent the scores of each class; in a binary classification problem, the logits represent the scores of a sample belonging to the positive class or the negative class.
[0070] Optionally, the covariance of the regression uncertainty obtained based on the regression variance of the model output is calculated as follows:
[0071] a) Variance prediction to Cholesky decomposition
[0072] Convert the diagonal variance output by the regression network into a Cholesky decomposition matrix. If only diagonal elements are included (representing the variance of each sample for this regression variable ), then calculate the standard deviation of the diagonal variance and embed it on the diagonal of the Cholesky decomposition matrix. Finally, a Cholesky decomposition matrix L of (N, 7×7) is obtained.
[0073] b) Construct a multivariate normal distribution and sample;
[0074] Through Cholesky decomposition of matrix L and mean prediction values Construct a multivariate normal distribution for generating sampled data. In the actual operation process, a multivariate normal distribution can be created based on MultivariateNormal using the PyTorch framework. Through the reparameterization trick (ReparameterizationTrick, whose purpose is to
[0075] introduce randomness into the network's calculation process while maintaining the differentiability of the sampling process with respect to the gradient. Through this trick, samples can be drawn from the standard normal distribution and transformed into samples that conform to the target distribution
[0076] This. The actual code implementation can use.resample() to sample 100 groups of data and generate sampling results of (N, k, S).
[0077] N is the number of samples, K is the dimension (e.g., 7, representing the 7 parameters of the 3D regression box mentioned above); S is the number of sampling times (e.g., 100).
[0078] c) Decode the sampling results into the coordinate space;
[0079] Decode the box regression parameters (offsets relative to the anchor points) obtained by sampling and transform them into the coordinate space. The shape of the decoded sampling box is (N, k, S). For the specific implementation, reference can be made to PointPillars, and this patent will not provide a detailed explanation here.
[0080]
[0081] d) Calculate the mean and covariance of the regression (final prediction);
[0082] The calculation process is as follows:
[0083]
[0084] where x i represents the nth sample after sampling. For the input point cloud or image, y and Σ are respectively the predicted mean and covariance matrix, that is, the predicted 3D detection box and its corresponding uncertainty representation. S takes 100.
[0085] Optionally, the step of matching the perception results of the vehicle end and the road end in this method specifically includes:
[0086] Through the Hungarian algorithm, match the detection results of the vehicle end with the detection results of the road end, and correspond the detection results of the vehicle end and the road end to the same object. The detection results of the vehicle end and the road end before matching are respectively and The detection results of the vehicle end and the road end after matching are respectively and Similarly, the covariance of the vehicle-end detection results and the road-end detection regression before matching are ∑ i and ∑ j respectively. After matching, the vehicle-end detection results and the road-end detection results are ∑ i′ and ∑ j′ .
[0087] (That is Figure 1 , and ∑ i represent the 3D bounding box prediction results of the vehicle end and the corresponding regression covariance results; and ∑ j represent the 3D bounding box prediction results of the road end and the corresponding regression covariance results)
[0088] Optionally, according to the matching result of the Hungarian algorithm in this method, the calculation method for calculating the difference between the Gaussian distributions of the vehicle-end and road-end detections is as follows:
[0089] Calculate the difference between two Gaussian distributions based on the Mahalanobis distance, and the calculation formula is:
[0090]
[0091] where δ is the collaborative perception post-fusion uncertainty proposed by this method.
[0092] Optionally, the uncertainty quantification of this method is also applicable to other 3D detection models, such as Second, VoxelNet, PV-RCNN, etc.
[0093] Optionally, during the regression loss design process of this method, in order to make the model converge stably and avoid potential 0-division problems, instead of directly predicting it predicts The loss function can be expressed as:
[0094]
[0095] Optionally, the uncertainty distance obtained by this method is further processed by min-max normalization.
[0096] Optionally, the logits in this method refer to the output results of the linear regression obtained by mapping the input data to the output space after the deep learning model has undergone forward propagation. In multi-classification problems, logits represent the scores for each class; in binary classification problems, logits represent the scores for a sample belonging to the positive or negative class.
[0097] Compared with the prior art, the advantages of the present invention are:
[0098] The present invention quantifies the uncertainty of the cooperative perception scenario through a direct modeling-based method, without the need for large-scale modification of the model. Only an additional detection head needs to be added to the current model structure. At the same time, the present invention can operate in real time, and the obtained uncertainty results have a certain scene generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] Other features, objects, and advantages of the present invention will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:
[0100] Figure 1 Schematic diagram of the vehicle-road cooperative perception uncertainty quantification model structure provided by the embodiment of the present invention;
[0101] Figure 2 Flow chart of the uncertainty quantification method for the backward fusion method in the vehicle-road cooperative perception scenario provided by the embodiment of the present invention;
[0102] Figure 3 Network structure diagram of the detection head. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0103] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0104] Figure 1 Schematic diagram of the vehicle-road cooperative perception scenario provided by the embodiment of the present invention, which shows the scenario of obtaining environmental information through point cloud and image. In this embodiment, the vehicle cooperates with road infrastructure (such as millimeter-wave radar, surveillance camera, etc.) to jointly achieve the acquisition and fusion of perception information. Through the cooperative perception between vehicles, the dynamic information of the traffic environment can be obtained more comprehensively, providing higher safety protection for autonomous driving.
[0105] Figure 2 Flow chart of the uncertainty quantification method for the backward fusion method in the cooperative perception scenario provided by the embodiment of the present invention. As Figure 2 shown, the uncertainty quantification method for the backward fusion method in the cooperative perception scenario of the present invention includes the following steps:
[0106] S1: Perform uncertainty modeling on a single agent;
[0107] Two new regression terms are added to the detection head to predict the variance of classification and the covariance of bounding box regression, so as to provide uncertainty estimation for classification and regression tasks. See Figure 3 .
[0108] a) Classification variance regression:
[0109] In the original classification head (used to predict the category of an object), a new regression head is added to predict
[0110] the classification variance. The structure of this regression term is exactly the same as that of the original classification detection head. Therefore, its implementation method is consistent with traditional classification tasks, except that its goal becomes to estimate the variance of the category, indicating the uncertainty of classification.
[0111] b) Bounding box covariance regression:
[0112] For bounding box regression (including the prediction of the center position, length, width, height, and orientation of the bounding box), a new regression head is added to regress the covariance. Assuming that each 3D bounding box regression variable (including position, size, and orientation) is independent and each regression variable follows a univariate Gaussian distribution, by regressing the variance, the random uncertainty of each regression variable (x, y, z, l, w, h, θ) can be quantified.
[0113] c) Modification of the model architecture.
[0114] In the implementation process, these two additional regression terms (classification variance and regression covariance) are both implemented through convolutional heads. The convolutional heads enable the model to predict these uncertainty values at each position while maintaining the simplicity of the original structure. The classification and regression variance heads are exactly the same as the original convolutional heads and do not require re - design. Only a simple modification is made to predict the variance; for the regression covariance regression head, the new convolutional head outputs the variance estimate of each regression variable.
[0115] The input - output encoding method is as follows:
[0116] The input sample x is the original sensor data (such as LiDAR point cloud data), which is the basic information for object detection. The 3D proposal z is generated by the LiDAR network and is used to initially estimate the category and position of the object.
[0117] This proposal includes:
[0118] c z : The category label of the object. For example, for a person, a car, and a bicycle, they are 0, 1, and 2 respectively.
[0119] s z : The softmax score of the category, indicating the confidence of the category prediction.
[0120] b z : The 3D position of the proposal, which is a 7 - dimensional vector representing the position and size of the object in three - dimensional space. Specifically, it includes:
[0121] d x ,d y : The position offset on the horizontal plane (usually refers to the x and y coordinate offsets in the two-dimensional plane).
[0122] d z : The height offset of the target bottom (i.e., the displacement along the z-axis).
[0123] log(l), log(w), log(h): Represent the logarithmic scales of the length, width, and height of the target respectively, aiming to better handle scale changes.
[0124] θ: The rotation angle of the target
[0125] The task of the fusion network is to further optimize and predict the final category and position of the target based on the 3D proposal z generated by LiDAR, combined with other information (such as context or prior information). Its output is y, including:
[0126] c y : The final target category label (e.g., 0, 1, 2), similar to c z and is used to indicate what category the target belongs to.
[0127] s y : The softmax score of the target category, indicating the confidence of the fusion network's prediction of the category.
[0128] b y : The predicted 3D bounding box position, which is a 7-dimensional vector representing the position offset relative to the original 3D proposal b z That is to say, b y represents the correction or optimization of b z .
[0129] S2: Conduct supervised learning by combining the true classification regression labels;
[0130] Assume there are pre-trained LiDAR and camera networks, which can generate 3D Proposal z and RoI (Region of Interest, i.e., the region of interest) features for fusion. From the perspective of maximum likelihood estimation, a set of network weights w is learned to maximize the observed likelihood of the training data. The negative log-likelihood is minimized by setting the following loss function:
[0131]
[0132] In the context of classification problems, p(y∣x,z) usually refers to the multinomial probability mass function, It is widely known as cross - entropy loss. Focal loss further adapts to the problem of positive and negative sample imbalance by introducing a modulation factor.
[0133] For the deterministic regression problem, assume that \(p(y|x,z)\) is a Gaussian density function with a fixed variance, and the corresponding loss function is the \(L2\) loss.
[0134] In this work, the proposed distribution \(p(z|x)\) is introduced into the loss function, and the formula is as follows:
[0135]
[0136] Since there is no analytical solution, this loss function is approximately calculated by sampling:
[0137] Sample \(z'\sim p(z|x)\),
[0138]
[0139] The process of designing the loss function for supervised learning by combining the true classification labels in this method is as follows:
[0140] Assume that the distribution of softmax logits is a Gaussian distribution, that is:
[0141] where, represents the mean, corresponding to the standard output of the network (the predicted value of logits).
[0142] represents the variance, representing the classification noise, which is regressed by adding an additional output layer at the LiDAR network head.
[0143] However, directly learning this variance variable is difficult.
[0144] For this reason, the reparameterization trick is adopted to sample the logits, and the specific process is as follows:
[0145] a) Sample a standard normal distribution variable
[0146] b) Use the formula to generate logits:
[0147]
[0148] The generated logit l z Subsequently, it is converted into a softmax score, and then the Focal loss is calculated (see Figure 3)。
[0149] The process of designing the loss function for supervised learning by combining real regression labels in this method is described as follows:
[0150] For the sake of simplicity, this patent uses scalar notation instead of vector notation to introduce the proposed method.
[0151] For example, b z represents a single regression variable in the regression vector b z .
[0152] Assume that each regression variable follows a Gaussian distribution, that is
[0153] where represents the mean of the distribution, corresponding to the standard output of the network (i.e., the regression prediction value); represents the variance of the distribution (regression variance), represents regression noise, and is regarded as an auxiliary regression variable for network prediction.
[0154]
[0155] Explicitly supervise the bounding box regression in the loss function while learning the corresponding uncertainty variance representation. The specific derivation process and principle of the loss function are prior art.
[0156] S3: Input the point cloud or image, and obtain the classification and regression variances from the classification variance and regression variance detection heads respectively. The calculation process from the classification variance to the classification uncertainty representation is as follows:
[0157] a) Classification distribution construction:
[0158] Construct a classification distribution based on the normal distribution through torch.distributions.normal.Normal
[0159] Refer to Figure 3 . In the specific implementation process, use torch.clamp to constrain the lower bound of the variance to 10 -6 to avoid numerical instability.
[0160] b) Classification score sampling:
[0161] Use the rsample method to sample the constructed normal distribution to obtain the uncertainty representation of the classification score. The sampling process generates 10 samples, constituting an approximation of the classification score distribution.
[0162] c) Final representation of classification uncertainty.
[0163] Perform Sigmoid activation processing on the sampled classification scores to map the scores to the probability range
[0164] [0, 1]. Finally, perform a mean operation on the 10 activated sample scores to represent the classification uncertainty approximately, with a range of 0 to 1.
[0165] The specific calculation process of the covariance from regression variance to regression uncertainty is as follows:
[0166] a) Variance prediction to Cholesky decomposition
[0167] Convert the diagonal variance output by the regression network into a Cholesky decomposition matrix.
[0168] If it only contains diagonal elements (representing the variance of each sample for this regression variable, e.g., ), then calculate the standard deviation of the diagonal variance and embed it on the diagonal of the Cholesky decomposition matrix. Finally, obtain the Cholesky decomposition matrix L of (N, 7×7).
[0169] b) Construct a multivariate normal distribution and sample;
[0170] Construct a multivariate normal distribution through the Cholesky decomposition matrix L and the mean prediction value to generate sampled data.
[0171] In the actual operation process, a multivariate normal distribution can be created based on MultivariateNormal using the pytorch framework. Through the reparameterization trick (ReparameterizationTrick, whose purpose is to introduce randomness into the network calculation process while maintaining the differentiability of the sampling process with respect to the gradient. Through this trick,
[0172] samples can be drawn from the standard normal distribution and transformed into samples that conform to the target distribution.
[0173] The actual code implementation can sample 100 groups of data using.resample() to generate the sampling result of (N, k, S). N is the number of samples, K is the dimension (e.g., 7, representing the 7 parameters of the 3D regression box mentioned above
[0174] ); S is the number of sampling times (e.g., 100).
[0175] c) Decode the sampling result into the coordinate space;
[0176] Decode the sampled bounding box regression parameters (offsets relative to the anchor points) and convert them into the coordinate space. The shape of the decoded sampled bounding box is (N, k, S). For the specific implementation, refer to PointPillars, and this patent will not explain it in detail here.
[0177] d) Calculate the mean and covariance (final prediction);
[0178] The calculation process is as follows:
[0179]
[0180] where x n represents the nth sample after sampling. For the input point cloud or image, y and Σ are obtained as the predicted mean and covariance matrix respectively, that is, the predicted 3D detection bounding box and its corresponding uncertainty representation. S takes 100.
[0181] S4: Match the perception results of the vehicle end and the road end;
[0182] The step of matching the perception results of the vehicle end and the road end in this method specifically includes:
[0183] Through the Hungarian algorithm, match the detection results of the vehicle end and the road end, and correspond the detection results of the vehicle end and the road end to the same object. The detection results of the vehicle end and the road end before matching are and The detection results of the vehicle end and the road end after matching are and Similarly, the regression covariances of the vehicle end detection results and the road end detection results before matching are ∑ i and ∑ j , and the detection results of the vehicle end and the road end after matching are ∑ i′ and ∑ j′ . The specific algorithm process of the Hungarian algorithm matching is not specifically explained in this method.
[0184] (that is Figure 1 , and ∑ i represent the 3D bounding box prediction results of the vehicle end and the corresponding regression covariance results; and ∑ j represent the 3D bounding box prediction results of the road end and the corresponding regression covariance results).
[0185] S5: Calculate the fused uncertainty after collaborative perception according to the matching results.
[0186] According to the matching results of the Hungarian algorithm, the calculation method for the difference between the Gaussian distributions of the vehicle end and road end detections is:
[0187] The difference between two Gaussian distributions is calculated based on the Mahalanobis distance, and the calculation formula is:
[0188]
[0189] where
[0190] δ is the collaborative perception post-fusion uncertainty proposed by this method.
[0191] Compared with the prior art, the present invention realizes the representation of perception reliability in abnormal environments through the quantification of vehicle-road collaborative perception uncertainty, provides method support for the deployment of safe vehicle-road collaborative algorithms, and reduces potential safety hazards caused by algorithm failures.
[0192] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A method for quantifying uncertainty in a backward fusion mode in a vehicle-road cooperative perception scenario, characterized in that: The method comprises the following steps: Modeling uncertainty for individual agents; Combine true classification regression labels for supervised learning; Calculate the classification uncertainty and regression uncertainty of the vehicle-side and road-side perception results; Matching 3D detection results between multiple agents; and Calculate the difference between the multivariate Gaussian distributions of the 3D detection frames of the matched vehicle-road detection frames.
2. The uncertainty quantification method for backward fusion in the vehicle-road cooperative perception scenario according to claim 1 is characterized in that: Uncertainty quantification for random uncertainty.
3. The uncertainty quantification method for backward fusion in the vehicle-road cooperative perception scenario according to claim 1 is characterized in that: The uncertainty model structure for a single agent includes the following steps: Two new regression heads are added to the detection head, one for predicting the variance of classification and the other for predicting the covariance of bounding box regression.
4. The uncertainty quantification method for backward fusion in the vehicle-road cooperative perception scenario according to claim 1 is characterized in that: The supervised learning is achieved by designing the loss function of the added regression head.
5. The uncertainty quantification method for backward fusion in the vehicle-road cooperative perception scenario according to claim 1 is characterized in that: The process of supervised learning is: Set the loss function to minimize the negative log-likelihood: p(y|x,z) - the multinomial probability mass function, - Cross entropy loss; For deterministic regression problems, assuming that p(y|x,z) is a Gaussian density function with a fixed variance, the corresponding loss function is L2 loss; Introducing the proposed distribution p(z|x) into the loss function, the formula is as follows: Samplez ′ ~p(z∣x), 6. The uncertainty quantification method for backward fusion in the vehicle-road cooperative perception scenario according to claim 1 is characterized in that: The loss function design process includes the following steps: Assume that the distribution of softmaxlogits is Gaussian, that is in, represents the mean, corresponding to the standard output of the network, i.e. the predicted value of logits; represents variance, which represents classification noise; The generated logitl z It is then converted into a softmax score and then the Focalloss is calculated.
7. The uncertainty quantification method for backward fusion in the vehicle-road cooperative perception scenario according to claim 1 is characterized in that: The loss function design process also includes the following steps: Assume that each regressor follows a Gaussian distribution, that is in, Represents the mean of the distribution, corresponding to the standard output of the network, i.e., the regression prediction value; represents the variance of the regression, represents the regression noise, and is considered as an auxiliary regressor for network prediction; 8. The uncertainty quantification method for backward fusion in the vehicle-road cooperative perception scenario according to claim 7 is characterized in that: Not directly predictable And prediction 9. The uncertainty quantification method for backward fusion in the vehicle-road cooperative perception scenario according to claim 1 is characterized in that: The specific calculation process of the final 3D detection box output and regression uncertainty is: where x i Represents the nth sample after sampling, inputs the point cloud or image, and obtains μ and Σ; y-predicted mean, that is, the predicted 3D detection box; Σ - Covariance matrix, which is a representation of the uncertainty in the regression.
10. The uncertainty quantification method for backward fusion in the vehicle-road cooperative perception scenario according to claim 9 is characterized in that: The specific steps of matching the 3D detection results between the multiple intelligent agents; and calculating the difference between the multivariate Gaussian distributions of the 3D detection frames of the matched vehicle-road detection frames include: The Hungarian algorithm is used to match the vehicle-side detection results with the road-side detection results, and the vehicle-side detection results and the road-side detection results correspond to the same object. The vehicle-side detection results and the road-side detection results before matching are respectively and The vehicle-side detection results and road-side detection results after matching are and According to the matching results of the Hungarian algorithm, the difference between the Gaussian distributions of vehicle-side and road-side detections is calculated based on the Mahalanobis distance.