Target detection credibility estimation method fusing data and model uncertainty
By integrating target detection confidence estimation methods that combine data and model uncertainties, and utilizing Bayesian inference and logistic regression mechanisms, a multidimensional confidence evaluation framework is constructed. This solves the uncertainty modeling problem in multimodal perception scenarios under autonomous driving environments, and improves the robustness and interpretability of the perception system.
Patent Information
- Application Number
- CN202511155766.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-11-07
AI Technical Summary
Existing object detection methods lack joint uncertainty modeling for multimodal perception scenarios in autonomous driving environments, and cannot effectively estimate the credibility of image detection and point cloud detection results, resulting in insufficient robustness of the perception system in complex environments.
By integrating data and model uncertainties in target detection confidence estimation methods, and utilizing Bayesian inference and logistic regression mechanisms, combined with image quality scores, weather conditions, target distance, point cloud quantity, and structural consistency, a multi-dimensional confidence evaluation framework is constructed to output a confidence score for each target detection result.
It significantly improves the robustness of perception results and the safety of decision-making in autonomous driving systems, enhances the interpretability and dynamic adaptability of perception systems, and enables accurate assessment of the credibility of detection results in complex environments.
Smart Images

Figure CN120913013A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and particularly relates to a target detection credibility estimation method fusing data and model uncertainty. BACKGROUND
[0002] With the rapid development of automatic driving technology, the environmental perception system as the front link of the whole vehicle intelligent decision and control, the accuracy and reliability of its perception results have a decisive role on driving safety. Among them, target detection as one of the core tasks of environmental perception, mainly through multi-modal sensors such as cameras and laser radars to identify and locate traffic participants in the road environment, and provide key inputs for path planning, trajectory prediction and risk assessment modules. However, in real complex traffic environment, sensor noise, environmental interference, algorithm model generalization ability and other factors often lead to uncertainty in target detection results, if the detection results are directly used for downstream system, it is easy to cause system-level decision errors. Therefore, accurately estimating the credibility of target detection results, that is, carrying out detection uncertainty modeling and quantitative analysis, has important significance for improving the robustness of the perception system and realizing safe and controllable automatic driving.
[0003] Current target detection uncertainty modeling methods can be mainly divided into two categories: the first category of methods analyzes uncertainty from the data-driven perspective, that is, by modeling the influence of environmental factors on detection performance, such as image blur, light change, weather factors such as rain and snow, target distance change, and occlusion degree, the detection confidence decay law under different data conditions is quantified, and the original detection results are post-processed. This method is easy to implement and has good generalization, but it usually ignores the uncertainty factors in the model internal reasoning process and cannot describe the performance fluctuation of the model on unknown samples or boundary samples. The second category of methods analyzes uncertainty from the model ontology perspective, such as introducing Bayesian neural networks, MC Dropout and other means to model the model parameters or output distribution, and then obtain the prediction variance and uncertainty score. This method can effectively estimate the inherent prediction bias of the model, but due to its high computational complexity, high deployment cost, and easy to be affected by sample feature distribution changes, its practicability is limited in multi-sensor fusion systems.
[0004] In summary, existing methods often only consider modeling uncertainties from one aspect of the data or model, lacking a systematic framework that unifies and integrates both types of uncertainty sources. This is especially true in multimodal perception scenarios, where a method for joint inference of the credibility of image detection and point cloud detection results is still lacking. To address this, this invention proposes a target detection credibility estimation method that integrates data uncertainty and model uncertainty. It systematically integrates methods for correcting 2D detection confidence based on image quality scores and weather conditions (reflecting data uncertainty), the impact of target distance and environmental factors on 3D detection performance (reflecting the coupling of data and model uncertainty), the support relationship between the number of point clouds within the detection box and the target's existence (reflecting the uncertainty caused by data sparsity), and the consistency confidence of the class size prior and the actual structure matching (reflecting the uncertainty of the model's generalization ability). This method achieves quantitative estimation of the credibility of multimodal detection results through a fusion mechanism of Bayesian inference and logistic regression, enhancing the interpretability and dynamic adaptability of target detection results, and providing a robust and reliable basis for evaluating the quality of perception results for autonomous driving systems. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides a method for estimating the credibility of target detection by integrating data and model uncertainties. This method aims to improve the credibility assessment capability of multimodal target detection results in autonomous driving systems and enhance the interpretability and robustness of perception systems. The method fully integrates environmental change factors from image and point cloud data with model output features. Through Bayesian correction and confidence fusion mechanisms, it outputs a credibility score for each target detection result.
[0006] The technical means employed in this invention are as follows: A target detection credibility estimation method that integrates data and model uncertainty includes: Acquire real-time point cloud data from the vehicle-mounted LiDAR sensor and image data from the vehicle-mounted camera; Image data is input into a quality assessment model to obtain an image quality assessment score; Image data is input into a two-dimensional object detection model, which then performs inference on the image and outputs the two-dimensional detection box of the target and its original confidence score. Image data is input into the weather perception model to obtain the weather information of the current vehicle location; By combining image quality assessment scores and weather information, the original confidence scores are corrected based on the Bayesian posterior principle, and a two-dimensional confidence score that takes into account the influence of image and weather factors is output. Point cloud data is input into a 3D target detection model, and the 3D target detection model performs inference on the point cloud data to output the 3D detection box in the current traffic scene and its original confidence score. Based on the target distance and weather conditions, the original confidence of the 3D detection box is corrected by posterior confidence to obtain the 3D confidence of environmental adaptability; The number of points contained within the detection box is counted, and based on the prior statistical relationship between the number of points and the detection accuracy, Bayesian inference is used to calculate the confidence score of the number of points. Based on the degree of consistency between the spatial geometry of the detection box and the typical prior size distribution of the target category, a structural consistency matching degree is constructed, and the structural consistency confidence degree is output. Using five confidence indices—image quality assessment score, two-dimensional confidence, three-dimensional confidence, point confidence, and structural consistency confidence—as input features, a regression estimation function based on logistic regression is constructed. Combined with true confidence labels based on true IoU, category matching, and detection uniqueness, a confidence regression model is trained, and the final uncertainty estimate of the multimodal fusion target detection result is output.
[0007] Furthermore, the step of inputting image data into a quality assessment model to obtain an image quality assessment score specifically includes: A no-reference image quality assessment model is built based on the open-source IQA-PyTorch framework. During the training phase, the LIVE publicly available image quality assessment dataset was used, and the ARNIQA neural network was employed to predict image quality. The input image is a color three-channel image with a resolution equal to the original image size. The output is a normalized image quality score. A higher score indicates higher image quality.
[0008] Furthermore, the step of inputting image data into a two-dimensional target detection model, performing inference on the image through the two-dimensional target detection model, and outputting the two-dimensional detection box of the target and its original confidence score specifically includes: YOLOv11 is used as the backbone network for image object detection, and real-time inference is performed on the images. The output format is as follows:
[0009] in, The coordinates of the center and the width and height of the two-dimensional detection frame; Category labels; This represents the original confidence level.
[0010] Furthermore, the step of inputting image data into the weather perception model to obtain the current weather information of the vehicle specifically includes: The WeatherNet image weather recognition network is used to perform multi-classification on the input image and output the weather category. .
[0011] Furthermore, the process of combining image quality assessment scores with weather information, and correcting the original confidence score based on Bayesian posterior principles, outputs a two-dimensional confidence score that considers the influence of image and weather factors. Specifically, this includes: On a pre-defined validation dataset, different image quality score ranges and different weather conditions were statistically analyzed. mAP detection accuracy under these conditions; For each combination The model accuracy was fitted to obtain a joint mapping relationship with image quality and weather conditions, and a two-dimensional confidence correction coefficient was constructed. , is represented as:
[0012] in, The detection accuracy under the current weather and image quality conditions; The detection accuracy under sunny conditions and high image quality within the dataset; correction coefficients. This indicates the extent of the decline in testing capacity; The original confidence score is corrected using a Bayesian posterior approach, expressed as:
[0013] in, Two-dimensional confidence level to account for the influence of image and weather factors.
[0014] Furthermore, the step of inputting point cloud data into a 3D object detection model, performing inference on the point cloud data through the 3D object detection model, and outputting the 3D detection bounding box and its original confidence score in the current traffic scene specifically includes: Construct a 3D target detection model based on the PointPillars architecture; The input is a LiDAR point cloud, which is processed through voxel encoding, a backbone network, and a detection head module to output a 3D detection bounding box. ,in, The coordinates of the center of the 3D detection frame are its width, length, and height. The heading angle of the 3D detection box. This represents the original three-dimensional confidence level.
[0015] Furthermore, the step of correcting the original confidence of the 3D detection box based on the target distance and weather conditions to obtain the environmental adaptability 3D confidence specifically includes: In the dataset, the accuracy metrics of the 3D detector were statistically analyzed for different distance intervals (<10m, 10-20m, 20-30m, 30-40m, ..., >100m) and under different weather conditions. The confidence level reduction pattern table was obtained; Fitting function , which represents the fidelity of the model in the weather and distance , and is calculated as:
[0016] wherein, is the detection accuracy in the weather and distance ; is the reference detection accuracy in sunny day and close distance (<10m); is the three-dimensional confidence correction coefficient; The final corrected three-dimensional confidence is:
[0017] wherein, is the three-dimensional confidence, is the original three-dimensional detection output confidence.
[0018] Further, the number of point clouds contained in the detection box is counted, and based on the prior statistical relationship between the number of point clouds and the detection accuracy, the point number confidence is calculated by Bayesian inference, specifically including: In the standard verification data set, for each detected target, the number of point clouds in the three-dimensional detection box is recorded , the point number is divided into several intervals (<5, 5-20, 20-40, 40-60, >60), and the detection accuracy mAP of the corresponding interval is counted; The mapping relationship between the number of point clouds and the detection accuracy is constructed to obtain the point number confidence correction coefficient; The interval with the most point number (>60) is taken as the ideal reference accuracy , the correction coefficient corresponding to each point number interval is calculated, which is expressed as:
[0019] wherein, is the correction coefficient corresponding to each point number interval; is the detection accuracy corresponding to the number of point clouds in the three-dimensional bounding box, is the ideal reference accuracy; According to the corresponding relationship of the interval and its correction coefficient, a point number confidence correction function is constructed based on the cubic spline interpolation method; Let the confidence of the original three-dimensional detector output be , and the point number correction coefficient be , then the point number confidence correction result is expressed as:
[0020] wherein, is the point number correction coefficient, represents the confidence level of the detection confidence under the current point number; is the original three-dimensional detection output confidence; is the point number confidence after combining the point number Bayesian correction.
[0021] Further, the point cloud number contained in the statistical detection frame is included, and based on the prior statistical relationship between the point cloud number and the detection accuracy, the point number confidence is calculated by using Bayesian inference, specifically including: In the training set or validation set, the true three-dimensional size (length, width, height) of each target is collected, and the mean and covariance matrix are calculated respectively to establish the prior distribution, as follows:
[0022] wherein, is the Gaussian distribution of the three-dimensional size of the target class, is the target class; represents the average size of the corresponding target class; is the covariance matrix, which describes the fluctuation range of the size of the class; simplify the features as independent Gaussians, represented as: , ,
[0023] wherein, is the length of the target class, is the Gaussian distribution of the length conforming to the mean and the variance, is the average value of the length of the target class, is the variance of the length of the target class in all training samples, is the width of the target class, is the Gaussian distribution of the width conforming to the mean and the variance, is the average value of the width of the target class, is the variance of the width of the target class, is the height of the target class, is the Gaussian distribution of the height conforming to the mean and the variance, is the average value of the height of the target class is the variance of the height of the target class; For the target output by the detection model, the class label is , the size is , the distance measure using Mahalanobis distance to measure the degree of deviation of three-dimensional size in the target category statistical space is calculated Confidence of the target detected by the detection model in the size distribution of the category to which it belongs, denoted as:
[0024] wherein, is the confidence of the target detected by the detection model in the size distribution of the category to which it belongs; ; The Mahalanobis distance is mapped to the confidence score based on the Gaussian likelihood form, denoted as:
[0025] Assuming that the original detection model output confidence , the structure consistency correction coefficient is defined as , the structure consistency confidence is represented as:
[0026] wherein, is the structure consistency confidence.
[0027] Further, the image quality evaluation score, two-dimensional confidence, three-dimensional confidence, point number confidence and structure consistency confidence are used as input features to construct a regression estimation function based on logistic regression, and a credibility regression model is trained by combining real credibility labels based on real IoU, category matching and detection uniqueness, and the uncertainty estimation of the final multi-modal fusion target detection result is output, specifically including: Based on the target detection dataset and the credibility label, a rule is constructed, and a credibility dataset is constructed; For each predicted box , the matching ground truth (GT) is , the IoU score is defined, and the expression is:
[0028] The category consistency score is defined, and the expression is:
[0029] If multiple boxes match the same ground truth (GT), only the one with the highest IoU is retained, and the others are considered redundant (repeated detection). The uniqueness penalty term is defined, and the expression is:
[0030] wherein, ; Based on the IoU score , category consistency score and uniqueness penalty term , calculate the final real trust label, the expression is:
[0031] wherein, is the final real trust label; Take five confidence indicators as input features, including two-dimensional detection confidence based on image quality and weather factor correction , three-dimensional detection confidence based on distance and weather condition correction , support confidence based on point cloud quantity and target existence , structure consistency confidence , image quality score , construct a logistic regression model, the expression is:
[0032] wherein, Sigmoid function; is the regression weight parameter to be learned; is the bias term; indicates the fusion confidence score output by the model; For enhanced physical interpretability, the regression coefficients are non-negative normalized to obtain the fusion weight, the expression is:
[0033] Use weighted average method to fuse into the final confidence score, the expression is:
[0034] wherein, is the fusion weight; is , , , , Five confidence indicators.
[0035] Compared with the prior art, the present application has the following advantages: 1. The target detection confidence estimation method provided by the application can significantly improve the robustness and decision safety of environmental perception results. Compared with single confidence relying on original model output, the multi-dimensional correction mechanism proposed by the application can more accurately capture the confidence fluctuation caused by complex environmental changes, avoiding false confidence and misjudgment in the case of image blur, low visibility in rainy days or long-distance targets. At the same time, the introduction of structural consistency and point cloud support makes the reliability of three-dimensional detection results further enhanced, providing reliable input for downstream tasks such as multi-target tracking and risk prediction.
[0036] 2. The target detection confidence estimation method provided by the application can be embedded in mainstream target detection and sensor fusion systems, has good universality and engineering deployment ability, and has important significance for building a "verifiable and controllable" safe automatic driving perception system, and is one of the key supporting technologies in the realization of high-level automatic driving landing process.
[0037] 3. The target detection confidence estimation method provided by the application focuses on the confidence modeling of target detection results in the automatic driving scene, and solves the problems of confidence information loss and model reliability difficulty in quantification in the current perception system in complex environment, which has clear engineering application value and practical significance.
[0038] 4. The target detection confidence estimation method provided by the application adopts image quality score, weather identification, target distance measurement, point cloud support analysis, structure prior modeling and other input information, which is easy to obtain in the existing automatic driving perception platform, and the calculation overhead of the Bayesian inference and logistic regression fusion strategy involved in the method is controllable, which has good system integration and actual deployment ability.
[0039] 5. The target detection confidence estimation method provided by the application systematically fuses multi-source information, first constructs a five-dimensional confidence index system based on image, point cloud and structure prior, and realizes confidence estimation based on Bayesian hierarchical inference and linear weighting mechanism, which provides an interpretable, traceable and controllable confidence expression form for the perception system.
[0040] 6、The application provides a target detection confidence estimation method fusing data and model uncertainty, a two-dimensional confidence correction method of image quality and weather perception guidance is proposed, which can effectively improve the reliability of the perception system under low-quality images, rain, snow and other adverse conditions; a three-dimensional detection confidence correction model based on distance and environmental factors fully considers the characteristics of spatial decay of model performance; the introduction of point cloud quantity and target category structure prior further improves the fine-grained description ability of confidence; the confidence output mechanism fusing logistic regression and Bayesian inference provides stable and quantifiable input guarantee for multi-task cooperation (such as perception-prediction-planning) of the automatic driving system, which helps to improve the safety redundancy and risk control ability of the overall system.
[0041] Based on the above reasons, the application can be widely promoted in the field of automatic driving. BRIEF DESCRIPTION OF DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0043] Figure 1 The method flowchart of the present application. DETAILED DESCRIPTION
[0044] In order to make the person skilled in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings of the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.
[0045] It should be noted that the terms "include" and "have" and any variations thereof in the specification and claims of the present application and the above drawings are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0046] As shown in Figure 1 The present application provides a target detection confidence estimation method fusing data and model uncertainty, which comprises: Obtain real-time point cloud data of a vehicle-mounted laser radar sensor and image data of a vehicle-mounted camera; Input the image data into a quality assessment model to obtain an image quality assessment score; Input the image data into a two-dimensional target detection model, infer the image through the two-dimensional target detection model, and output a two-dimensional detection frame of the target and an original confidence thereof; Input the image data into a weather perception model to obtain weather information of a current vehicle; Combine the image quality assessment score and the weather information, correct the original confidence based on a Bayesian posterior principle, and output a two-dimensional confidence considering the influence of the image and the weather factors; Input the point cloud data into a three-dimensional target detection model, infer the point cloud data through the three-dimensional target detection model, and output a three-dimensional detection frame in a current traffic scene and an original confidence thereof; Based on the target distance and the weather condition, correct the original confidence of the three-dimensional detection frame based on a posterior confidence, and obtain an environmental adaptability three-dimensional confidence; Statistically count the number of point clouds contained in the detection frame, and based on a prior statistical relationship between the number of point clouds and detection accuracy, calculate a point number confidence by using Bayesian inference; According to the consistency degree between the spatial geometric size of the detection frame and the typical prior size distribution of the target category, construct a structure consistency matching degree, and output a structure consistency confidence; Take the image quality assessment score, the two-dimensional confidence, the three-dimensional confidence, the point number confidence and the structure consistency confidence as input features, construct a regression estimation function based on logistic regression, and combine a real credibility label with a real IoU, category matching and detection uniqueness as standards to train a credibility regression model, and output an uncertainty estimation of a final multi-modal fusion target detection result.
[0047] In specific implementation, as a preferred embodiment of the present application, the inputting of the image data into the quality assessment model to obtain the image quality assessment score specifically includes: Based on an open source IQA-PyTorch framework, a no-reference image quality assessment model is constructed; In the training stage, a LIVE public image quality assessment dataset is used, and an ARNIQA neural network is used to predict the image quality; The input image is a color three-channel image, the resolution is the original image size, and the output is a normalized image quality score The higher the score, the higher the image quality.
[0048] In specific implementation, as a preferred embodiment of the present application, the image data is input into the two-dimensional target detection model, the image is inferred by the two-dimensional target detection model, and a two-dimensional detection frame of the target and its original confidence are output, which specifically includes: YOLOv11 is used as the image target detection backbone network to perform real-time inference on the image, and the output format is:
[0049] Among them, is the two-dimensional detection frame center coordinate and height and width; is the category label; is the original confidence.
[0050] In specific implementation, as a preferred embodiment of the present application, the image data is input into the weather perception model to obtain weather information of the current vehicle, which specifically includes: The input image is multi-classified using the WeatherNet image weather recognition network, and the weather category is output.
[0051] In specific implementation, as a preferred embodiment of the present application, the image quality evaluation score and the weather information are combined, the original confidence is corrected based on the Bayesian posterior principle, and a two-dimensional confidence considering the influence of the image and the weather factor is output, which specifically includes: The mAP detection accuracy under different image quality score intervals and different weather conditions is respectively counted on a preset verification data set; For each combination , a joint mapping relationship between the model accuracy and the image quality and the weather condition is fitted to construct a two-dimensional confidence correction coefficient , which is expressed as:
[0052] Among them, is the detection accuracy under the current weather and image quality; is the detection accuracy under sunny weather and high image quality in the data set; the correction coefficient represents the decline range of the detection ability; The original confidence is corrected in a Bayesian posterior manner, which is expressed as:
[0053] Among them, is the two-dimensional confidence considering the influence of the image and the weather factor.
[0054] In specific implementation, as a preferred embodiment of the present application, the point cloud data is input into a three-dimensional target detection model, the point cloud data is inferred by the three-dimensional target detection model, and a three-dimensional detection frame and its original confidence in the current traffic scene are output, which specifically includes: a three-dimensional target detection model based on a PointPillars architecture is constructed; the input is a laser radar point cloud, which is encoded by a voxel, a backbone network and a detection head module, and a three-dimensional detection frame is output wherein, is a three-dimensional detection frame center coordinate and width, length and height, is a three-dimensional detection frame heading angle, is an original three-dimensional confidence.
[0055] In specific implementation, as a preferred embodiment of the present application, the original confidence of the three-dimensional detection frame is corrected by a posteriori confidence based on the target distance and the weather condition, and an environment adaptability three-dimensional confidence is obtained, which specifically includes: in the data set, the precision index of the 3D detector under different distance intervals (<10m, 10-20m, 20-30m, 30-40m, …, >100m) and different weather conditions is counted , to obtain a confidence reduction law table; a fitting function , which represents the performance fidelity of the model under the weather and the distance , and the calculation method is:
[0056] wherein, is the detection accuracy under the weather and the distance ; is the reference detection accuracy under sunny weather and close distance (<10m); is a three-dimensional confidence correction coefficient, a cubic spline interpolation method based on experimental priori is adopted to construct a smooth and continuous correction function from the detection accuracy statistics under different distance conditions, which avoids the confidence shock problem caused by the discontinuity of the segmented function; the final corrected three-dimensional confidence is:
[0057] wherein, is a three-dimensional confidence, is an original three-dimensional detection output confidence.
[0058] In a specific implementation, as a preferred embodiment of the present application, the number of point clouds contained in the statistical detection box is counted, and based on the prior statistical relationship between the number of point clouds and the detection accuracy, the point number confidence is calculated by using Bayesian inference, which specifically includes: In the standard verification data set, for each detected target, the number of point clouds in its three-dimensional detection box is recorded The point number is divided into several intervals (<5, 5-20, 20-40, 40-60, >60), and the detection accuracy mAP of the corresponding interval is counted. The mapping relationship between the number of point clouds and the detection accuracy is constructed to obtain the point number confidence correction coefficient. The interval with the most point numbers (>60) is the ideal reference accuracy The correction coefficient corresponding to each point number interval is calculated and expressed as:
[0059] Among them, is the correction coefficient corresponding to each point number interval; is the detection accuracy corresponding to the number of point clouds in the three-dimensional bounding box, is the ideal reference accuracy; According to the corresponding relationship of the interval and its correction coefficient, the point number confidence correction function is constructed based on the cubic spline interpolation method. Let the confidence of the original three-dimensional detector output be , and the point number correction coefficient be , then the point number confidence correction result is expressed as:
[0060] Among them, is the point number correction coefficient, represents the degree of confidence that supports the detection confidence under the current point number; is the original three-dimensional detection output confidence; is the point number confidence combined with the Bayesian correction.
[0061] In a specific implementation, as a preferred embodiment of the present application, the number of point clouds contained in the statistical detection box is counted, and based on the prior statistical relationship between the number of point clouds and the detection accuracy, the point number confidence is calculated by using Bayesian inference, which specifically includes: In the training set or verification set, the real three-dimensional size (length, width, height) of each target is collected, and the mean and covariance matrix are calculated respectively to establish the prior distribution, as follows:
[0062] Among them, is the Gaussian distribution of the target class three-dimensional size, for the target class; is the average size of the corresponding target class; is the covariance matrix, describing the fluctuation range of the size of the class; simplify the feature to an independent Gaussian, denoted as: , ,
[0063] wherein, is the length of the target class, is a Gaussian distribution of the length conforming to the mean and the variance, is the average value of the length of the target class, is the variance of the length of the target class in all training samples, is the width of the target class, is a Gaussian distribution of the width conforming to the mean and the variance, is the average value of the width of the target class, is the variance of the width of the target class, is the height of the target class, is a Gaussian distribution of the height conforming to the mean and the variance, is the average value of the height of the target class is the variance of the height of the target class; For the target output by the detection model, the class label is , the size is , and the Mahalanobis distance is used to measure the distance measure of the deviation of the three-dimensional size in the target class statistical space. The confidence degree of the target output by the detection model with the size of in the size distribution of the belonging class is calculated, denoted as:
[0064] wherein, is the confidence degree of the target output by the detection model in the size distribution of the belonging class; ; The Mahalanobis distance is mapped to the confidence score based on the Gaussian likelihood form, denoted as:
[0065] Assuming that the original detection model output confidence is , the structural consistency correction coefficient is defined as The structural consistency confidence is expressed as:
[0066] wherein, is the structural consistency confidence.
[0067] In specific implementation, as a preferred embodiment of the present application, the five confidence indicators of the image quality evaluation score, the two-dimensional confidence, the three-dimensional confidence, the point number confidence and the structural consistency confidence are taken as input features to construct a regression estimation function based on logistic regression, and a credibility regression model is trained by combining the real credibility labels based on the real IoU, the category matching and the detection uniqueness, and the uncertainty estimation of the final multi-modal fusion target detection result is output, which specifically includes: A rule is constructed based on the target detection data set and the credibility label, and a credibility data set is constructed; For each predicted box , the matching ground truth (GT) is , the IoU score is defined, and the expression is:
[0068] The category consistency score is defined, and the expression is:
[0069] If multiple boxes match the same ground truth (GT), only the one with the highest IoU is retained, and the others are considered as redundant (repeated detection), and the uniqueness penalty term is defined, and the expression is:
[0070] wherein, ; Based on the IoU score , the category consistency score and the uniqueness penalty term , the final real credibility label is calculated, and the expression is:
[0071] wherein, is the final real credibility label; The five confidence indicators are taken as input features, including the two-dimensional detection confidence based on the image quality and weather factor correction , the three-dimensional detection confidence based on the distance and weather condition correction , the support confidence based on the point cloud number and target existence , the structural consistency confidence , and the image quality score , a logistic regression model is constructed, and the expression is:
[0072] wherein, is a Sigmoid function; is a regression weight parameter to be learned; is a bias term; indicates a fusion confidence score of the model output; In order to enhance the physical interpretability, the regression coefficients are non-negative normalized to obtain the fusion weight, and the expression is:
[0073] The weighted average method is used for fusion to obtain the final confidence score, and the expression is:
[0074] wherein, is a fusion weight; is , , , , five confidence indicators.
[0075] In summary, the target detection confidence estimation method for fusing data and model uncertainty provided in the embodiment aims to solve the problem of lack of fine-grained confidence evaluation in the existing target detection system. In the automatic driving environment, the perception system faces challenges such as unstable input of multi-source sensors, complex environmental interference (such as weather like rain, snow and fog), and insufficient model generalization ability. The traditional target detection model only outputs a confidence score, which often cannot truly reflect the reliability of the detection result, and is easy to cause false decision input, affecting the safety boundary of the system. The present application starts from two aspects of "data uncertainty" and "model uncertainty", and constructs a systematic and multi-factor fusion target detection confidence estimation method. By introducing image quality score, weather perception, detection distance, point cloud support degree and structural consistency, etc. information source, using Bayesian inference, Mahalanobis distance mapping and logistic regression fusion mechanism, a multi-dimensional confidence modeling and joint inference framework is established, which realizes the quantitative confidence evaluation of each detection result, effectively improves the interpretability, controllability and dynamic adaptability of the detection confidence.
[0076] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for target detection confidence estimation fusing data and model uncertainty, characterized in that, The method comprises the following steps: acquiring real-time point cloud data of a vehicle-mounted laser radar sensor and image data of a vehicle-mounted camera; inputting the image data into a quality assessment model to obtain an image quality assessment score; inputting the image data into a two-dimensional target detection model, performing inference on the image by the two-dimensional target detection model, and outputting a two-dimensional detection frame of a target and an original confidence thereof; inputting the image data into a weather perception model to obtain weather information of a current vehicle; combining the image quality assessment score and the weather information, correcting the original confidence based on a Bayesian posterior principle, and outputting a two-dimensional confidence considering the influence of the image and the weather factors; inputting the point cloud data into a three-dimensional target detection model, performing inference on the point cloud data by the three-dimensional target detection model, and outputting a three-dimensional detection frame in a current traffic scene and an original confidence thereof; based on a target distance and a weather condition, performing posterior confidence correction on the original confidence of the three-dimensional detection frame to obtain an environment adaptability three-dimensional confidence; counting the number of point clouds contained in the detection frame, and based on a prior statistical relationship between the number of point clouds and detection accuracy, calculating a point number confidence by Bayesian inference; constructing a structure consistency matching degree according to the consistency degree between the spatial geometric size of the detection frame and the typical prior size distribution of the target category, and outputting a structure consistency confidence; taking the image quality assessment score, the two-dimensional confidence, the three-dimensional confidence, the point number confidence and the structure consistency confidence as input features, constructing a regression estimation function based on logistic regression, and combining a real confidence label with true IoU, category matching and detection uniqueness as standards to train a confidence regression model, and outputting an uncertainty estimation of a final multi-modal fusion target detection result.
2. The method of claim 1, wherein The image data is input into the quality assessment model to obtain the image quality assessment score, specifically comprising: based on the open source IQA-PyTorch framework, a no-reference image quality assessment model is constructed; in the training stage, the LIVE public image quality assessment dataset is used, and the ARNIQA neural network is used to predict the image quality; The input image is a color three-channel image with resolution of the original image size, and the output is a normalized image quality score The higher the score, the higher the image quality.
3. The method of claim 1, wherein The image data is input into the two-dimensional target detection model, the two-dimensional target detection model is used to infer the image, and the two-dimensional detection frame of the target and the original confidence thereof are outputted, specifically comprising: YOLOv11 is used as the image target detection backbone network to perform real-time inference on the image, and the output format is: wherein, is the two-dimensional detection box center coordinate and width and height; is the category label; is the original confidence.
4. The method of claim 1, wherein The image data is input into the weather perception model to obtain the weather information of the current vehicle, specifically comprising: using the WeatherNet image weather recognition network to perform multi-classification on the input image, outputting a weather category .
5. The method of claim 1, wherein The image quality assessment score and the weather information are combined, the original confidence is corrected based on the Bayesian posterior principle, and the two-dimensional confidence considering the influence of the image and the weather factors is outputted, specifically comprising: The mAP detection accuracies under different image quality score intervals and different weather conditions are respectively counted on a preset verification data set ; For each combination , the model accuracy is fitted to obtain the joint mapping relationship of image quality, weather conditions, and construct a two-dimensional confidence correction coefficient , expressed as: wherein, is the detection accuracy under the current weather and image quality; is the detection accuracy under sunny weather and high image quality within the dataset; correction factor represents the magnitude of the decrease in detection capability; The original confidence is corrected by the Bayesian posterior method, which is represented as: wherein, is a two-dimensional confidence taking into account the image and weather factors.
6. The method of claim 1, wherein The point cloud data is input into the three-dimensional target detection model, the three-dimensional target detection model is used to infer the point cloud data, and the three-dimensional detection frame in the current traffic scene and the original confidence thereof are outputted, specifically comprising: a three-dimensional target detection model based on the PointPillars architecture is constructed; The input is a laser radar point cloud, which is encoded by a voxel, a backbone network and a detection head module, and a three-dimensional detection frame is output wherein, is a three-dimensional detection frame center coordinate and width, length and height, is a three-dimensional detection frame heading angle, is an original three-dimensional confidence.
7. The method of claim 1, wherein The original confidence of the three-dimensional detection frame is post-conformance modified based on the target distance and weather conditions to obtain an environment-adaptive three-dimensional confidence, and the method specifically comprises the following steps: In the data set, the accuracy indicators of the 3D detector in different distance intervals and different weather conditions are counted , and a confidence reduction rule table is obtained. Fitting function represents the fidelity of the model's performance under that weather and distance is calculated as: wherein, is the detection accuracy under weather and distance ; is the reference detection accuracy under sunny weather, close distance; is the three-dimensional confidence correction coefficient; The final modified three-dimensional confidence is: wherein, is a three-dimensional confidence, is an original three-dimensional detection output confidence.
8. The method of claim 1, wherein The number of point clouds contained in the detection frame is counted, and the point number confidence is calculated based on the prior statistical relationship between the number of point clouds and the detection accuracy by using Bayesian inference, and the method specifically comprises the following steps: In the standard verification dataset, for each detected target, the number of point clouds within its three-dimensional detection box is recorded The point number is divided into several intervals, and the detection accuracy mAP of the corresponding interval is counted A mapping relationship between the number of point clouds and the detection accuracy is constructed to obtain a point number confidence correction coefficient; The interval with the most points is taken as the ideal reference accuracy The correction coefficient corresponding to each point interval is calculated and expressed as: wherein, is the correction coefficient corresponding to each point interval; is the detection accuracy corresponding to the number of point clouds in the three-dimensional bounding box, is the ideal reference accuracy; According to the corresponding relationship between the interval and the correction coefficient, a point number confidence correction function is constructed based on the cubic spline interpolation method; Let the confidence of the original three-dimensional detector output be , and the point number correction coefficient be , then the point number confidence correction result is expressed as: wherein, is a point correction factor, represents the degree of confidence that the detection confidence is supported at the current point; is an original three-dimensional detection output confidence; is a point confidence combined with the point correction after Bayes correction.
9. The method of claim 1, wherein The number of point clouds contained in the detection frame is counted, and the point number confidence is calculated based on the prior statistical relationship between the number of point clouds and the detection accuracy by using Bayesian inference, and the method specifically comprises the following steps: In the training set or the verification set, the real three-dimensional size of each type of target is collected, the mean and the covariance matrix of each type of target are calculated respectively, and a prior distribution is established, as follows: wherein, is a Gaussian distribution of the target class three-dimensional size, is the target class; denotes the average size of the corresponding target class; is a covariance matrix describing the fluctuation range of the class size; The feature is simplified as an independent Gaussian, and is expressed as: , , wherein, is the length of the class of objects, is a Gaussian distribution with mean and variance, is the mean of the length of the class of objects, is the variance of the length of the class of objects in all training samples, is the width of the class of objects, is a Gaussian distribution with mean and variance, is the mean of the width of the class of objects, is the variance of the width of the class of objects, is the height of the class of objects, is a Gaussian distribution with mean and variance, is the mean of the height of the class of objects is the variance of the height of the class of objects; For a target detected by the detection model output, its category label is , the size is , the distance metric using Mahalanobis distance to measure the deviation degree of three-dimensional size in the target category statistical space, the confidence degree of the target detected by the detection model output in the size distribution of the category to which it belongs is calculated, which is represented as: wherein, to detect a confidence level of a target output by the model in a class size distribution; ; The Mahalanobis distance is mapped to a confidence score based on the Gaussian likelihood form, and a structure consistency correction coefficient is calculated, and the formula is as follows: Assuming the original detection model outputs a confidence , the structure consistency correction coefficient is The structure consistency confidence is represented as: wherein, is a structural consistency confidence.
10. The method of claim 1, wherein The image quality evaluation score, the two-dimensional confidence, the three-dimensional confidence, the point number confidence and the structure consistency confidence are used as input features to construct a regression estimation function based on logistic regression, and a credibility regression model is trained by combining the real credibility label with the real IoU, the category matching and the detection uniqueness as standards, and the uncertainty estimation of the final multi-modal fusion target detection result is output, and the method specifically comprises the following steps: A rule is constructed based on the target detection data set and the credibility label, and a credibility data set is constructed; For each prediction box The IoU score is defined as , where the expression is A category consistency score is defined, and the expression is as follows: If multiple frames match the same true value frame, only the frame with the highest IoU is retained, and the others are considered as redundant, a uniqueness penalty term is defined, and the expression is as follows: wherein ; is a truth box; Based on the IoU score , a class consistency score , and a uniqueness penalty term , a final ground-truth plausibility label is computed, expressed as: wherein, is the final authenticity label; Five confidence indicators are used as input features, including two-dimensional detection confidence based on image quality and weather factor correction , three-dimensional detection confidence based on distance and weather condition correction , support confidence based on the number of point clouds and the existence of the target , structure consistency confidence , image quality score , and a logistic regression model is constructed, expressed as: wherein, is a Sigmoid function; is a regression weight parameter to be learned; is a bias term; denotes the fused confidence score of the model output; In order to enhance the physical interpretability, the regression coefficient is processed by non-negative normalization to obtain a fusion weight, and the expression is as follows: The weighted average method is used for fusion to obtain the final credibility score, and the expression is as follows: wherein, is a fusion weight; is , , , , five confidence indicators.
Citation Information
Cited By
Object detection method and device, equipment and medium
CN121582539A
Confidence evaluation method, device, equipment, medium and product
CN122156869A