Multi-modal target detection integration method based on joint probability and entropy normalization
By adopting an integrated method of joint probability and entropy normalization in multimodal object detection, the problem of reducing detection accuracy in traditional methods in the case of external parameter calibration error and modal loss is solved, and higher detection accuracy and robustness are achieved, and it is suitable for cost-constrained engineering tasks.
Patent Information
- Application Number
- CN202510037131.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-09
AI Technical Summary
In multimodal object detection tasks, traditional integrated methods reduce detection accuracy due to external parameter calibration error or fixed IoU threshold, and it is difficult to effectively deal with the situations of mode missing and external parameter calibration unsatisfactory.
A multimodal object detection integration method based on joint probability and entropy normalization is adopted, and the confidence and detection frames of different object detection models are integrated through joint probability methods, and an entropy normalization is introduced to process external parameter calibration errors to improve detection accuracy and robustness.
It improves the accuracy and robustness of object detection, reduces the cost of model training, enhances the reliability and accuracy of object detection in cost-constrained engineering tasks, and reduces the risk of false detection and missed detection.
Smart Images

Figure CN119963813A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of target detection in computer vision, and relates to a target detection method, and specifically to a multimodal target detection integration method based on joint probability and entropy normalization. Background Art
[0002] With the rapid development of computer technology, target detection technology is increasingly used in various environmental perception tasks to achieve various tasks more safely and efficiently. At present, the target detection method based on deep learning has developed to a relatively mature stage and has achieved quite good accuracy in various data sets. However, the demand for time and hardware resources for training and optimizing complex models on increasingly large data sets continues to rise, making it more difficult to actually train and deploy deep learning models in low-cost projects for specific fields. How to reduce the cost of model training and improve the accuracy of model detection is currently a pain point in target detection in engineering applications. The ensemble method combines the advantages of multiple models to obtain better generalization performance, and improves robustness while reducing the bias of model training.
[0003] Traditional integration methods usually rely on a pre-set IoU threshold to determine whether overlapping objects are the same instance. In the task of multimodal target detection, bumps and vibrations may cause the accuracy of external parameter calibration between sensors to decrease over time. This decrease in accuracy will further cause the integration method to mistakenly determine that the detection boxes provided by different models are different instances when processing multimodal data due to the small IoU between them, thus causing false detection. This problem that frequently occurs in the engineering field still has no proper solution. Summary of the invention
[0004] In order to reduce the learning and training costs of target detection models and solve the problem of reduced detection accuracy in traditional integrated methods due to external parameter calibration errors or fixed IoU thresholds in multimodal target detection tasks, the present invention provides a multimodal target detection integration method based on joint probability and entropy normalization. This method integrates the confidence and detection boxes output by different target detection models through a joint probability method, uses probabilistic marginalization to handle possible modal missing, and introduces entropy normalization to reduce false detections due to external parameter calibration errors, thereby improving the overall target detection accuracy and robustness, which is of great significance for improving the reliability and accuracy of target detection in cost-constrained engineering tasks.
[0005] The objective of the present invention is achieved through the following technical solutions:
[0006] A multimodal target detection integration method based on joint probability and entropy normalization includes the following steps:
[0007] Step 1: Initialize the detector output set D and the integrated detection set F, d i ={s i ,c i}∈D,d i ′={s i ′,c i ′}∈F, Among them, s i ,c i Represents the confidence and detection box coordinates output by the detector, s i ′,c i ′ represents the integrated confidence and detection box coordinates, and the elements in D are divided by s i Arrange them from largest to smallest;
[0008] Step 2: Inter-modal projection: According to the external parameter matrix, the detection frame coordinate points output by the lidar detector are converted from the 3D lidar coordinate system to the 2D camera coordinate system. Then, the lidar points are projected into pixels according to the focal length of the camera in the x-axis and y-axis directions and the origin of the camera coordinate system. The specific projection method is as follows:
[0009] If you want to convert a lidar point Projected to the camera plane, first transform the lidar coordinate system to the camera coordinate system:
[0010]
[0011] in, Indicates point L i 3D coordinates in the laser radar coordinate system, Indicates point L i The 3D coordinates in the camera coordinate system, T is the external parameter matrix, and then the 3D coordinates are converted to 2D coordinates:
[0012]
[0013] Among them, (u i ,v i ) is the pixel coordinate of the lidar point on the camera plane, (f x ,f y ) is the focal length of the camera in the x-axis and y-axis directions in pixels, (c x ,c y ) is the origin of the camera coordinate system. So far, the projection of the laser radar point to the camera pixel point is completed;
[0014] Step 3: Entropy normalization:
[0015] Find s from D i Maximum detection d max, and all detections d that intersect with its detection box i Add to set I and normalize entropy according to the confidence ratio of elements in I:
[0016]
[0017] Among them, H γ is the expected entropy value, H1 is the entropy value when γ = 1, γ is the value to be solved, and p i is the probability distribution of the output result of the i-th detector, and E represents the expectation;
[0018] Step 4: Confidence integration:
[0019] For all the values in I and d max The detections whose IoU exceeds the threshold are compared with their confidence scores d max integrated:
[0020] c i ∝(p(f|x1)p(f|x2)) γ
[0021] Among them, p(f|x i ) represents the conditional probability distribution of the i-th detector;
[0022] Step 5: Detection frame integration:
[0023] Integrate the coordinates of the detection boxes that are greater than the IoU threshold:
[0024]
[0025] Where f represents a continuous random variable consisting of the centroid, width and height of the detection box, μ i Represents the detection box coordinates, is the variance of the distribution;
[0026] Step 6: Loop:
[0027] Remove dmax from D and replace d i ′={s i ′,c i ′} is added to F, and steps 1 to 5 are repeated until D is empty.
[0028] Compared with the prior art, the present invention has the following advantages:
[0029] In view of the increasing training cost of deep learning methods and the unsatisfactory performance of traditional integration methods on multimodal data, the present invention introduces the joint probability of confidence and detection box during independent distribution detection, and combines entropy normalization to approximate the probability distribution between different modes, and proposes a multimodal target detection integration method based on joint probability and entropy normalization. By integrating the output of trained detectors, the present invention provides more accurate detection results than traditional single-modal detectors without requiring additional training costs, and properly handles the situation where traditional integration methods may fail when the modality is missing. In addition, when the external parameter calibration of the multimodal detector is not ideal, it provides higher detection accuracy than traditional integration methods, which plays a positive role in improving the reliability and safety of target detection tasks and reducing the risk of accidents caused by false detection and missed detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 The overall flow chart of the multimodal target detection integration method based on joint probability and entropy normalization. DETAILED DESCRIPTION
[0031] The technical solution of the present invention is further described below in conjunction with the accompanying drawings, but is not limited thereto. Any modification or equivalent replacement of the technical solution of the present invention without departing from the spirit and scope of the technical solution of the present invention should be included in the protection scope of the present invention.
[0032] The present invention provides a multimodal target detection integration method based on joint probability and entropy normalization, comprising the following steps:
[0033] Step 1: Projection between different modes:
[0034] Step 1. Calculate the external parameter matrix between modes:
[0035] The extrinsic matrix describes the transformation from the world coordinate system to the camera coordinate system, defining the position and orientation of the sensor in the world coordinate system. The form of the extrinsic matrix T is generally similar to:
[0036]
[0037] Among them, (U, V, W) is the world coordinate system, (R, T) is the rotation and translation transformation;
[0038] Step 1 and 2: Projection of points between different modes:
[0039] If you want to convert a lidar point Projected to the camera plane, first transform the lidar coordinate system to the camera coordinate system:
[0040]
[0041] in, Indicates point L i 3D coordinates in the laser radar coordinate system, Indicates point L i The 3D coordinates in the camera coordinate system, T is the external parameter matrix, and then the 3D coordinates are converted to 2D coordinates:
[0042]
[0043] Among them, (u i ,v i ) is the pixel coordinate of the lidar point on the camera plane, (f x ,f y ) is the focal length of the camera in the x-axis and y-axis directions in pixels, (c x ,c y ) is the origin of the camera coordinate system. So far, the projection of the laser radar point to the camera pixel point is completed;
[0044] Step 2: Entropy normalization to process external parameter errors:
[0045] Step 21. Definition of entropy:
[0046] For the distribution of n detector results p1,p2,…,p n , their entropy H is defined as:
[0047] H=-Σp i logp i =E[-logp i ]
[0048] Among them, p i is the probability distribution of the output result of the i-th detector, and E represents the expectation;
[0049] P i Use a power transformation:
[0050]
[0051] in, is the power transformation p i The probability distribution of , γ is the value to be solved, which can make the entropy a specified value;
[0052] The power transformation ensures the monotonicity of the distribution while maintaining the main information of the distribution, and the entropy decreases monotonically with respect to γ;
[0053] Step 22: Iteration of entropy:
[0054] Define the expected entropy value H γ :
[0055]
[0056] Expand it at γ = 1:
[0057]
[0058] Right now:
[0059]
[0060] Among them, H1 is the entropy value when γ = 1;
[0061] Starting from γ=1, the expected entropy value γ can be obtained by iterating the above formula, then the updated joint probability distribution can be expressed as:
[0062]
[0063] in, represents the joint probability distribution of n detector results, p(f|x i ) represents the conditional probability distribution of the i-th detector, represents the conditional probability distribution of the i-th detector after power transformation;
[0064] Entropy normalization approximates the distribution between different detectors by softening overconfident estimates, thereby alleviating the degradation of target detection accuracy when the inter-modality extrinsic calibration is not ideal;
[0065] Step 3: Confidence integration:
[0066] Step 31: Integration object screening:
[0067] Suppose that for the same label y, there are n detectors that output x1, x2, ..., x n , select the detection x with the highest confidence score max , according to a pre-set threshold (usually 0.5), and x max Detection results whose IoU is less than this threshold do not participate in the integration process;
[0068] Step 32: Give the conditional independence formula for different detector outputs:
[0069] Since each detector is trained independently, their conditions are independent. So for the two detector outputs x1, x2:
[0070]
[0071] Where p(y) is the prior distribution for label y;
[0072] Extend it to n detectors participating in the ensemble:
[0073]
[0074] This method multiplies the conditional probability of each given individual detector and divides it by the class prior distribution, and normalizes the final result to produce the final score confidence. In this process, this method handles the missing modality through probabilistic marginalization, effectively solving the failure problem that traditional integration methods may encounter when facing modality missing;
[0075] Step 3. Calculate the final score based on the logit score:
[0076] Use logit score to represent the softmax posterior distribution:
[0077]
[0078]
[0079] Among them, s i [k] represents the logit score of the k-th class label under the i-th detector; the above formula gives the calculation method of confidence integration, which is equivalent to adding the logit scores of each detector. When more detectors are consistent, this method can make the detection more convincing and use more stable values for comparison between detectors;
[0080] Step 4: Detection frame integration:
[0081] Step 41: Parameterization of the detection box:
[0082] For the detection results involved in the confidence integration, f is used to represent the continuous random variable consisting of the centroid, width and height of the detection box, and it is assumed that the detector provides a posterior normal distribution where μ i Represents the detection box coordinates, is the variance of the distribution;
[0083] Step 4.2: Calculation of detection box integration:
[0084] The calculation formula for the detection box integration is as follows:
[0085]
[0086] This method integrates the detection box coordinates by calculating the weighted average, where the weight is the inverse covariance, generally using The denominator is the conditional probability distribution of the result of the i-th detector when the label is y, which means that when the detection box is integrated, the detection with higher confidence has a higher weight.
[0087] like Figure 1As shown, the specific steps of the method are as follows:
[0088] Step 1: Initialize the detector output set D and the integrated detection set F, d i ={s i ,c i}∈D,d i ′={s i ′,c i ′}∈F, Among them, s i ,c i Represents the confidence and detection box coordinates output by the detector, s i ′,c i ′ represents the confidence and detection box coordinates after integration of this method. Each element in the set contains two attributes: confidence and detection box coordinate distribution, and F is an empty set at this time. The elements in D are sorted by s i Arrange them from largest to smallest.
[0089] Step 2: Inter-modal projection: According to the extrinsic matrix, the detection frame coordinate points output by the lidar detector are converted from the 3D lidar coordinate system to the 2D camera coordinate system, and then the lidar points are projected into pixels according to the length of the camera's focal length in the x-axis and y-axis directions and the origin of the camera coordinate system.
[0090] Step 3: Entropy normalization:
[0091] Find s from D i Maximum detection d max , and all detections d that intersect with its detection box i Add to set I and normalize entropy according to the confidence ratio of elements in I:
[0092]
[0093] Step 4: Confidence integration:
[0094] For all the values in I and d max The detections whose IoU exceeds the threshold (usually set to 0.5) are compared with their confidence max integrated:
[0095] c i ∝(p(f|x1)p(f|x2)) γ
[0096] Step 5: Detection frame integration:
[0097] Integrate the coordinates of the detection boxes that are greater than the IoU threshold:
[0098]
[0099] Step 6: Loop:
[0100] Remove dmax from D and replace d i ′={s i ′,c i ′} is added to F, and steps 1 to 5 are repeated until D is empty.
[0101] Embodiment 1:
[0102] This example verifies the method of the present invention on the single-modal public dataset VOC dataset, and selects five common labels on the road: bicycle, bus, car, motorbike, person, and a total of 10,870 pictures for experimentation. The method uses Mask R-CNN and Faster R-CNN as the integrated benchmark and uses mAP as the evaluation index. The implementation results are shown in Table 1:
[0103] Table 1 Implementation results on the VOC dataset
[0104]
[0105] The results show that in single-modal datasets, the detection accuracy of the proposed method is significantly improved compared to the traditional integrated method. In addition, the inference speed of this method on Nvidia AGX Xavier can reach 20fps, which meets the general real-time requirement of 10fps for target detection.
[0106] Embodiment 2:
[0107] This example verifies the method of the present invention on the multimodal public dataset KITTI dataset, and selects three common labels on the road: Car, Pedestrian, and Cyclist, with a total of 7518 frames of data for the experiment. The method uses PointPillars and Faster R-CNN as the integrated benchmark and uses mAP as the evaluation index. The implementation results are shown in Table 2:
[0108] Table 2 Implementation results on the KITTI dataset
[0109]
[0110]
[0111] The results show that in multimodal datasets, the detection accuracy of the proposed method is significantly improved compared with the traditional integrated method. In addition, the inference speed of this method on Nvidia AGX Xavier can reach 14fps, which meets the general real-time requirement of 10fps for target detection.
[0112] It can be seen from the above embodiments that the method of the present invention improves the problem of false detection and missed detection in the field of target detection compared with traditional integrated methods, has the ability to resist modality loss, and has higher accuracy and stronger robustness.
Claims
1. A multimodal target detection ensemble method based on joint probability and entropy normalization, characterized in that The method comprises the following steps: Step 1: Initialize the detector output set D and the integrated detection set F, d i ={s i ,c i }∈D,d i ′={s i ′,c i ′}∈F, Among them, s i ,c i Represents the confidence and detection box coordinates output by the detector, s i ′,c i ′ represents the integrated confidence and detection box coordinates, and the elements in D are divided by s i Arrange them from largest to smallest; Step 2: Inter-modal projection: According to the external parameter matrix, the detection frame coordinate points output by the lidar detector are converted from the 3D lidar coordinate system to the 2D camera coordinate system, and then the lidar points are projected into pixels according to the focal length of the camera in the x-axis and y-axis directions and the origin of the camera coordinate system; Step 3: Entropy normalization: Find s from D i Maximum detection d max , and all detections d that intersect with its detection box i Add to set I and normalize entropy according to the confidence ratio of elements in I: Among them, H γ is the expected entropy value, H1 is the entropy value when γ = 1, γ is the value to be solved, and p i is the probability distribution of the output result of the i-th detector, and E represents the expectation; Step 4: Confidence integration: For all the values in I and d max The detections whose IoU exceeds the threshold are compared with their confidence scores d max integrated: c i ∝(p(f|x1)p(f|x2)) γ Among them, p(f|x i ) represents the conditional probability distribution of the i-th detector; Step 5: Detection frame integration: Integrate the coordinates of the detection boxes that are greater than the IoU threshold: Where f represents a continuous random variable consisting of the centroid, width and height of the detection box, μ i Represents the detection box coordinates, is the variance of the distribution; Step 6: Loop: Remove dmax from D and replace d i ′={s i ′,c i ′} is added to F, and steps 1 to 5 are repeated until D is empty.
2. The multimodal target detection integration method based on joint probability and entropy normalization according to claim 1 is characterized in that In step 2, the extrinsic matrix describes the transformation from the world coordinate system to the camera coordinate system, and defines the position and direction of the sensor in the world coordinate system. The form of the extrinsic matrix T is as follows: Among them, (U, V, W) is the world coordinate system, and (R, T) is the rotation and translation transformation.
3. The multimodal target detection integration method based on joint probability and entropy normalization according to claim 1 is characterized in that In the step 2, the specific projection method is as follows: If you want to convert a lidar point Projected to the camera plane, first transform the lidar coordinate system to the camera coordinate system: in, Indicates point L i 3D coordinates in the laser radar coordinate system, Indicates point L i The 3D coordinates in the camera coordinate system, T is the external parameter matrix, and then the 3D coordinates are converted to 2D coordinates: Among them, (u i ,v i ) is the pixel coordinate of the lidar point on the camera plane, (f x ,f y ) is the focal length of the camera in the x-axis and y-axis directions in pixels, (c x ,c y ) is the origin of the camera coordinate system. At this point, the projection of the lidar point to the camera pixel point is completed.
4. The multimodal target detection integration method based on joint probability and entropy normalization according to claim 1 is characterized in that In the step five, The denominator is the conditional probability distribution of the result of the i-th detector when the label is y.
Citation Information
Patent Citations
Improved YOLOv3 algorithm based on EIOU
CN112418212A
Collected cultural relic image data semantic association method based on multi-label image classification
CN117392420A
Dynamic target detection and tracking method based on camera and laser radar data fusion
CN118711030A
Target detection method based on multi-modal data fusion
CN118840637A
Systems and methods for camera-lidar fused object detection with lidar-to-image detection matching
US20220128701A1