Multi-modal target detection integration method based on joint probability and entropy normalization

By adopting an integrated method of joint probability and entropy normalization in multimodal object detection, the problem of reducing detection accuracy in traditional methods in the case of external parameter calibration error and modal loss is solved, and higher detection accuracy and robustness are achieved, and it is suitable for cost-constrained engineering tasks.

CN119963813AActive Publication Date: 2025-05-09HARBIN INST OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510037131.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-09
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

In multimodal object detection tasks, traditional integrated methods reduce detection accuracy due to external parameter calibration error or fixed IoU threshold, and it is difficult to effectively deal with the situations of mode missing and external parameter calibration unsatisfactory.

Method used

A multimodal object detection integration method based on joint probability and entropy normalization is adopted, and the confidence and detection frames of different object detection models are integrated through joint probability methods, and an entropy normalization is introduced to process external parameter calibration errors to improve detection accuracy and robustness.

Benefits of technology

It improves the accuracy and robustness of object detection, reduces the cost of model training, enhances the reliability and accuracy of object detection in cost-constrained engineering tasks, and reduces the risk of false detection and missed detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963813A_ABST
    Figure CN119963813A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal target detection integration method based on joint probability and entropy normalization. The method comprises the following steps: 1, initializing a detector output set and an integrated detection set; 2, performing inter-modal projection; 3, carrying out entropy normalization; 4, confidence integration; 5, integrating a detection frame; and 6, circulating. According to the method, the output of the trained detector is integrated, the detection result which is more accurate than that of a traditional single-mode detector is given while extra training cost is not needed, the situation that a traditional integration method may fail when the mode is missing is well handled, and in addition, the method has the advantage that the detection accuracy is improved. When external parameter calibration of the multi-mode detector is not ideal, detection precision higher than that of a traditional integration method is provided, the method plays a positive role in improving reliability and safety of a target detection task, and the risk of accidents caused by false detection and missing detection is reduced. The method is of great significance in improving the reliability and accuracy of target detection in cost-limited engineering tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of target detection in computer vision, and relates to a target detection method, and specifically to a multimodal target detection integration method based on joint probability and entropy normalization. Background Art

[0002] With the rapid development of computer technology, target detection technology is increasingly used in various environmental perception tasks to achieve various tasks more safely and efficiently. At present, the target detection method based on deep learning has developed to a relatively mature stage and has achieved quite good accuracy in various data sets. However, the demand for time and hardware resources for training and optimizing complex models on increasingly large data sets continues to rise, making it more difficult to actually train and deploy deep learning models in low-cost projects for specific fields. How to reduce the cost of model training and improve the accuracy of model detection is currently a pain point in target detection in engineering applications. The ensemble method combines the advantages of multiple models to obtain better generalization performance, and improves robustness while reducing the bias of model training.

[0003] Traditional integration methods usually rely on a pre-set IoU threshold to determine whether overlapping objects are the same instance. In the task of multimodal target detection, bumps and vibrations may cause the accuracy of external parameter calibration between sensors to decrease over time. This decrease in accuracy will further cause the integration method to mistakenly determine that the detection boxes provided by different models are different instances when processing multimodal data due to the small IoU between them, thus causing false detection. This problem that frequently occurs in the engineering field still has no proper solution. Summary of the invention

[0004] In order to reduce the learning and training costs of target detection models and solve the problem of reduced detection accuracy in traditional integrated methods due to external parameter calibration errors or fixed IoU thresholds in multimodal target detection tasks, the present invention provides a multimodal target detection integration method based on joint probability and entropy normalization. This method integrates the confidence and detection boxes output by different target detection models through a joint probability method, uses probabilistic marginalization to handle possible modal missing, and introduces entropy normalization to reduce false detections due to external parameter calibration errors, thereby improving the overall target detection accuracy and robustness, which is of great significance for improving the reliability and accuracy of target detection in cost-constrained engineering tasks.

[0005] The objective of the present invention is achieved through the following technical solutions:

[0006] A multimodal target detection integration method based on joint probability and entropy normalization includes the following steps:

[0007] Step 1: Initialize the detector output set D and the integrated detection set F, d i ={s i ,c i}∈D,d i ′={s i ′,c i ′}∈F, Among them, s i ,c i Represents the confidence and detection box coordinates output by the detector, s i ′,c i ′ represents the integrated confidence and detection box coordinates, and the elements in D are divided by s i Arrange them from largest to smallest;

[0008] Step 2: Inter-modal projection: According to the external parameter matrix, the detection frame coordinate points output by the lidar detector are converted from the 3D lidar coordinate system to the 2D camera coordinate system. Then, the lidar points are projected into pixels according to the focal length of the camera in the x-axis and y-axis directions and the origin of the camera coordinate system. The specific projection method is as follows:

[0009] If you want to convert a lidar point Projected to the camera plane, first transform the lidar coordinate system to the camera coordinate system:

[0010]

[0011] in, Indicates point L i 3D coordinates in the laser radar coordinate system, Indicates point L i The 3D coordinates in the camera coordinate system, T is the external parameter matrix, and then the 3D coordinates are converted to 2D coordinates:

[0012]

[0013] Among them, (u i ,v i ) is the pixel coordinate of the lidar point on the camera plane, (f x ,f y ) is the focal length of the camera in the x-axis and y-axis directions in pixels, (c x ,c y ) is the origin of the camera coordinate system. So far, the projection of the laser radar point to the camera pixel point is completed;

[0014] Step 3: Entropy normalization:

[0015] Find s from D i Maximum detection d max, and all detections d that intersect with its detection box i Add to set I and normalize entropy according to the confidence ratio of elements in I:

[0016]

[0017] Among them, H γ is the expected entropy value, H1 is the entropy value when γ = 1, γ is the value to be solved, and p i is the probability distribution of the output result of the i-th detector, and E represents the expectation;

[0018] Step 4: Confidence integration:

[0019] For all the values ​​in I and d max The detections whose IoU exceeds the threshold are compared with their confidence scores d max integrated:

[0020] c i ∝(p(f|x1)p(f|x2)) γ

[0021] Among them, p(f|x i ) represents the conditional probability distribution of the i-th detector;

[0022] Step 5: Detection frame integration:

[0023] Integrate the coordinates of the detection boxes that are greater than the IoU threshold:

[0024]

[0025] Where f represents a continuous random variable consisting of the centroid, width and height of the detection box, μ i Represents the detection box coordinates, is the variance of the distribution;

[0026] Step 6: Loop:

[0027] Remove dmax from D and replace d i ′={s i ′,c i ′} is added to F, and steps 1 to 5 are repeated until D is empty.

[0028] Compared with the prior art, the present invention has the following advantages:

[0029] In view of the increasing training cost of deep learning methods and the unsatisfactory performance of traditional integration methods on multimodal data, the present invention introduces the joint probability of confidence and detection box during independent distribution detection, and combines entropy normalization to approximate the probability distribution between different modes, and proposes a multimodal target detection integration method based on joint probability and entropy normalization. By integrating the output of trained detectors, the present invention provides more accurate detection results than traditional single-modal detectors without requiring additional training costs, and properly handles the situation where traditional integration methods may fail when the modality is missing. In addition, when the external parameter calibration of the multimodal detector is not ideal, it provides higher detection accuracy than traditional integration methods, which plays a positive role in improving the reliability and safety of target detection tasks and reducing the risk of accidents caused by false detection and missed detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 The overall flow chart of the multimodal target detection integration method based on joint probability and entropy normalization. DETAILED DESCRIPTION

[0031] The technical solution of the present invention is further described below in conjunction with the accompanying drawings, but is not limited thereto. Any modification or equivalent replacement of the technical solution of the present invention without departing from the spirit and scope of the technical solution of the present invention should be included in the protection scope of the present invention.

[0032] The present invention provides a multimodal target detection integration method based on joint probability and entropy normalization, comprising the following steps:

[0033] Step 1: Projection between different modes:

[0034] Step 1. Calculate the external parameter matrix between modes:

[0035] The extrinsic matrix describes the transformation from the world coordinate system to the camera coordinate system, defining the position and orientation of the sensor in the world coordinate system. The form of the extrinsic matrix T is generally similar to:

[0036]

[0037] Among them, (U, V, W) is the world coordinate system, (R, T) is the rotation and translation transformation;

[0038] Step 1 and 2: Projection of points between different modes:

[0039] If you want to convert a lidar point Projected to the camera plane, first transform the lidar coordinate system to the camera coordinate system:

[0040]

[0041] in, Indicates point L i 3D coordinates in the laser radar coordinate system, Indicates point L i The 3D coordinates in the camera coordinate system, T is the external parameter matrix, and then the 3D coordinates are converted to 2D coordinates:

[0042]

[0043] Among them, (u i ,v i ) is the pixel coordinate of the lidar point on the camera plane, (f x ,f y ) is the focal length of the camera in the x-axis and y-axis directions in pixels, (c x ,c y ) is the origin of the camera coordinate system. So far, the projection of the laser radar point to the camera pixel point is completed;

[0044] Step 2: Entropy normalization to process external parameter errors:

[0045] Step 21. Definition of entropy:

[0046] For the distribution of n detector results p1,p2,…,p n , their entropy H is defined as:

[0047] H=-Σp i logp i =E[-logp i ]

[0048] Among them, p i is the probability distribution of the output result of the i-th detector, and E represents the expectation;

[0049] P i Use a power transformation:

[0050]

[0051] in, is the power transformation p i The probability distribution of , γ is the value to be solved, which can make the entropy a specified value;

[0052] The power transformation ensures the monotonicity of the distribution while maintaining the main information of the distribution, and the entropy decreases monotonically with respect to γ;

[0053] Step 22: Iteration of entropy:

[0054] Define the expected entropy value H γ :

[0055]

[0056] Expand it at γ = 1:

[0057]

[0058] Right now:

[0059]

[0060] Among them, H1 is the entropy value when γ = 1;

[0061] Starting from γ=1, the expected entropy value γ can be obtained by iterating the above formula, then the updated joint probability distribution can be expressed as:

[0062]

[0063] in, represents the joint probability distribution of n detector results, p(f|x i ) represents the conditional probability distribution of the i-th detector, represents the conditional probability distribution of the i-th detector after power transformation;

[0064] Entropy normalization approximates the distribution between different detectors by softening overconfident estimates, thereby alleviating the degradation of target detection accuracy when the inter-modality extrinsic calibration is not ideal;

[0065] Step 3: Confidence integration:

[0066] Step 31: Integration object screening:

[0067] Suppose that for the same label y, there are n detectors that output x1, x2, ..., x n , select the detection x with the highest confidence score max , according to a pre-set threshold (usually 0.5), and x max Detection results whose IoU is less than this threshold do not participate in the integration process;

[0068] Step 32: Give the conditional independence formula for different detector outputs:

[0069] Since each detector is trained independently, their conditions are independent. So for the two detector outputs x1, x2:

[0070]

[0071] Where p(y) is the prior distribution for label y;

[0072] Extend it to n detectors participating in the ensemble:

[0073]

[0074] This method multiplies the conditional probability of each given individual detector and divides it by the class prior distribution, and normalizes the final result to produce the final score confidence. In this process, this method handles the missing modality through probabilistic marginalization, effectively solving the failure problem that traditional integration methods may encounter when facing modality missing;

[0075] Step 3. Calculate the final score based on the logit score:

[0076] Use logit score to represent the softmax posterior distribution:

[0077]

[0078]

[0079] Among them, s i [k] represents the logit score of the k-th class label under the i-th detector; the above formula gives the calculation method of confidence integration, which is equivalent to adding the logit scores of each detector. When more detectors are consistent, this method can make the detection more convincing and use more stable values ​​for comparison between detectors;

[0080] Step 4: Detection frame integration:

[0081] Step 41: Parameterization of the detection box:

[0082] For the detection results involved in the confidence integration, f is used to represent the continuous random variable consisting of the centroid, width and height of the detection box, and it is assumed that the detector provides a posterior normal distribution where μ i Represents the detection box coordinates, is the variance of the distribution;

[0083] Step 4.2: Calculation of detection box integration:

[0084] The calculation formula for the detection box integration is as follows:

[0085]

[0086] This method integrates the detection box coordinates by calculating the weighted average, where the weight is the inverse covariance, generally using The denominator is the conditional probability distribution of the result of the i-th detector when the label is y, which means that when the detection box is integrated, the detection with higher confidence has a higher weight.

[0087] like Figure 1As shown, the specific steps of the method are as follows:

[0088] Step 1: Initialize the detector output set D and the integrated detection set F, d i ={s i ,c i}∈D,d i ′={s i ′,c i ′}∈F, Among them, s i ,c i Represents the confidence and detection box coordinates output by the detector, s i ′,c i ′ represents the confidence and detection box coordinates after integration of this method. Each element in the set contains two attributes: confidence and detection box coordinate distribution, and F is an empty set at this time. The elements in D are sorted by s i Arrange them from largest to smallest.

[0089] Step 2: Inter-modal projection: According to the extrinsic matrix, the detection frame coordinate points output by the lidar detector are converted from the 3D lidar coordinate system to the 2D camera coordinate system, and then the lidar points are projected into pixels according to the length of the camera's focal length in the x-axis and y-axis directions and the origin of the camera coordinate system.

[0090] Step 3: Entropy normalization:

[0091] Find s from D i Maximum detection d max , and all detections d that intersect with its detection box i Add to set I and normalize entropy according to the confidence ratio of elements in I:

[0092]

[0093] Step 4: Confidence integration:

[0094] For all the values ​​in I and d max The detections whose IoU exceeds the threshold (usually set to 0.5) are compared with their confidence max integrated:

[0095] c i ∝(p(f|x1)p(f|x2)) γ

[0096] Step 5: Detection frame integration:

[0097] Integrate the coordinates of the detection boxes that are greater than the IoU threshold:

[0098]

[0099] Step 6: Loop:

[0100] Remove dmax from D and replace d i ′={s i ′,c i ′} is added to F, and steps 1 to 5 are repeated until D is empty.

[0101] Embodiment 1:

[0102] This example verifies the method of the present invention on the single-modal public dataset VOC dataset, and selects five common labels on the road: bicycle, bus, car, motorbike, person, and a total of 10,870 pictures for experimentation. The method uses Mask R-CNN and Faster R-CNN as the integrated benchmark and uses mAP as the evaluation index. The implementation results are shown in Table 1:

[0103] Table 1 Implementation results on the VOC dataset

[0104]

[0105] The results show that in single-modal datasets, the detection accuracy of the proposed method is significantly improved compared to the traditional integrated method. In addition, the inference speed of this method on Nvidia AGX Xavier can reach 20fps, which meets the general real-time requirement of 10fps for target detection.

[0106] Embodiment 2:

[0107] This example verifies the method of the present invention on the multimodal public dataset KITTI dataset, and selects three common labels on the road: Car, Pedestrian, and Cyclist, with a total of 7518 frames of data for the experiment. The method uses PointPillars and Faster R-CNN as the integrated benchmark and uses mAP as the evaluation index. The implementation results are shown in Table 2:

[0108] Table 2 Implementation results on the KITTI dataset

[0109]

[0110]

[0111] The results show that in multimodal datasets, the detection accuracy of the proposed method is significantly improved compared with the traditional integrated method. In addition, the inference speed of this method on Nvidia AGX Xavier can reach 14fps, which meets the general real-time requirement of 10fps for target detection.

[0112] It can be seen from the above embodiments that the method of the present invention improves the problem of false detection and missed detection in the field of target detection compared with traditional integrated methods, has the ability to resist modality loss, and has higher accuracy and stronger robustness.

Claims

1. A multimodal target detection ensemble method based on joint probability and entropy normalization, characterized in that The method comprises the following steps: Step 1: Initialize the detector output set D and the integrated detection set F, d i ={s i ,c i }∈D,d i ′={s i ′,c i ′}∈F, Among them, s i ,c i Represents the confidence and detection box coordinates output by the detector, s i ′,c i ′ represents the integrated confidence and detection box coordinates, and the elements in D are divided by s i Arrange them from largest to smallest; Step 2: Inter-modal projection: According to the external parameter matrix, the detection frame coordinate points output by the lidar detector are converted from the 3D lidar coordinate system to the 2D camera coordinate system, and then the lidar points are projected into pixels according to the focal length of the camera in the x-axis and y-axis directions and the origin of the camera coordinate system; Step 3: Entropy normalization: Find s from D i Maximum detection d max , and all detections d that intersect with its detection box i Add to set I and normalize entropy according to the confidence ratio of elements in I: Among them, H γ is the expected entropy value, H1 is the entropy value when γ = 1, γ is the value to be solved, and p i is the probability distribution of the output result of the i-th detector, and E represents the expectation; Step 4: Confidence integration: For all the values ​​in I and d max The detections whose IoU exceeds the threshold are compared with their confidence scores d max integrated: c i ∝(p(f|x1)p(f|x2)) γ Among them, p(f|x i ) represents the conditional probability distribution of the i-th detector; Step 5: Detection frame integration: Integrate the coordinates of the detection boxes that are greater than the IoU threshold: Where f represents a continuous random variable consisting of the centroid, width and height of the detection box, μ i Represents the detection box coordinates, is the variance of the distribution; Step 6: Loop: Remove dmax from D and replace d i ′={s i ′,c i ′} is added to F, and steps 1 to 5 are repeated until D is empty.

2. The multimodal target detection integration method based on joint probability and entropy normalization according to claim 1 is characterized in that In step 2, the extrinsic matrix describes the transformation from the world coordinate system to the camera coordinate system, and defines the position and direction of the sensor in the world coordinate system. The form of the extrinsic matrix T is as follows: Among them, (U, V, W) is the world coordinate system, and (R, T) is the rotation and translation transformation.

3. The multimodal target detection integration method based on joint probability and entropy normalization according to claim 1 is characterized in that In the step 2, the specific projection method is as follows: If you want to convert a lidar point Projected to the camera plane, first transform the lidar coordinate system to the camera coordinate system: in, Indicates point L i 3D coordinates in the laser radar coordinate system, Indicates point L i The 3D coordinates in the camera coordinate system, T is the external parameter matrix, and then the 3D coordinates are converted to 2D coordinates: Among them, (u i ,v i ) is the pixel coordinate of the lidar point on the camera plane, (f x ,f y ) is the focal length of the camera in the x-axis and y-axis directions in pixels, (c x ,c y ) is the origin of the camera coordinate system. At this point, the projection of the lidar point to the camera pixel point is completed.

4. The multimodal target detection integration method based on joint probability and entropy normalization according to claim 1 is characterized in that In the step five, The denominator is the conditional probability distribution of the result of the i-th detector when the label is y.

Citation Information

Patent Citations

  • Improved YOLOv3 algorithm based on EIOU

    CN112418212A

  • Collected cultural relic image data semantic association method based on multi-label image classification

    CN117392420A

  • Dynamic target detection and tracking method based on camera and laser radar data fusion

    CN118711030A

  • Target detection method based on multi-modal data fusion

    CN118840637A

  • Systems and methods for camera-lidar fused object detection with lidar-to-image detection matching

    US20220128701A1