3D Target Detection Method Based on Multi-Camera and Millimeter-Wave Radar Fusion

By constructing feature alignment, information gain, and task performance models, the problem of inaccurate evaluation in multimodal feature fusion is solved, achieving efficient and accurate fusion of multimodal features and improving the overall performance and robustness of 3D object detection.

CN121144964BActive Publication Date: 2026-01-30JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511667524.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-01-30
Estimated Expiration
2045-11-14

AI Technical Summary

Technical Problem

Existing 3D target detection methods lack a unified quantitative evaluation of the consistency of multimodal feature distribution, spatial alignment accuracy, and information redundancy, resulting in blind fusion strategies and failure to establish an end-to-end evaluation chain from feature alignment to task performance, making it difficult to achieve efficient deployment.

Method used

We construct a feature alignment model, an information gain model, a task performance model, and an efficiency optimization model. By using indicators such as feature distribution similarity, spatial feature consistency, information redundancy, positioning accuracy gain, and classification confidence, we can quantitatively evaluate and optimize the multimodal fusion process, and achieve dynamic matching and resource balance of feature alignment, information gain, and task performance.

Benefits of technology

It improves the compatibility and consistency of multimodal feature fusion, maximizes the extraction of complementary information, ensures that the fusion results have a clear performance gain in actual detection tasks, achieves an adaptive balance between detection accuracy and computing resources, and improves the system's 3D target detection performance in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121144964B_ABST
    Figure CN121144964B_ABST
Patent Text Reader

Abstract

This invention discloses a 3D target detection method based on the fusion of multi-camera and millimeter-wave radar, belonging to the field of target detection technology. The method includes the following steps: constructing a feature alignment model and outputting feature alignment coefficients; constructing information gain coefficients; constructing a task performance model and outputting task performance coefficients; constructing an alignment-performance matching model and outputting alignment-performance matching coefficients; and constructing an efficiency optimization model based on the alignment-performance matching coefficients, information gain coefficients, and the current fusion computation efficiency ratio, outputting the target fusion computation efficiency ratio. This invention achieves unified measurement of the feature layer, information layer, and task layer through multi-level coefficient collaborative evaluation and dynamic optimization, effectively improving the accuracy and system efficiency of 3D target detection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of target detection, and particularly relates to a 3D target detection method based on multi-camera and millimeter wave radar fusion. BACKGROUND

[0002] With the rapid development of automatic driving, intelligent transportation and other fields, higher requirements are put forward for accurate 3D detection of target objects in the environment. Due to the inherent limitations, a single sensor is difficult to meet the sensing needs in complex scenes, and multi-sensor fusion technology has become a key way to improve system robustness and accuracy. The fusion method based on multi-camera and millimeter wave radar has attracted much attention because it combines visual semantic information and radar ranging capability.

[0003] Existing 3D target detection fusion methods are mostly focused on data-level or decision-level fusion, such as simply concatenating image and radar features, or post-fusing respective detection results. These methods often ignore the feature distribution difference between multi-modal data, spatial alignment error and information complementary characteristics, resulting in limited fusion effect. Some researches try to introduce attention mechanism or shallow feature interaction, but still lack systematic modeling of the coupling relationship between feature alignment quality, information gain degree and task efficiency.

[0004] The existing technology mainly has the following defects: first, there is a lack of unified quantitative evaluation of multi-modal feature distribution consistency, spatial alignment accuracy and information redundancy, resulting in blind fusion strategy; second, it fails to establish an end-to-end evaluation chain from feature alignment to task efficiency, making it impossible to realize closed-loop optimization of the fusion process; third, it ignores the dynamic balance between computational efficiency and detection accuracy, making it difficult to achieve efficient deployment in actual systems.

[0005] Therefore, there is an urgent need for a method that can systematically evaluate and optimize the multi-modal fusion process to improve the overall performance and practicality of 3D target detection. SUMMARY

[0006] In view of the deficiencies of the prior art, the present application provides a 3D target detection method based on multi-camera and millimeter wave radar fusion, which solves the above problems.

[0007] To achieve the above purpose, the present application realizes the following technical scheme: a 3D target detection method based on multi-camera and millimeter wave radar fusion, comprising the following steps:

[0008] constructing a feature alignment model based on feature distribution similarity, spatial feature consistency and feature space alignment error to output a feature alignment coefficient;

[0009] constructing an information gain coefficient based on inter-modal information redundancy, feature dimension effectiveness and unique information contribution;

[0010] constructing a task performance model based on the positioning accuracy gain and the classification confidence promotion degree to output a task performance coefficient;

[0011] constructing an alignment-performance matching model based on the feature response intensity ratio, the feature alignment coefficient under the detection quality promotion degree, and the task performance coefficient to output an alignment-performance matching coefficient;

[0012] constructing an efficiency optimization model based on the alignment-performance matching coefficient, the information gain coefficient, and the current fusion calculation efficiency ratio to output a target fusion calculation efficiency ratio.

[0013] On the basis of the above technical solutions, the application further provides the following optional technical solutions:

[0014] A further technical solution is that the efficiency optimization model is represented as:

[0015]

[0016] wherein, the target fusion calculation efficiency ratio is represented as, the current fusion calculation efficiency ratio is represented as, the alignment-performance matching coefficient is represented as, the information gain coefficient is represented as, the performance threshold is represented as.

[0017] A further technical solution is that the step of constructing the alignment-performance matching model based on the feature response intensity ratio, the feature alignment coefficient under the detection quality promotion degree, and the task performance coefficient to output the alignment-performance matching coefficient is:

[0018] performing maximum-minimum normalization processing on the detection quality promotion degree to obtain a detection quality promotion degree index;

[0019] obtaining the feature response intensity ratio, performing ratio processing on the absolute difference between the feature response intensity ratio and the optimal feature response intensity ratio and the allowed deviation value of the optimal feature response intensity ratio to obtain a feature response intensity ratio index;

[0020] constructing the alignment-performance matching model based on the feature response intensity ratio index and the feature alignment coefficient under the detection quality promotion degree index and the task performance coefficient to obtain the alignment-performance matching coefficient, wherein the alignment-performance matching model is represented as:

[0021]

[0022] wherein, the alignment-performance matching coefficient is represented as, the feature alignment coefficient is represented as, the task performance coefficient is represented as, represents the prevention of zero decimals (generally 10 -6 ), represents the characteristic response intensity ratio index, represents the detection quality improvement degree, and the The greater the value, the higher the degree of coordination between the feature alignment and the task performance.

[0023] Further technical solutions: the step of constructing a task performance model based on the positioning accuracy gain and the classification confidence improvement degree to output a task performance coefficient is:

[0024] Obtain the positioning accuracy gain and the classification confidence, and perform maximum-minimum normalization processing on the two, to obtain a positioning accuracy gain index and a classification confidence index;

[0025] Construct a task performance model based on the positioning accuracy gain index and the classification confidence index, and obtain a task performance coefficient, wherein the task performance model is represented as:

[0026]

[0027] Among them, represents the task performance coefficient, represents the positioning accuracy gain index, represents the classification confidence index, and the The greater the value, the better the task performance.

[0028] Further technical solutions: the step of constructing an information gain coefficient based on inter-modal information redundancy, feature dimension effectiveness, and unique information contribution degree is:

[0029] Obtain the inter-modal information redundancy, the feature dimension effectiveness, and the unique information contribution degree, and perform maximum-minimum normalization processing to obtain a redundancy index, an effectiveness index, and a contribution degree index;

[0030] Construct an information gain model based on the redundancy index, the effectiveness index, and the contribution degree index, and obtain an information gain coefficient, wherein the information gain model is represented as:

[0031]

[0032] Among them, represents the information gain coefficient, represents the redundancy index, represents the effectiveness index, represents the contribution degree index, and the The greater the value, the higher the fusion value of image features and radar features (the more significant the information gain).

[0033] Further technical solution: The steps for constructing a feature alignment model based on feature distribution similarity, spatial feature consistency, and feature space alignment error, and outputting feature alignment coefficients, are as follows:

[0034] The feature distribution similarity, spatial feature consistency, and feature space alignment error are obtained and then subjected to maximum-minimum normalization to obtain the feature distribution similarity index, feature space alignment error index, and spatial feature consistency index.

[0035] A feature alignment model is constructed based on the feature distribution similarity index, the feature space alignment error index, and the spatial feature consistency index, and the feature alignment coefficient is output. The feature alignment model is expressed as follows:

[0036]

[0037] in, Represents the feature alignment coefficient. Represents the similarity index of feature distributions. This represents the feature space alignment error index. Indicators representing spatial feature consistency indices Represents the weight coefficient and The Furthermore, the larger the value, the more consistent the image features are with the radar features.

[0038] Further technical solutions: fusion computing efficiency ratio , This represents the total computation time of the fusion process. Indicates the image branching time. This indicates the radar branch processing time.

[0039] The 3D target detection method based on the fusion of multi-camera and millimeter-wave radar is adopted.

[0040] This invention provides a 3D target detection method based on the fusion of multi-camera and millimeter-wave radar, which has the following advantages compared with the prior art:

[0041] 1. This invention constructs a feature alignment model to quantitatively evaluate the alignment degree between image and radar features in terms of distribution, space, and transformation, providing a high-quality feature foundation for fusion and improving the compatibility and consistency of intermodal information fusion.

[0042] 2. This invention can comprehensively evaluate the information redundancy between modalities, the effectiveness of feature dimensions, and the contribution of unique information by using the information gain coefficient, so as to ensure that the fusion process maximizes the extraction of complementary information, avoids information overlap or loss, and improves the value and efficiency of feature fusion.

[0043] 3. This invention uses a task performance model to uniformly evaluate the improvement in positioning accuracy and classification confidence, ensuring that the fusion results have a clear performance gain in actual detection tasks and improving the overall detection capability of the system.

[0044] 4. This invention constructs an alignment-performance matching model to dynamically correlate the alignment quality at the feature level with the performance at the task level, thereby achieving effective synergy between feature optimization and task requirements and improving the overall coordination of the system.

[0045] 5. This invention can utilize an efficiency optimization model based on the alignment-performance matching coefficient, information gain, and current computational efficiency to dynamically output the target fusion computational efficiency ratio, thereby achieving an adaptive balance between detection accuracy and computational resources and improving the feasibility and efficiency of the system in actual deployment. Attached Figure Description

[0046] Figure 1 This is a schematic diagram of the process of the present invention. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0048] The specific implementation of the present invention will be described in detail below with reference to specific embodiments.

[0049] Please see Figure 1 The present invention provides a 3D target detection method based on multi-camera and millimeter-wave radar fusion, comprising the following steps:

[0050] A feature alignment model is constructed based on feature distribution similarity, spatial feature consistency, and feature spatial alignment error, and the feature alignment coefficient is output.

[0051] Information gain coefficients are constructed based on intermodal information redundancy, feature dimension effectiveness, and unique information contribution.

[0052] A task performance model is constructed based on the positioning accuracy gain and classification confidence improvement, and the task performance coefficient is output.

[0053] An alignment-performance matching model is constructed based on the feature response intensity ratio, the feature alignment coefficient under the detection quality improvement, and the task performance coefficient, and the alignment-performance matching coefficient is output.

[0054] An efficiency optimization model is constructed based on the alignment-efficiency matching coefficient, information gain coefficient, geometric consistency coefficient, and current fusion computing efficiency ratio, and the target fusion computing efficiency ratio is output.

[0055] Through the above technical solutions, this application effectively solves the problems of feature inaccuracy, information redundancy, and low computational efficiency in multimodal feature fusion. The feature alignment model eliminates distribution differences and spatial misalignment between modalities, improving the compatibility of feature fusion; the information gain model filters high-value features to reduce invalid computation, improving the effectiveness of the fusion process; the task performance model establishes a closed-loop feedback for detection performance, ensuring the reliability of system output; and the efficiency optimization model realizes dynamic allocation of computing resources, balancing real-time performance with detection accuracy requirements. This method forms a complete evaluation and optimization chain in the multimodal data fusion process, significantly improving 3D target detection performance in complex scenes.

[0056] Preferably, the steps for constructing a feature alignment model and outputting feature alignment coefficients based on feature distribution similarity (the degree of similarity between the image and radar feature probability distributions), spatial feature consistency (alignment accuracy after feature space transformation), and feature space alignment error (correlation of features at the same spatial location) are as follows:

[0057] The feature distribution similarity, spatial feature consistency, and feature space alignment error are obtained and then subjected to maximum-minimum normalization to obtain the feature distribution similarity index, feature space alignment error index, and spatial feature consistency index.

[0058] A feature alignment model is constructed based on the feature distribution similarity index, the feature space alignment error index, and the spatial feature consistency index, and the feature alignment coefficient is output. The feature alignment model is expressed as follows:

[0059]

[0060] in, Represents the feature alignment coefficient. Represents the similarity index of feature distributions. This represents the feature space alignment error index. Indicators representing spatial feature consistency indices Represents the weight coefficient and The Furthermore, the larger the value, the more consistent the image features are with the radar features;

[0061] Among them, feature distribution similarity ,in, This represents the probability distribution of image features, and this represents the probability distribution of radar features. Indicates the Jensen-Shannon divergence;

[0062] Among them, spatial feature consistency ,in, Denotes the effective total prime number. Indicates the first Image feature vectors of individual pixels Indicates the first Radar feature vectors of individual elements, This represents the Pearson correlation coefficient;

[0063] Among them, feature space alignment error , Indicates the quantity of feature pairs. Indicates the first Image feature samples, Indicates the first One radar feature sample, Represents the feature space transformation function ( ), Represents the weight matrix. (representing the bias vector) This represents the L2 norm (Euclidean distance).

[0064] Feature distribution similarity refers to the degree of similarity between the probability distributions of image and radar features. Specifically, the Jensen-Shannon divergence can be used to measure the difference between their distributions. By calculating the divergence between probability distributions and converting it into a similarity index, the matching degree of low-level features between modalities can be evaluated. Spatial feature consistency refers to the alignment accuracy after feature space transformation. Specifically, the Pearson correlation coefficient can be used to assess the linear correlation of feature vectors at the same spatial location. By calculating the average of the correlation coefficients of voxel-level feature vectors, the alignment accuracy of features in the spatial dimension can be ensured. Feature space alignment error refers to the correlation deviation of features at the same spatial location. Specifically, the L2 norm can be used to calculate the geometric deviation between the transformed radar features and image features. After mapping radar features to the image feature space using a feature space transformation function, the Euclidean distance is calculated to quantify the spatial alignment error.

[0065] Specifically, the feature alignment model eliminates dimensional differences between different indicators through normalization, transforming distribution similarity, alignment error, and consistency into comparable exponential forms. Jensen-Shannon divergence measures the difference in feature distribution between the image and radar, and its range characteristics facilitate conversion into a similarity index. The Pearson correlation coefficient assesses spatial alignment accuracy, reflecting the degree of matching in the spatial dimension through voxel-level feature vector correlation calculation. The feature space transformation function maps radar features to the image feature space and calculates geometric deviation using the L2 norm, directly quantifying the spatial alignment error. These three indices evaluate feature alignment quality from the dimensions of probability distribution, spatial correlation, and geometric error, respectively. The importance of each dimension is dynamically adjusted through weighting coefficients, allowing the fusion strategy to balance the priority of different evaluation dimensions according to scenario requirements, avoiding the bias caused by a single indicator.

[0066] Compared to existing technologies, traditional methods typically employ only a single metric to evaluate feature alignment quality, such as relying solely on geometric projection error or feature correlation. This fails to comprehensively reflect the degree of matching of multimodal features across distribution, space, and geometry. This proposed solution addresses the problem of blind fusion strategies caused by the single evaluation dimension in traditional methods by constructing a multidimensional evaluation system that comprehensively quantifies distributional differences, spatial alignment accuracy, and geometric errors. This provides a more comprehensive guide for optimizing the feature alignment process.

[0067] Through the above technical solutions, this application can systematically evaluate the matching quality of multimodal features in three dimensions: probability distribution, spatial alignment, and geometric error. By implementing the optimal feature alignment strategy in different scenarios through a dynamic weight adjustment mechanism, it effectively solves the problem of blind fusion caused by the single evaluation dimension in traditional methods, and improves the accuracy of multimodal feature fusion and the overall performance of the system.

[0068] Preferably, the steps for constructing the information gain coefficient based on intermodal information redundancy, feature dimension effectiveness, and unique information contribution are as follows:

[0069] The redundancy of intermodal information, the effectiveness of feature dimensions, and the contribution of unique information are obtained and then subjected to maximum-minimum normalization to obtain the redundancy index, effectiveness index, and contribution index.

[0070] An information gain model is constructed based on the redundancy index, effectiveness index, and contribution index to obtain the information gain coefficient. The information gain model is expressed as follows:

[0071]

[0072] in, Represents the information gain coefficient. Indicates the redundancy index. Indicating an effectiveness index, The contribution index represents the percentage of contribution. Furthermore, the larger the value, the higher the value of fusing image features and radar features (the more significant the information gain).

[0073] Among them, intermodal information redundancy , This represents the mutual information between the image and the radar. Entropy, representing image features Entropy representing radar characteristics;

[0074] Among them, feature dimension effectiveness ,in, The rank of the fused feature matrix is ​​represented by the rank of the feature matrix. The rank of the image feature matrix is ​​represented by... The rank of the radar characteristic matrix is ​​represented by... Indicates the minimum feature dimension;

[0075] Among them, unique information contribution ,in, This indicates the detection accuracy of the fusion system. This represents the detection accuracy of a purely image-based system. This indicates the detection accuracy of a pure radar system.

[0076] Specifically, this method first calculates the intermodal information redundancy by using the ratio of mutual information to entropy, reflecting the degree of overlap between image and radar features; higher redundancy indicates lower fusion value. Second, it evaluates the effectiveness of feature dimensions by comparing the rank difference of the feature matrices before and after fusion; a larger rank increase indicates stronger expressive power of the fused feature space. Further, it quantifies the contribution of unique information to the detection results by calculating the accuracy gain ratio of the fused system relative to the single-modal system. After normalizing the above three indicators, a logistic function is used to fuse the redundancy index, effectiveness index, and contribution index into a unified information gain coefficient. This coefficient dynamically reflects the value of multimodal fusion by suppressing redundant information, enhancing effective feature expression, and maximizing the contribution of unique information, thereby guiding the system to select the optimal fusion strategy.

[0077] Compared with existing technologies, traditional methods do not quantitatively evaluate information redundancy, feature dimension effectiveness, and unique information contribution, resulting in a lack of basis for selecting fusion strategies. This solution systematically evaluates the complementarity and fusion potential of multimodal data by introducing mutual information, feature matrix rank analysis, and accuracy gain calculation, thus solving the problem of inaccurate fusion value assessment.

[0078] Through the above technical solution, this application can accurately quantify the information complementarity characteristics of multimodal features, avoid resource waste caused by redundant information, enhance the effectiveness of feature expression and make full use of unique information, thereby improving the efficiency of multi-camera and millimeter-wave radar fusion detection and optimizing the resource allocation and decision accuracy of fusion strategy.

[0079] Preferably, the steps for constructing a task performance model based on positioning accuracy gain and classification confidence improvement, and then outputting task performance coefficients, are as follows:

[0080] The positioning accuracy gain and classification confidence are obtained, and the two are subjected to maximum-minimum normalization to obtain the positioning accuracy gain index and classification confidence index.

[0081] A task performance model is constructed based on the positioning accuracy gain index and the classification confidence index, and the task performance coefficients are obtained. The task performance model is expressed as follows:

[0082]

[0083] in, Indicates the task effectiveness coefficient. This represents the positioning accuracy gain index. Represents the classification confidence index, the Furthermore, the higher the value, the better the task performance;

[0084] Among them, positioning accuracy gain , This represents the expected average displacement error of the fused system. This represents the expected average displacement error of the optimal single-mode.

[0085] Among them, classification confidence , This represents the expected confidence level of the fusion system in predicting the correct category. This represents the expected confidence level of the best single modality in predicting the correct category.

[0086] Specifically, this technical solution calculates the raw value of positioning accuracy gain by collecting positioning error data from the fusion system and the optimal single-modality system. Simultaneously, it extracts the expected values ​​of prediction confidence for the correct category from both systems and calculates the raw value of classification confidence. These two raw indices are then subjected to max-min normalization to eliminate dimensional differences, resulting in standardized exponents. The two exponents are input into a task performance model for nonlinear fusion. The exponential function design in the model ensures that the task performance coefficient exhibits accelerated growth when both positioning accuracy and classification confidence improve simultaneously, effectively capturing the performance leap brought about by multi-indicator synergistic optimization. The task performance coefficient output by this model serves as a quantifiable evaluation indicator, providing a basis for the dynamic adjustment of subsequent fusion strategies. For example, when the coefficient falls below a set threshold, feature alignment optimization or computational resource reallocation can be triggered.

[0087] Compared with existing technologies, traditional methods typically optimize positioning accuracy or classification confidence separately, lacking joint modeling of the two types of task indicators, and the linear weighting method fails to accurately reflect the nonlinear coupling relationship between indicators. This solution, by constructing an exponential task performance model, not only achieves quantitative fusion of multiple task indicators but also accurately characterizes the nonlinear change law of performance as indicators improve, thus solving the problems of one-sidedness and linearity in task performance evaluation in existing technologies.

[0088] Through the above technical solution, this application realizes dynamic quantitative evaluation of task performance in a multimodal fusion system, which can accurately reflect the synergistic optimization effect of positioning accuracy and classification confidence, providing a reliable basis for real-time adjustment of fusion strategies. This solution, by establishing a correlation mechanism between task performance coefficient and fusion computation efficiency, supports the system in achieving a dynamic balance between detection accuracy and resource consumption, effectively solving the problem of blind fusion strategies caused by the lack of performance evaluation in traditional methods.

[0089] Preferably, the steps for constructing the alignment-performance matching model and outputting the alignment-performance matching coefficient based on the feature response intensity ratio, the feature alignment coefficient under the detection quality improvement, and the task performance coefficient are as follows:

[0090] The detection quality improvement degree is subjected to maximum-min normalization to obtain the detection quality improvement degree index.

[0091] Obtain the characteristic response intensity ratio, and then perform a ratio operation between the absolute difference between the characteristic response intensity ratio and the optimal characteristic response intensity ratio and the allowable deviation from the optimal characteristic response intensity ratio to obtain the characteristic response intensity ratio index;

[0092] An alignment-performance matching model is constructed based on the feature alignment coefficient and the detection quality improvement coefficient under the feature response intensity ratio index and the detection quality improvement index, as well as the task performance coefficient. The alignment-performance matching model is expressed as follows:

[0093]

[0094] in, This represents the alignment-performance matching coefficient. Represents the feature alignment coefficient. Indicates the task effectiveness coefficient. Indicates preventing decimals except zero (usually 10 -6 ), The index represents the ratio of characteristic response intensity. Indicates the improvement in detection quality, the Furthermore, the larger the value, the higher the synergy between feature alignment and task performance;

[0095] Among them, the characteristic response intensity ratio ; where represents the average L2 norm of the image features, The average L2 norm representing radar characteristics;

[0096] Among them, the improvement in testing quality , ( , The area under the Precision-Recall curve (AUC) represents the average precision of category c, and vice versa. This represents the average accuracy of the optimal single-mode system.

[0097] The detection quality improvement index is a standardized index generated by maximizing-minimum normalization to eliminate the impact of dimensional differences on the model. Specifically, it can be achieved by mapping the original detection quality improvement value to a range of 0 to 1, making the improvement comparable across different modalities or scenarios. This index serves as a weighting factor in the model, highlighting the actual improvement in detection performance. The feature response intensity ratio index dynamically assesses the alignment deviation between image and radar features in spatial distribution by calculating the deviation ratio from the optimal value. Specifically, it can be achieved by comparing the absolute difference with an allowable deviation threshold to avoid fusion failure due to excessively strong or weak features of a single modality. This index penalizes the deviation using an exponential decay function to ensure the feature response intensity remains within a reasonable range. The alignment-performance matching model is a mathematical model that combines the harmonic mean of the feature alignment coefficient and the task performance coefficient, and incorporates the detection quality improvement index and the feature response intensity ratio index for weighted calculation. Specifically, it can be achieved by using the harmonic mean to balance the relationship between feature alignment and task performance, combined with an exponential function to adjust the weight distribution. This model can adaptively adjust the fusion strategy to ensure a positive correlation between feature alignment accuracy and task performance improvement.

[0098] Specifically, the detection quality improvement index eliminates dimensional differences across different scenarios through normalization, enabling the model to evaluate the improvement in fusion performance across scenarios. The feature response intensity ratio index quantifies the alignment deviation between image and radar features in spatial distribution by dynamically calculating the deviation ratio, avoiding the dominance of a single modality feature in the fusion process. The alignment-performance matching model uses a harmonic mean to balance the relationship between feature alignment coefficients and task performance coefficients, and combines an exponential decay function to penalize feature response deviations. Simultaneously, the detection quality improvement index is used as a weighting factor, ensuring that the matching coefficients reflect both the balance between feature alignment and task performance while highlighting the actual improvement in detection performance. Therefore, this model can dynamically adjust the fusion strategy based on real-time feature response status and changes in detection performance, achieving synergistic optimization of feature alignment quality and task performance.

[0099] Compared to existing technologies, traditional methods typically evaluate feature alignment or task performance metrics independently, lacking quantitative modeling of their synergistic relationship. This leads to lags and blind adjustments in fusion strategies. Our proposed solution, however, constructs a matching model that incorporates feature response intensity deviation and detection quality improvement into a unified computational framework. This enables real-time, interconnected evaluation of feature alignment accuracy and task performance improvement, resolving the instability in fusion results caused by the separation of these two aspects in traditional methods.

[0100] Through the above technical solution, this application can dynamically balance the spatial alignment accuracy of multimodal features with the requirements for improving task performance, and automatically optimize the fusion strategy based on the real-time feature response status, effectively improving the synergy between feature alignment quality and task efficiency. This solution solves the problem of blind strategy adjustment caused by the disconnect between feature alignment and task efficiency in traditional fusion methods, realizes closed-loop optimization of the fusion process, and enhances the adaptability of the multimodal fusion system in complex scenarios.

[0101] Preferably, the efficiency optimization model is expressed as:

[0102]

[0103] in, This indicates the target fusion computational efficiency ratio. This indicates the current efficiency ratio of fused computing. This represents the alignment-performance matching coefficient. Represents the information gain coefficient. Indicates the performance threshold;

[0104] Among them, the efficiency ratio of fusion computing , This represents the total computation time of the fusion process. Indicates the image branching time. This indicates the radar branch processing time.

[0105] Among them, the target fusion computational efficiency ratio refers to the proportion of computational resources that the fusion system needs to achieve. Specifically, it can be achieved by real-time monitoring of the time consumption of the fusion process and single-modal processing. Its function is to provide a quantitative target for dynamically adjusting computational resources. The current fusion computational efficiency ratio refers to the ratio of the current actual computation time to the single-modal processing time. Specifically, it can be achieved by using a timestamp recording module to time and count each processing stage. Its function is to provide a benchmark reference for efficiency optimization. The alignment-performance matching coefficient refers to the synergy between feature alignment quality and task performance improvement. Specifically, it can be calculated by the product of the normalized feature response intensity ratio and the detection quality improvement. Its function is to provide a basis for evaluating the contribution of fusion quality to task performance. The information gain coefficient refers to the information complementarity value brought by multimodal feature fusion. Specifically, it can be calculated by combining mutual information, feature matrix rank, and accuracy improvement. Its function is to provide a quantitative standard for measuring the information effectiveness of the fusion process. The performance threshold refers to the critical condition parameter that triggers the adjustment of computational resources. Specifically, it can be determined by historical data statistics or experimental calibration. Its function is to set a stability boundary to prevent frequent fluctuations in computational efficiency.

[0106] Specifically, this technical solution dynamically adjusts the fusion computation efficiency ratio through a nonlinear function. When the product of the alignment-performance matching coefficient and the information gain coefficient exceeds a performance threshold, the hyperbolic tangent function outputs a positive value, and the system proportionally increases the investment in fusion computation resources to improve detection accuracy; conversely, it reduces resource allocation to lower computational overhead. The calculation of the fusion computation efficiency ratio correlates the fusion time with the single-modal processing time, quantifying the difference in computational load between the fusion process and independent processing through a ratio. For example, when the fusion time accounts for too high a proportion, the system can automatically reduce the computational complexity of the feature alignment module and optimize resource allocation.

[0107] Compared to existing technologies, traditional methods typically employ fixed resource allocation strategies, failing to dynamically adjust computational efficiency based on the synergistic effects of multimodal data. This proposed solution, however, introduces multi-dimensional metrics such as feature alignment quality and information gain to construct a quantifiable dynamic adjustment mechanism, thus resolving the resource waste or performance degradation issues caused by static strategies. For example, in scenarios with fluctuating sensor data quality, this solution can automatically suppress the fusion computation of low-information-gain features, avoiding ineffective resource consumption.

[0108] Through the above technical solution, this application achieves a dynamic balance between computational resources and task performance during multimodal fusion, solving the problem of low resource utilization caused by blind fusion in traditional methods. Furthermore, this solution provides an interpretable decision-making basis for the online optimization of the fusion system by quantifying the correlation between feature alignment quality and task performance.

[0109] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.

[0110] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A 3D object detection method based on multi-camera and millimeter wave radar fusion, characterized in that, Further comprising the following steps: constructing a feature alignment model based on the feature distribution similarity, the spatial feature consistency, and the feature space alignment error to output a feature alignment coefficient; constructing an information gain coefficient based on the inter-modal information redundancy, the feature dimension effectiveness, and the unique information contribution degree; constructing a task performance model based on the positioning accuracy gain and the classification confidence degree to output a task performance coefficient; constructing an alignment-performance matching model based on the feature response intensity ratio, the feature alignment coefficient, and the task performance coefficient under the detection quality improvement degree to output an alignment-performance matching coefficient; constructing an efficiency optimization model based on the alignment-performance matching coefficient, the information gain coefficient, and the current fusion calculation efficiency ratio to output a target fusion calculation efficiency ratio.

2. The 3D target detection method based on multi-camera and millimeter wave radar fusion according to claim 1, characterized in that, The efficiency optimization model is represented as: ; wherein, represents a target fusion computing efficiency ratio, represents a current fusion computing efficiency ratio, represents an alignment-performance matching coefficient, represents an information gain coefficient, represents a performance threshold.

3. The 3D target detection method based on multi-camera and millimeter wave radar fusion according to claim 2, characterized in that, The step of constructing an alignment-performance matching model based on the feature response intensity ratio, the feature alignment coefficient, and the task performance coefficient under the detection quality improvement degree to output an alignment-performance matching coefficient is: performing maximum-minimum normalization processing on the detection quality improvement degree to obtain a detection quality improvement degree index; obtaining the feature response intensity ratio, performing ratio processing on the absolute difference between the feature response intensity ratio and the optimal feature response intensity ratio and the allowed deviation from the optimal feature response intensity ratio value to obtain a feature response intensity ratio index; constructing an alignment-performance matching model based on the feature response intensity ratio index and the feature alignment coefficient and the task performance coefficient under the detection quality improvement degree to obtain an alignment-performance matching coefficient, the alignment-performance matching model being represented as: ; wherein, represents an alignment-performance matching coefficient, represents a feature alignment coefficient, represents a task performance coefficient, represents a prevention zero decimal, represents a feature response intensity ratio index, represents a detection quality improvement degree, and the and the greater the value, the higher the degree of synergy between the feature alignment and the task performance.

4. The 3D target detection method based on multi-camera and millimeter wave radar fusion according to claim 3, characterized in that, The step of constructing a task performance model based on the positioning accuracy gain and the classification confidence degree to output a task performance coefficient is: obtaining the positioning accuracy gain and the classification confidence degree and performing maximum-minimum normalization processing thereon to obtain a positioning accuracy gain index and a classification confidence degree index; constructing a task performance model based on the positioning accuracy gain index and the classification confidence degree index to obtain a task performance coefficient, the task performance model being represented as: ; wherein, represents a task performance coefficient, represents a positioning accuracy gain index, represents a classification confidence index, said and the greater the value the better the task performance.

5. The multi-camera and millimeter wave radar fusion based 3D target detection method of claim 3, wherein, The step of constructing an information gain coefficient based on the inter-modal information redundancy, the feature dimension effectiveness, and the unique information contribution degree is: obtaining the inter-modal information redundancy, the feature dimension effectiveness, and the unique information contribution degree and performing maximum-minimum normalization processing thereon to obtain a redundancy index, an effectiveness index, and a contribution degree index; constructing an information gain model based on the redundancy index, the effectiveness index, and the contribution degree index to obtain an information gain coefficient, the information gain model being represented as: ; wherein, represents an information gain coefficient, represents a redundancy index, represents an effectiveness index, represents a contribution index, the and the greater the value, the higher the value of the fusion of the image features and the radar features.

6. The multi-camera and millimeter wave radar fusion based 3D target detection method of claim 3, wherein, The step of constructing a feature alignment model based on the feature distribution similarity, the spatial feature consistency, and the feature space alignment error to output a feature alignment coefficient is: obtaining the feature distribution similarity, the spatial feature consistency, and the feature space alignment error and performing maximum-minimum normalization processing thereon to obtain a feature distribution similarity index, a feature space alignment error index, and a spatial feature consistency index; constructing a feature alignment model based on the feature distribution similarity index, the feature space alignment error index, and the spatial feature consistency index to output a feature alignment coefficient, the feature alignment model being represented as: ; wherein, represents a feature alignment coefficient, represents a feature distribution similarity index, represents a feature space alignment error index, represents a spatial feature consistency index, represents a weight coefficient and , the and the greater the value the more consistent the image features are with the radar features.

7. The multi-camera and millimeter wave radar fusion based 3D target detection method of claim 2, wherein, fusion computation efficiency ratio , denotes the total computation time of the fusion process, denotes the image branch processing time, denotes the radar branch processing time.

Citation Information

Patent Citations

  • Target tracking method and system based on radar data and video data fusion

    CN117949942A

  • Laser cutting device and control method thereof

    CN119216838A