Multi-modal target detection method based on feature enhancement and alignment fusion
Through the multimodal object detection method of feature enhancement and alignment fusion, the image alignment problem is solved, the feature expression and fusion capabilities are enhanced, and the detection performance and robustness are improved.
Patent Information
- Application Number
- CN202510430525.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, in visible light-infrared image object detection, image alignment is difficult to perfectly achieve, resulting in feature conflicts and information loss, and the cross-modal complementary information is not fully utilized, affecting detection performance, especially under complex backgrounds and lighting conditions.
The feature enhancement module (FEM), feature alignment module (FAM), feature fusion module (FFM) and cross-scale feature fusion module (SFFM) are used to improve detection performance by enhancing feature expression capabilities, correcting spatial dislocations between modes and optimizing feature fusion.
It effectively improves the object detection performance in the case of visible-infrared image misalignment, achieving better model performance and robustness.
Smart Images

Figure CN120339638A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-modal object detection method based on feature enhancement and alignment fusion, specifically to the problems of feature enhancement, alignment and fusion of images, and belongs to the field of object detection in computer vision. Background Art
[0002] Object detection is one of the core tasks in the field of computer vision, and its purpose is to accurately identify and locate target objects of interest in images or videos. With the development of technology, multi-modal object detection has gradually become a research hotspot, and visible-light-infrared image fusion object detection has attracted much attention due to its unique advantage of being able to utilize the image information of two modalities. Visible-light images have rich texture and color information, while infrared images can capture the thermal radiation characteristics of objects. Combining the advantages of both can effectively improve the robustness and accuracy of the object detection model.
[0003] In existing visible-light-infrared image object detection methods, many studies focus on how to effectively extract and fuse the features of the two modalities. However, most of these methods assume that the input images have been precisely aligned, but in actual application scenarios, it is often difficult to achieve perfect image alignment. When there is a spatial misalignment between the visible-light image and the infrared image, direct fusion will lead to feature conflicts and information loss, thus reducing the detection performance. In addition, many current methods may not fully utilize the complementary information between cross-modalities during the feature extraction process, which may limit the expressive ability of the fused features, and thus the effect is not ideal when dealing with object detection tasks under complex backgrounds and lighting conditions.
[0004] To address the above problems, the present invention proposes a multi-modal object detection method based on feature enhancement and alignment fusion, aiming to improve the object detection performance of the model in the case of misalignment between visible-light and infrared image pairs by enhancing the feature expression ability, correcting the spatial misalignment between modalities, and optimizing the feature fusion strategy. Summary of the Invention
[0005] The present invention proposes a multi-modal object detection method based on feature enhancement and alignment fusion, which mainly consists of a feature extraction backbone network, a Feature Enhancement Module (FEM), a Feature Alignment Module (FAM), a Feature Fusion Module (FFM), and a Cross-scale Feature Fusion Module (SFFM). Specifically, the feature enhancement module enhances the feature representations of the two modalities by combining the features of the two modalities at each stage of the backbone network for feature extraction; the feature alignment module is used to align the feature maps from different modalities at each scale globally and locally, laying a good foundation for feature fusion; the feature fusion module is designed to overcome the inherent differences between feature maps from different modalities and has a good effect of fusing features from different modalities; the cross-scale feature fusion module fuses feature maps of different scales and can integrate information from features of different scales. The careful design of each of the above modules enables the present invention to achieve excellent performance in the visible light-infrared image fusion object detection task.
[0006] A multi-modal object detection method based on feature enhancement and alignment fusion includes the following steps:
[0007] (1) Use the backbone network to extract the features of the visible light image and the infrared image. At multiple stages of the backbone network for image feature extraction, use FEM to better guide the backbone network to extract relevant features from the two modalities and enhance the feature representation;
[0008] (2) Use FAM to perform global-to-local alignment on the feature maps of the two modalities after feature enhancement;
[0009] (3) Use FFM to fuse the aligned feature maps from the two modalities;
[0010] (4) After shallow multi-modal feature fusion, use SFFM to further fuse the feature maps after deeper cross-modal fusion in (3);
[0011] (5) Input the fused feature maps at different scales into the detection head to obtain the results of object detection.
[0012] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0013] 1. The present invention designs a feature enhancement module based on cross-modal cross-attention, which enhances the feature expression ability during the feature extraction process;
[0014] 2. The present invention designs a feature alignment module based on global-to-local alignment, which can effectively align feature maps from two modalities, laying a solid foundation for the subsequent feature fusion process;
[0015] 3. The present invention constructs a frequency-aware cross-modal and multi-scale feature fusion module, which effectively alleviates the differences between multi-modal and multi-scale features, enhances the effectiveness of feature fusion, and thus achieves better model performance in the detection task. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is the overall model block diagram of the multi-modal object detection method based on feature enhancement and alignment fusion according to the present invention;
[0017] Figure 2 It is the network structure diagram of the feature enhancement module based on cross-modal cross-attention according to the present invention;
[0018] Figure 3 It is the structure diagram of the feature map alignment module based on global-to-local alignment according to the present invention;
[0019] Figure 4 It is the network structure diagram of the frequency-aware cross-modal feature fusion according to the present invention;
[0020] Figure 5 It is the network structure diagram of the cross-scale feature fusion according to the present invention.
[0021] Figure 6 It is the comparison diagram of the visual effects of the method according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0022] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0023] As Figure 1 shown, a multi-modal object detection method based on feature enhancement and alignment fusion mainly includes five parts: a feature extraction backbone network, a feature enhancement module (Feature Enhancement Module, FEM), a feature alignment module (Feature Alignment Module, FAM), a feature fusion module (Feature Fusion Module, FFM) and a cross-scale feature fusion module (Cross-scale Fusion Module, SFFM). 1. Overall network framework and processing process:
[0024] For the input visible light image and infrared image, first use a dual-branch backbone network to extract features at the Level1 level respectively. Input the two types of features into the FEM module for feature interaction and enhancement, and introduce them into the backbone network through residuals. Then use the backbone network to continue extracting deeper-level (Level2 and Level3) features. Similar to the features at the Level1 level, use FEM to enhance the features at the Level2 and Level3 levels. Then, use FAM to align the features of the two modalities enhanced by the FEM module globally and locally. Finally, fuse the aligned features of the two modalities using FFM. Among them, the features of the two modalities at the Level3 level are directly fused through FFM and sent to the detection head. Before the multi-modal fusion features at the Level2 and Level1 levels are input to the detection head, they are secondarily fused with the deeper-level fusion features through SFFM. Finally, the final fusion features from each level are sent to the FPN layer and the detection head to obtain the final detection result.
[0025] 2. Feature Enhancement Module (FEM) Based on Cross-Modal Cross-Attention
[0026] As Figure 2 shown, features from the two modalities are first pooled at different levels to generate feature maps with spatial resolutions of 1×1, 3×3, 6×6, and 8×8. Subsequently, these feature maps are flattened and input into a cross-attention module for further processing. In the cross-attention module, the feature maps of the two modalities are concatenated together, and then position embedding and layer normalization are performed. They are input into the fully connected layer (i.e., multiplied by the W q , W k and W v matrices) to generate query (Q rgb and Q ir ), key (K rgb and K ir ), and value (V rgb and V ir ) matrices for the visible light and infrared modalities. Then, calculate the cross-attention scores between the two modalities, namely Attention rgb and Attention ir , and the calculation formula is as follows:
[0027]
[0028] where d is the dimension of the K matrix. Finally, multiply Attention vis and Attention irInput the Multilayer Perceptron (MLP) and introduce residuals to obtain feature maps of size 8×8 for both modalities. These feature maps are upsampled to ensure that their sizes are consistent with the size of the feature maps input to the FEM.
[0029] 3. Feature Map Alignment Module (FAM) Based on Global-to-Local Alignment
[0030] The feature map alignment module based on global-to-local alignment can be divided into two steps: global alignment and local alignment. As Figure 3 shown in the left part of , the global alignment module is based on the affine transformation model and uses a parameter estimation network to estimate the position transformation parameters between the two modalities. First, subtract the feature maps of the two modalities. The feature map obtained through the subtraction operation can highlight the difference information between the two modalities. Then, use three average pooling and convolutional layers to predict the translation, scaling, and rotation transformation parameters respectively according to the subtracted feature map. In the method of the present invention, the infrared modality feature is used as the reference feature, and according to the parameters learned by the network, an affine transformation is applied to the visible light feature to obtain the visible light feature globally aligned with respect to the infrared feature. This process can be represented in the following matrix form:
[0031]
[0032] where (x', y') are the coordinates of the aligned visible light feature map, and (x, y) are the coordinates of the unaligned visible light feature map; a, b, c, d, tx, and ty are the learned affine transformation factors. Through the affine transformation, the visible light and infrared feature maps can be globally aligned in three dimensions of translation, scaling, and rotation.
[0033] As Figure 3 shown in the right part of , deformable convolution is further used to achieve the local alignment of the visible light feature map after global alignment. Taking the infrared feature as the reference, first subtract the globally aligned visible feature and the infrared feature, and then use a convolutional layer to learn the local offset. Then, based on this offset, correct and locally align the globally aligned visible feature. This process can be expressed as:
[0034]
[0035] where Y represents the result after local alignment, X represents the globally aligned visible light feature; n and N represent the index and total number of convolutional kernel weights respectively. w n represents the weight of the nth convolutional kernel; p, p n and Δp N are the center index, the fixed offset of the nth convolutional kernel, and the learnable offset respectively. In the method of the present invention, p nLimited within {(-1,1),(-1,0),...,(1,1)}.
[0036] 4. Frequency-Aware Feature Fusion Module (FFM)
[0037] As Figure 4 shown, the frequency-aware fusion module predicts high-pass and low-pass filter kernels from the input visible light and infrared feature maps, applies high-pass and low-pass filters to the two types of feature maps respectively, and then obtains the fused features by adding the two filtered feature maps. To achieve more effective fusion, the prediction of the filter kernels and the filtering process are performed twice. This process first compresses the channels of the two feature maps through 1×1 convolution. The two compressed feature maps are then concatenated and passed through two convolutional layers to learn the low-pass and high-pass filter kernels respectively. These filter kernels are used to filter the compressed feature maps. Taking the low-pass filtering as an example, the filtering process can be expressed as:
[0038]
[0039] V i,j = conv(concat(X vis ,X ir ))
[0040]
[0041] where X and Y are the input and output feature maps respectively; the indices i and j correspond to the height and width dimensions of the feature map respectively; S is the surrounding area of (i,j); W is the weighted weight of the filter; V i,j is the filter kernel of size L×L or H×H generated by the convolutional layer. The softmax function in Equation (6) can ensure that the processing of the feature map according to Equation (4) has a low-pass effect. Subtracting the low-pass filtering result of the feature map itself and introducing the residual can achieve the high-pass filtering of the feature map, and this process can be expressed as:
[0042]
[0043] where, X high_pass , X low_pass are the high-pass filtering result and the low-pass filtering result of the feature map X respectively. Then, the preliminarily denoised and enhanced feature maps are concatenated and input into the convolutional layer to predict the filter kernel for filtering the input feature map, and the process is similar to the processing of the compressed feature map. Finally, the fused features are obtained by adding the two filtered feature maps.
[0044] 5. Frequency-Aware Multi-Scale Feature Fusion Module (SFFM)
[0045] As Figure 5As shown, adding an upsampling module to the frequency-aware fusion module can be used for frequency-aware multi-scale feature fusion. Different from the frequency-aware fusion module, the input of SFFM is the features after multi-modal feature fusion at different scales. Therefore, the sizes of the two feature maps are different. During processing, the features from the deeper layer are upsampled and then concatenated with the features from the lower layer. Besides, there is no other difference between SFFM and the frequency-aware fusion module, so it will not be elaborated here.
[0046] 6. Loss Function
[0047] The method of the present invention uses the localization loss, classification loss, and confidence loss of the prediction box as the total loss, and the formula is as follows:
[0048] L total = L reg + L cls + L conf
[0049] Where L reg 、L cls 、L conf represent the regression loss, classification loss, and confidence loss of the prediction box respectively.
[0050] First of all, the method of the present invention uses the Generalized Intersection over Union (GIoU) loss
[21] as the localization loss. Different from the IoU loss that only focuses on the overlapping area, the GIoU loss not only considers the overlapping area but also pays attention to other non-overlapping areas, thus more comprehensively reflecting the overlapping degree between two boxes. The calculation formula of the localization loss is as follows:
[0051]
[0052] Where, S 2 represents the number of predicted image grids, M represents the number of prediction boxes in each grid, indicates whether the jth prediction box in the ith grid is a positive sample, represents the prediction box, represents the ground truth box, represents by and forms the area of the smallest bounding box.
[0053] The classification loss uses the cross-entropy loss to optimize the classification problem of the prediction box, and the calculation formula is as follows:
[0054]
[0055] Where, and p i(c) represents the probability and the true probability that the network prediction sample belongs to class c.
[0056] The confidence loss is used to measure the difference between the prediction confidence of the model for the detected object and the true label. The mean squared error loss is used to calculate the distance between the predicted value and the true value of the confidence score for each prediction box. The calculation formula is as follows:
[0057]
[0058] where c i and represent the true value and the value predicted by the network of the confidence respectively, indicates whether the j-th prediction box in the i-th grid is a negative sample. Through such a design, the loss function can effectively balance the localization accuracy, the object detection confidence, and the classification accuracy, thereby improving the accuracy and robustness of the model.
[0059] To verify the effectiveness of the multi-modal object detection method based on feature enhancement and alignment fusion described in the present invention, a detailed comparison will be made through experiments below.
[0060] The experimental environment is as follows: the operating system is Ubuntu20.04, the deep learning framework is PyTorch 1.10, and the Python version is 3.7. The present invention compares the mAP50 results of the model on the mainstream visible light-infrared image fusion object detection dataset DVTOD (the larger this value, the better the method effect), and visualizes some detection results. The mainstream object detection methods in recent years are selected for comparison, specifically:
[0061] YOLOv5: The method proposed by Jocher et al., reference "Jocher G, Stoken A, Borovec J. et al. “ultralytics / yolov5: v3.1 - Bug Fixes and Performance Improvements,” Oct. 2020. [Online]. Available: https: / / doi.org / 10.5281 / zenodo.4154370."
[0062] YOLOv11: The method proposed by Khanam et al., reference "Khanam R, Hussain M. Yolov11: An overview of the key architectural enhancements [EB / OL]. arXiv preprint arXiv:2410.17725, 2024."
[0063] CFT: The method proposed by Fang et al., reference "Fang Q Y, Han D P, Wang Z K. Cross-modality fusion transformer for multispectral object detection[EB / OL]. arXiv preprint arXiv:2111.00273, 2021."
[0064] CMA-Det: The method proposed by Song et al., reference "Song K, Xue X, Wen H, et al. Misaligned visible-thermal object detection: a drone-based benchmark and baseline[J]. IEEE Transactions on Intelligent Vehicles, 2024."
[0065] The test results are shown in Table 1. By comparing with other methods, it can be found that the method of the present invention has better detection performance. At the same time Figure 6 it also shows the objective visual effect comparison chart of the present invention on the DVTOD dataset. The columns from left to right are the visible light image input, infrared image input, YOLOv5, YOLOv11, CMA-Det, and the visualization results of the method of the present invention. It can be found that the method of the present invention can more effectively detect the objects located at the image edge and smaller objects and frame them with appropriate bounding boxes.
[0066] Table 1 Comparison of detection effects on the DVTOD dataset
[0067]
[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-modal object detection method based on feature enhancement and alignment fusion, characterized in that Including the following steps: (1) When using the backbone network to extract the features of visible light images and infrared images, in multiple stages of the backbone network for extracting image features, a Feature Enhancement Module (FEM) is used to guide the backbone network to extract relevant features from the two modalities and enhance the feature representation; (2) Using a Feature Alignment Module (FAM) to perform global-to-local alignment on the feature maps of the two modalities after feature enhancement; (3) Using a Feature Fusion Module (FFM) to fuse the feature maps from the two modalities after alignment; (4) After shallow multi-modal feature fusion, using a Cross-scale Feature Fusion Module (SFFM) to further fuse the feature maps from the deeper cross-modal fusion in (3); (5) Inputting the fused feature maps at different scales into the detection head to obtain the results of object detection.
2. The method according to claim 1, wherein In step (1), when using the backbone network to extract features, the features from the two modalities are first pooled at different levels to generate feature maps with spatial resolutions of 1×1, 3×3, 6×6, and 8×8; subsequently, these feature maps are flattened and input into a cross-attention module for further processing; In the cross-attention module, the feature maps of the two modalities are concatenated together, followed by position embedding and layer normalization, and then input into the fully connected layer (i.e., multiplied by the W q , W k and W v matrices) to generate the query (Q rgb and Q ir ), key (K rgb and K ir ), and value (V rgb and V ir ) matrices for the visible and infrared modalities; then, the cross-attention scores between the two modalities are calculated, i.e., Attention rgb and Attention ir , and the calculation formula is as follows: where d is the dimension of the K matrix; finally, Attention vis and Attention ir are input into a Multilayer Perceptron (MLP) and residuals are introduced to obtain two modality feature maps of size 8×8. These feature maps are upsampled to ensure that their sizes are consistent with the size of the feature maps input into the FEM.
3. The method according to claim 1, characterized in that, The global alignment module in step (2) is based on an affine transformation model, and a parameter estimation network is used to estimate the position transformation parameters between the two modalities; first, the feature maps of the two modalities are subtracted, and the feature map obtained through the subtraction operation can highlight the difference information between the two modalities; then, three average pooling and convolutional layers are used to predict the translation, scaling, and rotation transformation parameters respectively according to the subtracted feature map; in the present invention, the infrared modality features are used as reference features, and according to the parameters learned by the network, an affine transformation is applied to the visible light features to obtain visible light features globally aligned with respect to the infrared features; This process can be represented in the following matrix form: where (x', y') are the coordinates of the aligned visible light feature map, and (x, y) are the coordinates of the unaligned visible light feature map; a, b, c, d, tx, and ty are the learned affine transformation factors; through the affine transformation, the visible light and infrared feature maps can be globally aligned in three dimensions of translation, scaling, and rotation; Deformable convolution is further used to achieve local alignment of the globally aligned visible light feature map; taking the infrared features as a reference, first subtract the globally aligned visible features and infrared features, and then use a convolutional layer to learn the local offsets; then, based on this offset, correct and locally align the globally aligned visible features; this process can be expressed as: where Y represents the result after local alignment, and X represents the visible light features after global alignment; n and N respectively represent the index and total number of the convolutional kernel weights; w n represents the weight of the n-th convolutional kernel; p, p n and Δp N are the central index, the fixed offset, and the learnable offset of the nth convolutional kernel, respectively; in the present invention, p n is limited within the range of {(-1, 1), (-1, 0),..., (1, 1)}.
4. The method according to claim 1, characterized in that, In step (3), high-pass and low-pass filter kernels are predicted from the input visible light and infrared feature maps, and high-pass and low-pass filtering are respectively applied to the two types of feature maps, and then the fused features are obtained by adding the two filtered feature maps; to achieve more effective fusion, the prediction of the filter kernels and the filtering process are performed twice; this process first compresses the channels of the two feature maps through 1×1 convolution; the two compressed feature maps are then concatenated and passed through two convolutional layers to respectively learn the low-pass and high-pass filter kernels; These filter kernels are used to filter the compressed feature maps; Taking low-pass filtering as an example, the filtering process can be expressed as: V i,j = conv(concat(X vis , X ir )) where X and Y are the feature maps of the input and output respectively; the indices i and j correspond to the height and width dimensions of the feature map respectively; S is the surrounding area of (i, j); W is the weighted weight of the filter; V i,j is a filter kernel of size L×L or H×H generated by the convolutional layer; the softmax function in Equation (6) can ensure that the processing of the feature map according to Equation (4) has a low-pass effect; subtracting the low-pass filtering result of the feature map itself and introducing the residual can achieve high-pass filtering of the feature map, and this process can be expressed as: Among them, X high_pass and X low_pass are respectively the high-pass filtering result and the low-pass filtering result of the feature map X; then, the preliminarily denoised and enhanced feature maps are concatenated and input into a convolutional layer to predict the filter kernel for filtering the input feature map, and the process is similar to the processing of the compressed feature map; Finally, the fused features are obtained by adding the two filtered feature maps.
5. The method according to claim 1, wherein In step (4), an upsampling module is added to the frequency-aware fusion module for frequency-aware multi-scale feature fusion, and a cross-scale feature fusion module is used to perform secondary fusion on the features after shallow multi-modal fusion and deeper multi-modal features to obtain the features fed into the detection head; the input of the cross-scale feature fusion module is the features after multi-modal feature fusion at different scales. When concatenating the feature maps twice and finally adding the two feature maps to obtain the final output, upsampling is performed on the features from the deeper layer to facilitate concatenation with the features from the lower layer.
Citation Information
Cited By
Intelligent ampullaria gigas egg detection method based on image reconstruction and multi-modal fusion
CN120510156A
An intelligent detection method for golden apple snail eggs based on image reconstruction and multimodal fusion
CN120510156B
Infrared and visible light image fusion method and system combined with self-supervised feature alignment
CN121213374A
Visible light infrared image target detection method and system based on modal common characteristics
CN121437515A
Flame detection method suitable for low-quality double-spectrum image and equipment thereof
CN121937884A