Cross-space-frequency domain multi-modal remote sensing small target detection method
By designing a multimodal remote sensing small object detection method across space-frequency domains, using the YOLOv8n dual-flow backbone network and the cross-domain gated self-attention fusion module, the problems of high computational volume and weak detection capabilities of multimodal fusion detection in remote sensing images are solved, and high-precision and low-computation remote sensing small object detection are realized.
Patent Information
- Application Number
- CN202510734301.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The existing multimodal fusion detection methods do not fully utilize multimodal complementary information in remote sensing images, resulting in weaker small object detection capabilities, and high computing volumes are difficult to deploy on small drone platforms and satellite platforms.
A multimodal remote sensing small object detection method across space-frequency domain is designed, and the YOLOv8n dual-flow backbone network and a cross-domain gated self-attention fusion module (CGSA) are used to realize the adaptive fusion of multimodal complementary features through cross-space-frequency domain differential feature extraction, self-attention feature fusion and gating units.
It improves the accuracy and robustness of remote sensing small object detection, reduces the calculation amount, is suitable for satellite-based and airborne terminal deployment, and improves image processing accuracy and response speed.
Smart Images

Figure CN120279004A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing image detection and recognition, and particularly to a multi-modal remote sensing small target detection method across the spatial-frequency domain. Background Art
[0002] With the popularization of high-resolution satellites and drones, the target detection of remote sensing images has received extensive attention in various fields. Compared with natural image detection, the size of remote sensing images is much larger than that of natural images, and the target of interest only occupies a very small number of pixels, which requires a high small target detection ability of the model, and the model needs to have the ability of fast inference for large-size images. In addition, the target categories are diverse and the sizes are different, which requires the model to have the detection ability for targets of different scales. At the same time, remote sensing images usually contain complex backgrounds, such as vegetation, buildings, etc., which are easy to interfere with the target detection task, making it difficult for the algorithm to accurately identify and locate the target.
[0003] With the remarkable success of deep learning technology in the field of natural images, many studies have successfully applied deep learning algorithms to the remote sensing field. Multi-modal fusion detection has been proven to be an effective means to improve the perception ability of remote sensing small targets. However, the existing multi-modal fusion detection methods still face the following challenges: 1) The multi-modal complementary information is not fully utilized, and the detection ability for complex small targets in remote sensing images is weak; 2) The high recognition performance of multi-modal models often comes at the cost of high computational complexity, and it is difficult to deploy large-scale intelligent models on small unmanned aerial vehicle platforms and satellite platforms.
[0004] To solve the above problems, the present invention proposes an extremely lightweight multi-modal remote sensing small target detection method across the spatial-frequency domain. Summary of the Invention
[0005] To solve the above problems, the present application proposes a multi-modal remote sensing small target detection method across the spatial-frequency domain. First, the present application designs a two-stream backbone network based on YOLOv8n as the baseline model of the present application. Then, the present application designs a cross-domain gated self-attention fusion module (CGSA), which realizes the adaptive fusion of multi-modal complementary features across the spatial-frequency domain through cross-spatial-frequency domain differential feature extraction, self-attention feature fusion, and a gated unit. The specific content is as follows: A multi-modal remote sensing small target detection method across the spatial-frequency domain includes the following steps: S1. Obtain a remote sensing image, and preprocess the remote sensing image to obtain an initial image; S2. Process the initial image by using an improved multi-modal remote sensing small target detection model to obtain the target detection result of the remote sensing image; The improved multi-modal remote sensing small target detection model includes the YOLOv8n dual-stream backbone network, the improved cross-domain gated self-attention fusion module, the lightweight self-attention mechanism, and the YOLOv8n detection head.
[0006] Preferably, the specific content of using the improved multi-modal remote sensing small target detection model to process the initial image to obtain the target detection result of the remote sensing image in S2 is as follows: S201. Divide the initial image to obtain a visible light image and an infrared image; S202. Use the YOLOv8n dual-stream backbone network to perform multi-scale feature fusion on the visible light image and the infrared image respectively to obtain image features, where the image features include visible light features and infrared features ; S203. Use the improved cross-domain gated self-attention fusion module to integrate the frequency domain information and the spatial domain information of the feature image to obtain the information to be detected; S204. Use the YOLOv8n detection head to detect the information to be detected to obtain the target detection result, where the target detection result includes the target position and the target category.
[0007] Preferably, the specific content of the improved cross-domain gated self-attention fusion module CGSA in S203 includes: Cross-domain differential feature extraction CFE, which is used to promote the learning of complementary features in the frequency domain and the spatial domain during the process of integrating the frequency domain information and the spatial domain information; Self-attention feature fusion SAFF, which is used to guide the fusion of visible light and infrared features in the feature image and capture long-range dependencies; Adaptive gating AG unit, which is used to dynamically allocate the fusion weights of visible light and infrared features to achieve more sensitive and adaptive feature cross-fusion.
[0008] Preferably, the running steps of the cross-domain differential feature extraction CFE are as follows: Visible light features and infrared features can obtain the visible light frequency feature map and the infrared frequency feature map through FFT, and can be decomposed into the amplitude component and the phase component of visible light, and the phase component of infrared, which are respectively expressed as: ; ; Concatenate the amplitude components and phase components of the visible light modality and the infrared modality, and perform a feature enhancement operation to generate enhanced global frequency features. The amplitude component is , and the phase component is , and . The expressions of are respectively: ; Among them, Cat[.] represents concatenation at the channel dimension, and are bottleneck layers, including a 1×1 convolution for feature dimensionality reduction, a ReLU activation function, and a 1×1 convolution for feature dimensionality expansion; Apply the inverse FFT to convert the enhanced global frequency features back to the spatial domain to obtain the reconstructed feature map affected by the frequency domain ; ; represents the inverse FFT, represents the reconstructed feature; Use the reconstructed feature to subtract the original spatial domain features, namely the visible light feature and the infrared feature respectively, to obtain the cross-domain differential features of each modality. The cross-domain differential features of each modality include the visible light differential feature and the infrared differential feature , which are respectively expressed as: ; ; Among them is the visible light differential feature, is the infrared differential feature, and are 1×1 convolution operations.
[0009] Preferably, the running steps of the self-attention feature fusion SAFF are as follows: The cross-domain differential features are respectively generated in two branches Q and V composed of a 1×1 convolution and a Reshape operation and . On the Q branch, the features are fully folded, and on the V branch, the features remain at high resolution; Combine the Q branch results of the visible light modality and the infrared modality, and perform a Softmax operation to obtain the fused weight distribution, and multiply the fused weight distribution by the information-rich V branch features of the visible light modality and the infrared modality; The dual-channel attention weights are obtained successively through a 1×1 convolutional layer, a Layer Normalization layer, and a Sigmoid function. and ; Channel-wise multiplication is performed between the attention weights and the input features to generate a feature representation with less noise. The visible light output features are and the infrared output features are , to suppress redundant non-critical information therein while enhancing critical information that affects detection.
[0010] After filtering the feature representation, it is re-enhanced through cross-addition of the two modalities to obtain the final features. After channel concatenation and 1×1 channel-mixing convolution, the output of the CGSA module is obtained.
[0011] Preferably, the process of re-enhancing through cross-addition of the two modalities to obtain the final features is a cross-fusion process. The expression of the cross-fusion process is: ; ; where and represent the cross-fusion results of the two modalities, represents the adaptive gating unit AG.
[0012] Preferably, the output of the CGSA module is the fusion feature of the visible light and infrared modalities, expressed as: ; where represents the 1×1 convolution operation.
[0013] Preferably, the AG module adaptively determines the contribution degree of each modality feature to the final fusion result according to the characteristics of the input data by introducing a multiplicative gating mechanism and a non-linear transformation, and effectively filters out noise information.
[0014] Preferably, the expression of the AG module is: ; ; where represents the 1×1 convolution operation, tanh represents the non-linear transformation activation function, represents the Sigmoid function, represents the dynamically learnable weight.
[0015] In summary, compared with traditional technologies, the multi-modal remote sensing small target detection method across the spatial-frequency domain of the present invention designs a dual-modal feature-level fusion target detection architecture, distributes and extracts the features of each visible light and infrared channel, and proposes a cross-domain gated self-attention fusion module for the interaction and fusion of dual-modal features, improving the accuracy and robustness of remote sensing small target detection. At the same time, it has extremely low computational complexity, is suitable for mobile deployment on satellite-borne and airborne platforms, and improves the image processing accuracy and response speed.
[0016] The technical method of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0017] Figure 1 It is a flowchart of the steps of a multi-modal remote sensing small target detection method across the spatial-frequency domain of the present invention; Figure 2 It is a diagram of the improved multi-modal remote sensing small target detection model of the present invention; Figure 3 It is a schematic diagram of the principle of the cross-domain gated self-attention fusion module; Figure 4 It is a comparison diagram of the target detection results of different models under the DroneVehicle dataset. Figure 4 In (a), it is a pair of visible light-infrared images, where there are cases of dense targets and targets in the dark. Figure 4 In (b), it is a pair of visible light-infrared images, where there are cases of dense targets and targets under strong light illumination. Detailed Embodiments
[0018] The technical method of the present invention will be further described below with reference to the accompanying drawings and embodiments. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present application.
[0019] The description of at least one exemplary embodiment is merely illustrative in nature and in no way limits the present application, its application, or its use.
[0020] Technologies, systems, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, systems, and devices should be regarded as part of the specification.
[0021] In all examples shown and discussed herein, any specific values should be construed as merely exemplary and not as a limitation. Accordingly, other examples of the exemplary embodiments may have different values.
[0022] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings as understood by those of ordinary skill in the field to which the present invention pertains.
[0023] The present invention provides a multi-modal remote sensing small target detection method across the spatial-frequency domain, as Figure 1 shown. The present application redesigned the feature extraction network of YOLOv8n to form a two-stream backbone, and extracted depth features from visible light and infrared images respectively. Secondly, the present application designed a cross-domain gated self-attention fusion module (CGSA) to integrate frequency domain and spatial domain information, and introduced a lightweight self-attention mechanism to achieve long-range modeling without increasing complexity. Finally, the present application retained the detection head structure of the original network.
[0024] Specifically, it includes the following steps: S1. Obtain a remote sensing image, and preprocess the remote sensing image to obtain an initial image.
[0025] S2. Use the improved multi-modal remote sensing small target detection model to process the initial image to obtain the target detection result of the remote sensing image.
[0026] As Figure 2 shown, the improved multi-modal remote sensing small target detection model includes a YOLOv8n two-stream backbone network, an improved cross-domain gated self-attention fusion module, a lightweight self-attention mechanism, and a YOLOv8n detection head.
[0027] Preferably, the specific content of using the improved multi-modal remote sensing small target detection model to process the initial image to obtain the target detection result of the remote sensing image in S2 is: S201. Divide the initial image to obtain a visible light image and an infrared image.
[0028] S202. Use the YOLOv8n two-stream backbone network to perform multi-scale feature fusion on the visible light image and the infrared image respectively to obtain image features, and the image features include visible light features and infrared features .
[0029] S203. Use the improved cross-domain gated self-attention fusion module to integrate frequency domain information and spatial domain information of the feature image to obtain information to be detected.
[0030] S204. Use the YOLOv8n detection head to detect the information to be detected to obtain the target detection result, and the target detection result includes the target position and the target category.
[0031] Preferably, in current multi-modal feature fusion methods, an encoder module is usually designed to strengthen the dependence on global information, or convolution is combined with the encoder to achieve comprehensive learning of local and global information. Inevitably, these methods come at the cost of computational complexity. This application designs a lightweight CGSA cross-domain gated self-attention module for multi-modal feature fusion.
[0032] As Figure 3 shown, the specific content of improving the cross-domain gated self-attention fusion module CGSA in S203 includes: Cross-domain differential feature extraction CFE, which is used to promote the learning of complementary features in the frequency domain and the spatial domain during the process of integrating frequency domain information and spatial domain information.
[0033] The spectral convolution theorem in Fourier theory shows that processing information in the Fourier space can capture the global frequency representation, greatly reducing the computational complexity compared to the encoder structure. Assume a visible light feature map , after the fast Fourier transform (FFT), the frequency feature map is .
[0034] Where is the position coordinate in and respectively represent the frequency variables along and directions. In the case of a multi-channel image, the FFT transform is applied to each channel separately. For simplicity, the channel symbols in the formula are omitted in this application. The amplitude spectrum component is , the frequency spectrum component is , and the formula is: .
[0035] .
[0036] and are respectively the real part and the imaginary part of
[0037] Therefore, CFE affects the spatial domain features by extracting and reconstructing the global frequency information of the bimodal, and performs differential enhancement with the spatial local information, promoting the information flow between the frequency domain and the spatial domain to effectively filter and mix features of different modalities.
[0038] Preferably, the running steps of the cross-domain differential feature extraction CFE are: The visible light feature and the infrared feature can obtain the visible light frequency feature map through FFT and the infrared frequency feature map , the amplitude component of visible light can be decomposed and the phase component , the amplitude component of infrared and the phase component , which are respectively expressed as: .
[0039] .
[0040] Concatenate the amplitude components and phase components of the visible light modality and the infrared modality, and perform a feature enhancement operation to generate enhanced global frequency features. The amplitude component is , and the phase component is , and The expressions of are respectively: .
[0041] .
[0042] Among them, Cat[.] represents channel dimension concatenation, and are bottleneck layers, which contain a 1×1 convolution for feature dimensionality reduction, a RELU activation function, and a 1×1 convolution for feature dimension expansion.
[0043] Apply the inverse FFT to convert the enhanced global frequency features back to the spatial domain to obtain the reconstructed feature map affected by the frequency domain .
[0044] .
[0045] represents the inverse FFT, represents the reconstructed feature.
[0046] Use the reconstructed feature to subtract the original spatial domain features, namely the visible light feature and the infrared feature respectively, to obtain the cross-domain differential features of each modality. The cross-domain differential features of each modality include the visible light differential feature and the infrared differential feature , which are respectively expressed as: .
[0047] .
[0048] Among them is the visible light differential feature, is the infrared differential feature, and is a 1×1 convolution operation.
[0049] In view of the advantages that the self-attention mechanism has long-term modeling and is good at capturing the internal correlation of features, this application redesigned the polarization self-attention mechanism for guiding the fusion of visible light and infrared features. Since the visible light features and infrared features are largely spatially aligned, this application focuses on the channel dimension features of different modalities for feature fusion.
[0050] Self-attention Feature Fusion (SAFF) is used to guide the fusion of visible light and infrared features in the feature image and capture long-range dependencies.
[0051] Preferably, the operation steps of the self-attention feature fusion SAFF are as follows: Cross-domain differential features are generated in two branches Q and V composed of 1×1 convolution and Reshape operations respectively and . On the Q branch, the features are fully folded, and on the V branch, the features remain at high resolution.
[0052] Combine the results of the Q branches of the visible light modality and the infrared modality, and perform a Softmax operation to obtain the fused weight distribution. Since the features of the Q branch are compressed features, in order to achieve information enhancement, multiply the fused weight distribution by the information-rich V branch features of the visible light modality and the infrared modality.
[0053] Obtain the dual-channel attention weights through a 1×1 convolutional layer, a Layer Normalization layer, and a Sigmoid function in sequence and .
[0054] Perform channel-level multiplication between the attention weights and the input features to generate a feature representation with less noise. The visible light output feature is and the infrared output feature is to suppress the redundant non-critical information therein while enhancing the critical information affecting detection.
[0055] After filtering the feature representation, perform re-enhancement through cross-addition of the two modalities to obtain the final feature.
[0056] After channel concatenation and 1×1 channel mixing convolution, obtain the output of the CGSA module.
[0057] Preferably, the process of performing re-enhancement through cross-addition of the two modalities to obtain the final feature is a cross-fusion process, and the expression of the cross-fusion process is: 。
[0058] 。
[0059] Among them, and represent the cross - fusion results of two modalities, represents the adaptive gating unit AG.
[0060] The adaptive gating AG unit is used to dynamically allocate the weights of visible - light and infrared feature fusion, achieving more sensitive and adaptive feature cross - fusion.
[0061] It can be understood that considering that equal treatment of bimodal features for fusion may lead to sub - optimal results, this application designs an adaptive and learnable gating unit for multimodal feature cross - fusion. Specifically, the AG module adaptively determines the contribution degree of each modal feature to the final fusion result according to the characteristics of the input data by introducing a multiplicative gating mechanism and a non - linear transformation, and effectively filters out noise information. It can enhance the flexible and dynamic adjustment ability of the model to the importance of different modal features, thereby improving the fusion performance and robustness of the model in multimodal tasks.
[0062] Preferably, the output of the CGSA module is the fusion feature of the visible - light and infrared modalities, expressed as: 。
[0063] Among them represents a 1×1 convolution operation.
[0064] Preferably, the AG module adaptively determines the contribution degree of each modal feature to the final fusion result according to the characteristics of the input data by introducing a multiplicative gating mechanism and a non - linear transformation, and effectively filters out noise information.
[0065] Preferably, the expression of the AG module is: 。
[0066] 。
[0067] Among them, represents a 1×1 convolution operation, tanh represents a non - linear transformation activation function, represents the Sigmoid function, represents the dynamically learnable weight. 1. Bimodal feature - level fusion object - detection architecture.
[0068] To prove the effectiveness and lightweight degree of the algorithm of this application, the present invention conducts object - detection tests on the DroneVehicle dataset.
[0069] DroneVehicle is a recently released large RGB-IR remote sensing vehicle target detection dataset captured by drones. The dataset contains 28,439 pairs of visible-infrared images, including urban roads, residential areas, and parking lots, covering various complex shooting scenarios from day to night. It includes five vehicle types, namely 'car', 'truck', 'bus', 'van', and 'freight car'.
[0070] In the experiments of this application, 17,990 pairs of image pairs were selected as the training set, 1,469 pairs as the validation set, and 8,980 pairs as the test set.
[0071] The experiments were conducted using the Pycharm IDE on the Windows 11 system. The environment includes Python 3.8, Pytorch 1.8, and CUDA 11.1. The computing resources include an Intel Core i7-13700KF CPU and an NVIDIA RTX 4090 GPU. The training ran for 300 epochs with a batch size of 16. The initial learning rate of the SGD optimizer was 0.01, the momentum was 0.937, and the weight decay was 0.0005.
[0072] In this application, ablation experiments were conducted on the cross-domain gated self-attention fusion module (CGSA) designed in this application to judge the impact of this module on network performance, as shown in Table 1. The basic model of this application uses the method of Concat fusion in the middle layer of the dual-modal YOLOv8n (YOLO8n+MidFusionCAT). After using the self-attention feature fusion (SAFF) module to replace the Concat fusion method of the basic model, the mAP50 increased by 1.5% with only a slight increase in computational overhead. In addition, the cross-domain differential feature extraction (CFE) module and the adaptive gating (AG) module designed in this application were respectively added to the SAFF, and the recognition accuracy increased by 1.7% and 1.8% respectively, proving the effectiveness of introducing cross-domain differential features of frequency-domain global feature extraction and the adaptive gating unit to adaptively learn the feature fusion weights. Finally, introducing the complete CGSA module increased the mAP50 by 2.0% compared with the baseline model, proving that this module can improve the deep fusion ability of different modal features.
[0073] Table 1 Ablation experiments of the cross-domain gated self-attention fusion module on the DroneVehicle dataset ;
[0074] As shown in Table 2, it can be seen that under the DroneVehicle remote sensing small target dataset, compared with the single-mode or basic multi-mode fusion models of YOLOv8n and YOLOv11n, the present invention makes full use of cross-modal cross-space-frequency domain information, demonstrating higher recognition accuracy. The recognition accuracy mAP50 reaches 84.9%, the number of parameters is only 6.31M, and the computational volume is 13.9 GFLOPs. Compared with the YOLOv8n visible light single-mode model, the recognition accuracy mAP50 of the present invention is increased by 12.7%. Compared with the YOLOv8n infrared single-mode model, the recognition accuracy mAP50 of the present invention is increased by 5.3%.
[0075] Table 2 Comparative experiment on target detection results on the DroneVehicle dataset ;
[0076] To more intuitively display the detection performance of this method, the present application selects different methods and the model of the present application for qualitative comparison. Specifically, as Figure 4 shown in (a) of [reference], the freight car target above the night visible light image is in the dark area. Only the YOLOv8n infrared single-mode and the model of the present application achieve accurate detection. There are missed detections and misclassifications in both the visible light single-mode and the Baseline model. Although the infrared single-mode detection is not affected by light at night, due to the lack of texture information in the infrared image, it is easy to cause confusion for targets with similar contours. As Figure 4 shown in (b) of [reference], there are misclassification problems in the YOLOv8n infrared detection mode. This comparative experiment proves the robustness of the model of the present application for detecting complex remote sensing small targets.
[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical method of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical method of the present invention, and these modifications or equivalent replacements cannot make the modified technical method deviate from the spirit and scope of the technical method of the present invention.
Claims
1. A multi-modal remote sensing small target detection method across the spatial-frequency domain, characterized in that It includes the following steps: S1. Obtain a remote sensing image and preprocess the remote sensing image to obtain an initial image; S2. Use an improved multi-modal remote sensing small target detection model to process the initial image to obtain the target detection result of the remote sensing image; The improved multi-modal remote sensing small target detection model includes a YOLOv8n dual-stream backbone network, a cross-domain gated self-attention fusion module, and a YOLOv8n detection head.
2. A multi-modal remote sensing small target detection method across spatial-frequency domains according to claim 1, characterized in that The specific content of using the improved multi-modal remote sensing small target detection model to process the initial image in S2 to obtain the target detection result of the remote sensing image is as follows: S201. Divide the initial image to obtain a visible light image and an infrared image; S202. Use the YOLOv8n dual-stream backbone network to perform multi-scale feature fusion on the visible light image and the infrared image respectively to obtain image features, where the image features include visible light features and infrared features ; S203. Design a cross-domain gated self-attention fusion module to integrate the frequency domain information and spatial domain information of the visible light and infrared feature images, and generate a multi-scale fusion feature map; S204. Use the YOLOv8n detection head to detect different scales of the fusion feature images to obtain the target detection result, and the target detection result includes the target position, target category, and prediction confidence.
3. A multi-modal remote sensing small target detection method across the spatial-frequency domain according to claim 2, characterized in that, The specific content of the improved cross-domain gated self-attention fusion module CGSA in S203 includes: Cross-domain differential feature extraction CFE, which is used to promote the learning of complementary features in the frequency domain and spatial domain during the process of integrating frequency domain information and spatial domain information; Self-attention feature fusion SAFF, which is used to guide the fusion of visible light and infrared features in the feature image and capture long-range dependencies; Adaptive gating AG unit, which is used to dynamically allocate the weights of the fusion of visible light and infrared features to achieve more sensitive and adaptive feature cross-fusion.
4. A multi-modal remote sensing small target detection method across the spatial-frequency domain according to claim 3, characterized in that The running steps of the cross-domain differential feature extraction CFE are as follows: Visible light features and infrared features The visible light frequency feature map can be obtained by FFT and the infrared frequency feature map , and the amplitude component of the visible light can be decomposed and the phase component , the amplitude component of the infrared and the phase component , which are respectively expressed as: ; ; The amplitude components and phase components of the visible light modality and the infrared modality are concatenated and subjected to a feature enhancement operation to generate enhanced global frequency features. The amplitude component is , and the phase component is . and The expressions of are respectively: ; ; Among them, Cat[.] represents concatenation at the channel dimension level, and is a bottleneck layer, which includes a 1×1 convolution for feature dimensionality reduction, a ReLU activation function, and a 1×1 convolution for feature dimensionality increase; Apply the inverse FFT to transform the enhanced global frequency feature back to the spatial domain to obtain the reconstructed feature affected by the frequency domain: ; Represents the inverse FFT, Indicates the reconstructed feature; Using the reconstructed features Subtracting the original spatial domain features respectively gives the visible light features and the infrared features , to obtain the cross-domain differential features of each modality, and the cross-domain differential features of each modality include visible light differential features and infrared differential features , which are respectively expressed as: ; ; Among them is the visible light differential feature, is the infrared differential feature, and is the 1×1 convolution operation.
5. A multi-modal remote sensing small target detection method across the spatial-frequency domain according to claim 3, characterized in that, The running steps of the self-attention feature fusion SAFF are as follows: Cross-domain differential features are generated in two branches, Q and V, which are composed of 1×1 convolution and Reshape operations respectively. and , on the Q branch, the features are fully folded, while on the V branch, the features remain at high resolution. Combine the Q-branch results of the visible light modality and the infrared modality, and perform a Softmax operation to obtain the fused weight distribution, and multiply the fused weight distribution by the information-rich V-branch features of the visible light modality and the infrared modality; Obtain the dual-channel attention weights through a 1×1 convolutional layer, a Layer Normalization layer, and a Sigmoid function in sequence and ; Perform channel-wise multiplication between the attention weights and the input features to generate a feature representation. The visible light output feature is and the infrared output feature is , to suppress redundant non-critical information therein while enhancing critical information that affects detection; After filtering the feature representation, perform recalibration enhancement through cross-addition of the two modalities to obtain the final feature; After channel concatenation and 1×1 channel mixing convolution, obtain the output of the CGSA module.
6. A multi-modal remote sensing small target detection method in the cross space-frequency domain according to claim 5, characterized in that The process of performing recalibration enhancement through cross-addition of the two modalities to obtain the final feature is the cross-fusion process, and the expression of the cross-fusion process is: ; ; Among them, is the visible light output feature, is the infrared output feature, and represent the cross-fusion result of two modalities, represents the adaptive gating unit AG.
7. A multi-modal remote sensing small target detection method in the cross-space-frequency domain according to claim 5, characterized in that The output of the CGSA module is the fusion feature of the visible light and infrared modalities, which is expressed as: ; Among them represents a 1×1 convolution operation.
8. A multi-modal remote sensing small target detection method in the cross-space-frequency domain according to claim 3, characterized in that, The AG module adaptively determines the contribution degree of each modality feature to the final fusion result according to the characteristics of the input data by introducing a multiplicative gating mechanism and a non-linear transformation, and effectively filters out noise information.
9. A multi-modal remote sensing small target detection method in the cross-space-frequency domain according to claim 8, characterized in that, The expression of the AG module is: ; ; Among them, and are any two features to be fused, represents a 1×1 convolution operation, tanh represents a non-linear transformation activation function, represents the Sigmoid function, represents the dynamically learnable weight.
Citation Information
Patent Citations
Multispectral pedestrian detection method based on differential attention and frequency domain fusion
CN117710899A
Multi-spectral target detection method based on multi-modal interaction and fusion
CN118799832A
Double-dynamic cross-modal interaction remote sensing image target detection method
CN119445063A
Method for detecting infrared ship target based on improved yolov7
US20250078541A1
Cited By
Multi-scale remote sensing image target detection method and system based on frequency fusion
CN120563818A
Dual-band semantic segmentation method and device and medium
CN121767617A
A dual-band semantic segmentation method, device and medium
CN121767617B