Multi-scale remote sensing image target detection method and system based on frequency fusion

Through the multi-scale remote sensing image target detection method of frequency fusion, the ResNet network and the frequency interactive Transformer framework are used for feature interaction, which solves the problem of the existing technology that it is difficult to simultaneously capture the features of large and small targets, and improves the accuracy of remote sensing image detection.

CN120563818BActive Publication Date: 2025-09-30NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511048174.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-09-30
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

Existing remote sensing image target detection algorithms find it difficult to simultaneously capture the contour features of large targets and the fine details of small targets in a single framework. In addition, existing methods lack the full utilization of high- and low-frequency information, resulting in reduced accuracy in multi-scale target detection.

Method used

A multi-scale remote sensing image target detection method based on frequency fusion is adopted. The ResNet network is used to extract multi-scale features. Frequency feature interaction is performed through a shallow multi-branch representation fusion network and a frequency interaction Transformer framework, including frequency interaction in the channel direction and spatial direction, avoiding the up and down sampling operations of the traditional feature pyramid structure.

Benefits of technology

It improves the accuracy of multi-scale detection of remote sensing images, ensures the completeness of small target information and the fidelity of large target features, and improves the accuracy of target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120563818B_ABST
    Figure CN120563818B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-scale remote sensing image target detection method and system based on frequency fusion in the field of remote sensing image target detection technology. The method first uses a ResNet network to extract multi-scale features of the remote sensing image data to be measured and then performs channel adjustment to obtain a first feature map, a second feature map, a third feature map and a fourth feature map with decreasing scales. The first feature map and the second feature map are then spliced ​​in the channel direction and channel adjustment is performed to obtain a first fused feature map. The first fused feature map is then input into a shallow multi-branch representation fusion network to obtain a second fused feature map. Then, under the frequency interaction Transformer framework, the second fused feature map, the third feature map and the fourth feature map are serially subjected to channel direction frequency interaction and spatial direction frequency interaction to obtain a frequency fused multi-scale feature map. Finally, the frequency fused multi-scale feature map is input into a detection head network to obtain regression prediction results and classification prediction results. The method of the present invention can improve the accuracy of multi-scale detection of remote sensing images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image target detection, and in particular to a multi-scale remote sensing image target detection method and system based on frequency fusion. Background Art

[0002] In recent years, the rapid development of optical remote sensing technology has significantly improved the quality and quantity of remote sensing images. Various imaging platforms, such as satellites and aircraft, can now achieve multi-scale imaging, ranging from small areas to global coverage. This has made optical remote sensing a crucial means of acquiring large-scale, high-resolution ground object information, and has found widespread application in fields such as resource exploration, disaster prevention, urban planning, and ecological and environmental monitoring. However, due to the varying heights of remote sensing platforms, the scale of remotely sensed objects often varies significantly, making it difficult for existing detection algorithms to simultaneously capture the outline features of large objects and the fine details of small objects within a single framework.

[0003] High-resolution remote sensing images contain rich frequency information. Because high- and low-frequency information can effectively reflect the global structure and detailed features of objects of varying sizes, it can help detection models better capture the diverse characteristics of targets. This is crucial for addressing the multi-scale problem caused by the large size variations of remote sensing objects in remote sensing images. Introducing frequency information into target detection models can enhance the model's robustness to noise and background interference, further improving detection accuracy and stability. However, existing methods for incorporating frequency information into target detection models typically focus solely on extracting frequency features and fail to fully exploit the interactions between multi-band features. This results in insufficient exploration of the synergy between high- and low-frequency information, making it difficult to fully characterize the diverse characteristics of multi-scale targets. Furthermore, the feature pyramid structures used in existing target detection networks for multi-scale interaction often rely on repeated up- and down-sampling operations to fuse feature maps of different sizes. This further leads to information loss for small targets and feature distortion for large targets, resulting in reduced accuracy in multi-scale detection of remote sensing images. Summary of the Invention

[0004] Based on this, it is necessary to provide a multi-scale remote sensing image target detection method and system based on frequency fusion to address the above technical problems, which can improve the accuracy of multi-scale detection of remote sensing images.

[0005] In a first aspect, the present invention provides a multi-scale remote sensing image target detection method based on frequency fusion, comprising the following steps:

[0006] The ResNet network is used to extract multi-scale features of the remote sensing image data to be measured and then channel adjustment is performed to obtain the first feature map, the second feature map, the third feature map and the fourth feature map with decreasing scale;

[0007] Splicing the first feature map and the second feature map in the channel direction and performing channel adjustment to obtain a first fused feature map;

[0008] Input the first fused feature map into the shallow multi-branch representation fusion network to obtain the second fused feature map;

[0009] In the frequency interaction Transformer framework, the second fusion feature map, the third feature map and the fourth feature map are serially subjected to channel-wise frequency interaction and spatial-wise frequency interaction to obtain a frequency-fused multi-scale feature map.

[0010] The frequency fused multi-scale feature map is input into the detection head network to obtain the regression prediction results and classification prediction results of the remote sensing image data to be tested.

[0011] In one embodiment, the shallow multi-branch representation fusion network includes a serial multi-branch frequency enhancement sub-network and a frequency representation fusion sub-network;

[0012] The multi-branch frequency enhancement sub-network is a cross-stage partially connected structure. The part of the multi-branch frequency enhancement sub-network used to enhance the first fusion feature map includes parallel local branches, large branches and frequency feature enhancement branches.

[0013] In one embodiment, the local branch is a 1x1 depthwise convolutional layer with no padding and a stride of 1;

[0014] The large branch includes a 31x31 depth convolution layer, a 31x1 depth convolution layer, and a 1x31 depth convolution layer with a parallel stride of 1. The 31x31 depth convolution layer is padded with 15 in both width and height. The 31x1 depth convolution layer is padded with 15 in width, and the 1x31 depth convolution layer is padded with 15 in height.

[0015] The frequency feature enhancement branch includes a serial channel frequency enhancement structure and a spatial frequency enhancement structure. Both the channel frequency enhancement structure and the spatial frequency enhancement structure utilize fast Fourier transform and inverse fast Fourier transform to enhance the feature frequency representation.

[0016] In one embodiment, in a frequency interaction Transformer framework, serially performing channel-wise frequency interaction and spatial-wise frequency interaction on the second fused feature map, the third feature map, and the fourth feature map to obtain a frequency fused multi-scale feature map includes the following steps:

[0017] Spatial alignment is performed on the second fused feature map, the third feature map, and the fourth feature map to obtain a channel token sequence;

[0018] Use different window sizes to perform strip sliding window operations on the channel token sequence to obtain multiple groups of channel sub-token sequences;

[0019] Use the cross-layer channel frequency encoder to perform channel-wise frequency fusion on each group of channel sub-token sequences to obtain the intra-group fused token sequence;

[0020] All fused token sequences within the group are recombined to obtain channel-wise frequency-fused multi-scale feature maps;

[0021] Using different window sizes, rectangular sliding window operations are performed on the channel direction frequency fusion multi-scale feature map to obtain multiple sets of two-dimensional features;

[0022] The cross-layer spatial frequency encoder is used to perform frequency fusion of each group of two-dimensional features in the spatial direction to obtain the intra-group fused two-dimensional features;

[0023] All the fused two-dimensional features within the group are recombined to obtain the spatial direction frequency fusion multi-scale feature map;

[0024] The spatial direction frequency fusion multi-scale feature map is added to the second fusion feature map, the third feature map and the fourth feature map through the residual structure to obtain the frequency fusion multi-scale feature map.

[0025] In one embodiment, the cross-layer channel frequency encoder and the cross-layer spatial frequency encoder are both Transformer encoders, and both include a frequency attention module based on a multi-head attention mechanism;

[0026] The frequency attention module includes high-frequency interaction units and low-frequency interaction units, and the multi-head allocation ratio of high-frequency interaction units and low-frequency interaction units is α∈(0,1);

[0027] The high-frequency interaction unit adopts the self-attention mechanism, and the low-frequency interaction unit adopts the cross-attention mechanism.

[0028] In one embodiment, the input of the high-frequency interaction unit of the cross-layer channel frequency encoder and the Q input of the low-frequency interaction unit are sequences obtained by linearly transforming the channel sub-token sequences, and the KV input of the low-frequency interaction unit is a sequence obtained by linearly transforming and average pooling the channel sub-token sequences.

[0029] The input of the high-frequency interaction unit of the cross-layer spatial frequency encoder and the Q input of the low-frequency interaction unit are both spatial sub-token sequences obtained by linearly transforming the two-dimensional features of different scales after splicing in the spatial direction. The KV input of the low-frequency interaction unit is a spatial sub-token sequence obtained by averaging and merging the two-dimensional features of different scales and splicing in the spatial direction.

[0030] In one embodiment, the first feature map, the second feature map, the third feature map, and the fourth feature map are feature maps obtained by downsampling the remote sensing image data to be measured by 4 times, 8 times, 16 times, and 32 times, respectively, and then adjusting the channel direction.

[0031] In one embodiment, in the strip sliding window operation, the window sizes of the token sequences corresponding to the second fused feature map, the third feature map, and the fourth feature map are 1, 3, and 5, respectively, and the overlapping window steps are 1, 2, and 4, respectively;

[0032] In the rectangular sliding window operation, the window sizes of the two-dimensional features corresponding to the second fused feature map, the third feature map, and the fourth feature map are 1x1, 5x5, and 17x17, respectively, and the overlapping window steps are 1, 4, and 16.

[0033] In one embodiment, the detection head network is composed of a Transformer decoder and a feedforward neural network.

[0034] In a second aspect, the present invention further provides a multi-scale remote sensing image target detection system based on frequency fusion, the system comprising:

[0035] A feature extraction unit is used to extract multi-scale features of the remote sensing image data to be measured using a ResNet network and then perform channel adjustment to obtain a first feature map, a second feature map, a third feature map, and a fourth feature map with decreasing scales;

[0036] A feature splicing unit, configured to splice the first feature map and the second feature map in a channel direction and perform channel adjustment to obtain a first fused feature map;

[0037] A first fusion unit is used to input the first fused feature map into a shallow multi-branch representation fusion network to obtain a second fused feature map;

[0038] The second fusion unit is used to serially perform channel-wise frequency interaction and spatial-wise frequency interaction on the second fused feature map, the third feature map, and the fourth feature map under the frequency-interactive Transformer framework to obtain a frequency-fused multi-scale feature map;

[0039] The prediction unit is used to input the frequency fusion multi-scale feature map into the detection head network to obtain the regression prediction results and classification prediction results of the remote sensing image data to be tested.

[0040] The beneficial effects of the present invention are as follows: the multi-scale remote sensing image target detection method and system based on frequency fusion of the present invention can fully mine the rich multi-band feature information in the first fusion feature map by using a shallow multi-branch representation fusion network, thereby improving the final target detection accuracy. Furthermore, the present invention uses the attention mechanism of cross-layer channel interaction and spatial frequency interaction under the frequency interaction Transformer framework to fully serially interact with the second fusion feature map mined by the shallow multi-branch representation fusion network, i.e., the multi-band feature information, which can comprehensively characterize the diverse characteristics of multi-scale targets. In addition, the serial interaction avoids the interactive fusion feature map of the up-and-down sampling operation of the traditional feature pyramid structure, thereby ensuring the integrity of the small target information and the fidelity of the large target features in the feature interaction stage, further improving the target detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is one of the flow charts of the multi-scale remote sensing image target detection method based on frequency fusion provided by an embodiment of the present invention;

[0042] Figure 2 Schematic diagram of the structure of a shallow multi-branch representation fusion network provided by an embodiment of the present invention;

[0043] Figure 3 Schematic diagram of the network structure of the channel frequency enhancement structure provided by an embodiment of the present invention;

[0044] Figure 4 is a schematic diagram of the network structure of the spatial frequency enhancement structure provided by an embodiment of the present invention;

[0045] Figure 5 Schematic diagram of the structure of the frequency characterization fusion sub-network provided by an embodiment of the present invention;

[0046] Figure 6 This is a schematic diagram of a process for performing frequency interaction on a multi-scale feature map under a frequency interaction Transformer framework provided by an embodiment of the present invention;

[0047] Figure 7 Schematic diagram of channel frequency interaction attention provided by an embodiment of the present invention;

[0048] Figure 8 is a schematic diagram of spatial-frequency interactive attention provided by an embodiment of the present invention;

[0049] Figure 9 It is a structural diagram of a multi-scale remote sensing image target detection system based on frequency fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0051] In one embodiment, Figure 1 As shown, Figure 1 FIG1 is one of the flow charts of the multi-scale remote sensing image target detection method based on frequency fusion provided by an embodiment of the present invention. The multi-scale remote sensing image target detection method based on frequency fusion in this embodiment includes the following steps:

[0052] S101, using a ResNet network to extract multi-scale features of the remote sensing image data to be measured and then performing channel adjustment to obtain a first feature map, a second feature map, a third feature map, and a fourth feature map with decreasing scales.

[0053] The remote sensing image data to be measured is a remote sensing image obtained by extracting frames from a remote sensing video of a remote sensing image imaging platform.

[0054] Specifically, the scale of the remote sensing image data to be tested processed by the ResNet network is 640×640, and the first feature map, the second feature map, the third feature map, and the fourth feature map are feature maps obtained after the remote sensing image data to be tested is downsampled by 4 times, 8 times, 16 times, and 32 times and then the channel direction is adjusted, and the scales are 160×160, 80×80, 40×40, and 20×20, respectively.

[0055] The ResNet network performs convolution and pooling operations on the remote sensing image data to obtain feature map A1. Four consecutive residual block downsampling operations are performed on A1 to obtain feature maps A2 to A5. The channel dimensions of feature maps A2 to A5 are adjusted to the preset dimensions through 1×1 convolution to obtain the first feature map C2, the second feature map C3, the third feature map C4, and the fourth feature map C5. Specifically, in this embodiment, the preset dimension is 256.

[0056] S102: Splice the first feature map and the second feature map in the channel direction and perform channel adjustment to obtain a first fused feature map.

[0057] Specifically, in this embodiment, a 3x3 convolution is used to splice the first feature map C2 and the second feature map C3 in the channel direction and perform channel adjustment to obtain a first fused feature map D3.

[0058] S103: Input the first fused feature map into a shallow multi-branch representation fusion network to obtain a second fused feature map.

[0059] Specifically, such as Figure 2 As shown, Figure 2This is a schematic diagram of the structure of a shallow multi-branch representation fusion network provided by an embodiment of the present invention. The shallow multi-branch representation fusion network includes a serial multi-branch frequency enhancement sub-network and a frequency representation fusion sub-network. The output of the multi-branch frequency enhancement sub-network is the input of the frequency representation fusion sub-network.

[0060] The multi-branch frequency enhancement sub-network is a cross-stage partially connected structure. The part of the multi-branch frequency enhancement sub-network used to enhance the first fusion feature map includes parallel local branches, large branches and frequency feature enhancement branches.

[0061] In this embodiment, the multi-branch frequency enhancement sub-network is provided with multiple branches. By collaboratively representing the shallow features with sufficient information of small targets in the spatial and frequency domains, it is possible to mine the frequency features of small targets and perform feature enhancement. The frequency representation fusion sub-network FR-Fusion further integrates the features extracted from multiple branches to generate more discriminative comprehensive features, thereby improving the ability of the shallow multi-branch representation fusion network to perceive small targets with minimal computational overhead.

[0062] In general, the shallow multi-branch representation fusion network mines the rich frequency information of small targets in shallow features, thereby fully enhancing the feature expression of small targets in the first fusion feature map and improving the final target detection accuracy.

[0063] S104. Under the frequency interaction Transformer framework, serially perform channel-wise frequency interaction and spatial-wise frequency interaction on the second fused feature map, the third feature map, and the fourth feature map to obtain a frequency-fused multi-scale feature map.

[0064] In this embodiment, a frequency-interactive Transformer framework is used to serially perform channel-wise and spatial-directional frequency interaction on the multi-band feature information extracted by a shallow multi-branch representation fusion network. This fully leverages the interactive relationship between multi-band feature information and comprehensively characterizes the diverse characteristics of multi-scale targets. Furthermore, the serial channel-wise and spatial-directional interaction avoids the interactive fusion of feature maps using upsampling and downsampling operations in the traditional feature pyramid structure. This ensures that small target information is complete and large target feature fidelity is preserved during the feature interaction phase, improving target detection accuracy.

[0065] S105: Input the frequency fused multi-scale feature map into the detection head network to obtain the regression prediction result and the classification prediction result of the remote sensing image data to be tested.

[0066] Specifically, the detection head network in this embodiment consists of a Transformer decoder and a feedforward neural network. Furthermore, after obtaining regression and classification prediction results, detection boxes can be drawn in the smoke image to be tested and the categories can be labeled based on the two prediction results, enabling visualization of the detection results.

[0067] In one embodiment, the local branch is a 1x1 depthwise convolutional layer with no padding and a stride of 1.

[0068] The large branch consists of a parallel 31x31, 31x1, and 1x31 depthwise convolutional layers, all with a stride of 1. The 31x31 depthwise convolutional layer has padding of 15 in both width and height, the 31x1 depthwise convolutional layer has padding of 15 in width, and the 1x31 depthwise convolutional layer has padding of 15 in height. In the large branch, a 31x31 depthwise convolution is introduced to achieve a larger receptive field, thereby capturing more global information. Furthermore, to more efficiently extract contextual information about the stripe, 31x1 and 1x31 depthwise convolutions are designed in parallel with the block-shaped convolution. This combination enables modeling in both the horizontal and vertical directions, thereby improving the perception of fine-grained features. In the local branch, a simple 1x1 depthwise convolutional layer is used for modulation. In the frequency feature enhancement branch, a dual-domain processing mechanism is introduced to further enhance the representation of frequency features.

[0069] The frequency feature enhancement branch includes a serial channel frequency enhancement structure CFE and a spatial frequency enhancement structure SFE. Figure 3 and Figure 4 As shown, Figure 3 is a schematic diagram of the network structure of the channel frequency enhancement structure provided by an embodiment of the present invention, Figure 4 3 is a schematic diagram of the network structure of the spatial frequency enhancement structure provided by an embodiment of the present invention. Both the channel frequency enhancement structure and the spatial frequency enhancement structure utilize fast Fourier transform and inverse fast Fourier transform to enhance the characteristic frequency representation.

[0070] Specifically, the channel frequency enhancement structure CFE of this embodiment first performs a fast Fourier transform FFT on the branch features of the input frequency feature enhancement branch to convert them into frequency domain complex numbers, and the frequency domain complex numbers contain amplitude spectrum and phase spectrum information. At the same time, a 1x1 convolution layer is used to perform global average pooling on the branch features of the input frequency feature enhancement branch to generate channel-level attention weights, and then the frequency domain complex numbers and the channel-level attention weights are multiplied to achieve channel-direction frequency domain feature enhancement, and then the frequency domain features are converted to time domain space through the inverse fast Fourier transform IFFT, and then the time domain space features are input into another 1x1 convolution layer, and the output of the other 1x1 convolution layer is multiplied with the output of the bypass branch to obtain the final output of the frequency interaction structure CFE. The channel frequency enhancement structure CFE further optimizes the feature representation of the channel dimension, so that the model pays more attention to the channel frequency information that helps to distinguish small targets from the background.

[0071] The spatial frequency enhancement structure (SFE) simultaneously performs a fast Fourier transform (FFT) and global average pooling on the final output of the channel frequency enhancement structure (CFE). The resulting frequency-domain complex number is then multiplied by the spatial-level attention weight to enhance spatial frequency-domain features. The features are then converted to the time domain using an inverse fast Fourier transform (IFFT) to obtain the final output of the frequency feature enhancement branch. The spatial frequency enhancement structure (SFE) enhances the capture of spatial details of small objects, such as edges and textures, helping the model accurately locate small objects in complex backgrounds.

[0072] In the multi-branch frequency enhancement sub-network, a cross-stage partially connected structure, the local branch, large branch, and frequency feature enhancement branch each extract feature information at different levels of the first fused feature map, and fuse the outputs of each branch through an addition operation. The fused features of each branch are then channel-modulated through a 1x1 convolution to enrich the final feature representation, which is then concatenated with the output of the bypass branch to produce a multi-branch frequency enhancement feature map.

[0073] like Figure 5 As shown, Figure 5 It is a structural diagram of the frequency characterization fusion sub-network provided by an embodiment of the present invention. The frequency characterization fusion sub-network FR-Fusion splits the multi-branch frequency enhancement feature map in the channel direction with a ratio of 1:1, performs 1x1 convolution, batch normalization and Relu activation respectively, and then splices the two branches in the channel direction. The spliced ​​vector features are subjected to 3 layers of reparameterized convolution, one 1x1 convolution, batch normalization and Relu activation to obtain the final second fusion feature map D3'. By utilizing channel feature recombination and fusion, the features extracted from multiple branches can be further integrated. Subsequently, the multi-branch structure is simulated by reparameterized convolution in the training phase to improve feature discrimination, while it is equivalent to a single-path structure during inference to reduce computational overhead. Finally, 1x1 convolution further fuses and optimizes the features, so that the network can not only efficiently integrate the features of different branches, but also maintain a low computational cost through reparameterization technology, thereby achieving a balance between computational efficiency and improving the discriminability of small target features.

[0074] In one embodiment, Figure 6 As shown, Figure 6 This is a flow chart of frequency interaction of multi-scale feature maps under the frequency interaction Transformer framework provided by an embodiment of the present invention. In the frequency interaction Transformer framework, the second fused feature map, the third feature map, and the fourth feature map are serially subjected to channel direction frequency interaction and spatial direction frequency interaction to obtain a frequency fused multi-scale feature map, including the following steps:

[0075] S601. Perform spatial alignment on the second fused feature map, the third feature map, and the fourth feature map to obtain a channel token sequence.

[0076] S602: Perform a strip sliding window operation on the channel token sequence using different window sizes to obtain multiple groups of channel sub-token sequences.

[0077] Specifically, in the strip sliding window operation, the window sizes of the token sequences corresponding to the second fused feature map, the third feature map, and the fourth feature map are 1, 3, and 5, respectively, and the overlapping window steps are 1, 2, and 4, respectively.

[0078] S603: Use a cross-layer channel frequency encoder to perform frequency fusion in the channel direction on each group of channel sub-token sequences to obtain an intra-group fused token sequence.

[0079] In this embodiment, frequency fusion interaction in the channel direction is performed within different sliding windows, that is, within each group of channel sub-token sequences.

[0080] S604: Recombining all fused token sequences within the group to obtain a channel-wise frequency-fused multi-scale feature map.

[0081] S605 , performing a rectangular sliding window operation on the channel direction frequency fusion multi-scale feature map using different window sizes to obtain multiple groups of two-dimensional features.

[0082] Specifically, in the rectangular sliding window operation, the window sizes of the two-dimensional features corresponding to the second fused feature map, the third feature map, and the fourth feature map are 1x1, 5x5, and 17x17, respectively, and the overlapping window steps are 1, 4, and 16.

[0083] S606: Use a cross-layer spatial frequency encoder to perform frequency fusion in the spatial direction on each group of two-dimensional features to obtain intra-group fused two-dimensional features.

[0084] S607: Reorganize all the fused two-dimensional features within the group to obtain a spatial direction frequency fusion multi-scale feature map.

[0085] S608. Add the spatial direction frequency fusion multi-scale feature map to the second fusion feature map, the third feature map, and the fourth feature map through a residual structure to obtain a frequency fusion multi-scale feature map.

[0086] In this embodiment, the second fused feature map, the third feature map and the fourth feature map are subjected to frequency interaction in the channel direction and frequency interaction in the spatial direction in turn. The successive interaction in the channel direction and the spatial direction avoids the up and down sampling operations of the traditional feature pyramid structure to interact with the fused feature maps, thereby ensuring the completeness of small target information and the fidelity of large target features in the feature interaction stage.

[0087] In one embodiment, Figure 7 and Figure 8 As shown, Figure 7is a schematic diagram of channel frequency interaction attention provided by an embodiment of the present invention, Figure 8 This is a schematic diagram of the spatial-frequency interactive attention provided by an embodiment of the present invention. Both the cross-layer channel frequency encoder and the cross-layer spatial frequency encoder are Transformer encoders, and both include a frequency attention module based on a multi-head attention mechanism.

[0088] The frequency attention module includes high-frequency interaction units and low-frequency interaction units. The multi-head allocation ratio of high-frequency interaction units to low-frequency interaction units is α∈(0,1). High-frequency interaction units use a self-attention mechanism, while low-frequency interaction units use a cross-attention mechanism. Specifically, in this embodiment, α=0.5.

[0089] In one embodiment, the input of the high-frequency interaction unit of the cross-layer channel frequency encoder and the Q input of the low-frequency interaction unit are sequences obtained by linearly transforming the channel sub-token sequences after splicing, and the KV input of the low-frequency interaction unit is a sequence obtained by linearly transforming and average pooling the channel sub-token sequences after splicing.

[0090] The input of the high-frequency interaction unit of the cross-layer spatial frequency encoder and the Q input of the low-frequency interaction unit are both spatial sub-token sequences obtained by linearly transforming the two-dimensional features of different scales after splicing in the spatial direction. The KV input of the low-frequency interaction unit is a spatial sub-token sequence obtained by averaging and merging the two-dimensional features of different scales and splicing in the spatial direction.

[0091] In the attention mechanism, Q, K, and V are three core matrices, representing query, key, and value, respectively. Each head in the cross-layer channel frequency encoder and cross-layer spatial frequency encoder uses linear transformations to calculate Q, K, and V. The output of a single head is then calculated by performing a scaled dot product using Q, K, and V. The output of the high-frequency interaction unit is then concatenated and mapped across multiple heads.

[0092] It should be noted that the outputs of all attention mechanisms are consistent with the shape of the input, so the cross-layer channel direction frequency attention mechanism of this embodiment splices the outputs of the low-frequency interaction unit and the high-frequency interaction unit and divides them into the channel sub-token sequence shape, and the cross-layer spatial direction frequency attention mechanism splices the outputs of the low-frequency interaction unit and the high-frequency interaction unit and deforms them into a two-dimensional feature shape.

[0093] In a specific embodiment, the NWPU VHR-10 public remote sensing dataset has a wide range of target size distribution, which can reflect the detection capabilities of different target detection methods in multi-scale target detection. Therefore, this embodiment experimentally verifies the detection accuracy of a multi-scale remote sensing image target detection method based on frequency fusion based on the NWPU VHR-10 public remote sensing dataset. The NWPU VHR-10 dataset has a total of 3651 target instances labeled, covering 10 categories with huge scale differences. Specifically, they include airplanes (AP), ships (SP), oil tanks (ST), baseball fields (BD), tennis courts (TC), basketball courts (BC), track and field fields (GF), ports (HB), bridges (BR), and vehicles (VE).

[0094] The experimental environment is as follows: the GPU is an NVIDIA GeForce RTX 4090, the CPU is an Intel(R) Xeon(R) Gold 6271C CPU @ 2.60GHz, and the operating system is Linux. mAP:50, mAP:50-95, and the mAP value for each category are used as evaluation metrics. AP represents object detection accuracy, and mAP represents mean object detection precision. mAP is categorized into mAP50 and mAP50:95 based on different IoU thresholds. mAP50 is the average object detection accuracy for all categories at an IoU of 0.5, and mAP50:95 is the average of the average object detection accuracy for IoU thresholds ranging from 0.5 to 0.95.

[0095] The comparison methods used in the experiment include YOLOV8-M, YOLOV10-L, YOLOVX-M, FFCA-YOLO and RT-DETR.

[0096] In this example, the NWPU VHR-10 public remote sensing dataset was randomly divided into a training set and a test set in a ratio of 8:2. The training set was used to train the frequency fusion-based multi-scale remote sensing image target detection method and comparison method of this example. During training, the optimizer, batch size, initial learning rate, momentum, and weight decay were AdamW, 4, 0.0001, 0.9, and 0.0001, respectively. During training, the remote sensing images in the training set were preprocessed, including random flipping, cropping, and color change.

[0097] The test set was used to determine the mAP, mAP50, and mAP50-95 of the frequency fusion-based multi-scale remote sensing image object detection method of this embodiment and the comparison method. The final mAP:50, mAP:50-95, and mAP values ​​for each category for each method are shown in Table 1.

[0098] Table 1 Experimental results of different methods on the test set

[0099]

[0100] As shown in Table 1, the multi-scale remote sensing image target detection method based on frequency fusion of this embodiment has the best performance, with mAP:50 and mAP:50-95 reaching 93.3% and 61.5% respectively. In addition, each type of target maintains a high detection accuracy, which fully proves that the multi-scale remote sensing image target detection method based on frequency fusion of the present invention has a high detection accuracy.

[0101] Based on the same inventive concept, the embodiments of the present application also provide a frequency fusion-based multi-scale remote sensing image target detection system for implementing the frequency fusion-based multi-scale remote sensing image target detection method involved above. The implementation solution provided by this system is similar to the implementation solution described in the above method. Therefore, the specific limitations of one or more frequency fusion-based multi-scale remote sensing image target detection system embodiments provided below can be found in the above-mentioned limitations of the frequency fusion-based multi-scale remote sensing image target detection method, and will not be repeated here.

[0102] like Figure 9 As shown, Figure 9 : is a schematic diagram of the structure of a multi-scale remote sensing image target detection system based on frequency fusion provided by an embodiment of the present invention. The multi-scale remote sensing image target detection system based on frequency fusion of this embodiment includes:

[0103] A feature extraction unit is used to extract multi-scale features of the remote sensing image data to be measured using a ResNet network and then perform channel adjustment to obtain a first feature map, a second feature map, a third feature map, and a fourth feature map with decreasing scales;

[0104] A feature splicing unit, configured to splice the first feature map and the second feature map in a channel direction and perform channel adjustment to obtain a first fused feature map;

[0105] A first fusion unit is used to input the first fused feature map into a shallow multi-branch representation fusion network to obtain a second fused feature map;

[0106] The second fusion unit is used to serially perform channel-wise frequency interaction and spatial-wise frequency interaction on the second fused feature map, the third feature map, and the fourth feature map under the frequency-interactive Transformer framework to obtain a frequency-fused multi-scale feature map;

[0107] The prediction unit is used to input the frequency fusion multi-scale feature map into the detection head network to obtain the regression prediction results and classification prediction results of the remote sensing image data to be tested.

[0108] In the multi-scale remote sensing image target detection system based on frequency fusion of this embodiment, the first fusion unit uses a shallow multi-branch representation fusion network to fully mine the rich frequency feature information in the first fusion feature map. The second fusion unit uses the attention mechanism of cross-layer channel interaction and spatial frequency interaction under the frequency interaction Transformer framework to fully serially interact with the second fusion feature map mined by the shallow multi-branch representation fusion network, which can comprehensively characterize the diverse characteristics of multi-scale targets. In addition, the serial interaction avoids the interactive fusion feature maps of the up and down sampling operations of the traditional feature pyramid structure, thereby ensuring the completeness of small target information and the fidelity of large target features in the feature interaction stage, thereby improving the target detection accuracy.

[0109] The above-described embodiments merely represent several implementation methods of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A multi-scale remote sensing image target detection method based on frequency fusion, characterized in that: The following steps are involved: The ResNet network is used to extract multi-scale features of the remote sensing image data to be measured and then channel adjustment is performed to obtain the first feature map, the second feature map, the third feature map and the fourth feature map with decreasing scale; Splicing the first feature map and the second feature map in the channel direction and performing channel adjustment to obtain a first fused feature map; Inputting the first fused feature map into a shallow multi-branch representation fusion network to obtain a second fused feature map; Under the frequency interaction Transformer framework, channel-wise frequency interaction and spatial-wise frequency interaction are serially performed on the second fused feature map, the third feature map, and the fourth feature map to obtain a frequency fused multi-scale feature map; The frequency fusion multi-scale feature map is input into the detection head network to obtain the regression prediction results and classification prediction results of the remote sensing image data to be tested; The shallow multi-branch representation fusion network includes a serial multi-branch frequency enhancement sub-network and a frequency representation fusion sub-network; The multi-branch frequency enhancement sub-network is a cross-stage partially connected structure, and the part of the multi-branch frequency enhancement sub-network used for enhancing the first fusion feature map includes parallel local branches, large branches and frequency feature enhancement branches; The serially performing channel-wise frequency interaction and spatial-wise frequency interaction on the second fused feature map, the third feature map, and the fourth feature map under the frequency-interactive Transformer framework to obtain a frequency-fused multi-scale feature map includes the following steps: Performing spatial alignment on the second fused feature map, the third feature map, and the fourth feature map to obtain a channel token sequence; Use different window sizes to perform strip sliding window operations on the channel token sequence to obtain multiple groups of channel sub-token sequences; Use the cross-layer channel frequency encoder to perform channel-wise frequency fusion on each group of channel sub-token sequences to obtain the intra-group fused token sequence; All fused token sequences within the group are recombined to obtain channel-wise frequency-fused multi-scale feature maps; Using different window sizes to perform a rectangular sliding window operation on the channel direction frequency fusion multi-scale feature map to obtain multiple groups of two-dimensional features; The cross-layer spatial frequency encoder is used to perform frequency fusion of each group of two-dimensional features in the spatial direction to obtain the intra-group fused two-dimensional features; All the fused two-dimensional features within the group are recombined to obtain the spatial direction frequency fusion multi-scale feature map; The spatial direction frequency fusion multi-scale feature map is added to the second fusion feature map, the third feature map and the fourth feature map through the residual structure to obtain the frequency fusion multi-scale feature map.

2. The multi-scale remote sensing image target detection method based on frequency fusion according to claim 1 is characterized in that: The local branch is a 1x1 depth convolutional layer with no padding and a stride of 1; The large branch includes a 31x31 depth convolution layer, a 31x1 depth convolution layer, and a 1x31 depth convolution layer with a parallel stride of 1. The 31x31 depth convolution layer is padded with a size of 15 in both width and height. The 31x1 depth convolution layer is padded with a size of 15 in width, and the 1x31 depth convolution layer is padded with a size of 15 in height. The frequency feature enhancement branch includes a serial channel frequency enhancement structure and a spatial frequency enhancement structure, and both the channel frequency enhancement structure and the spatial frequency enhancement structure utilize fast Fourier transform and inverse fast Fourier transform to enhance the feature frequency representation.

3. The multi-scale remote sensing image target detection method based on frequency fusion according to claim 1, characterized in that: The cross-layer channel frequency encoder and the cross-layer spatial frequency encoder are both Transformer encoders, and both include a frequency attention module based on a multi-head attention mechanism; The frequency attention module includes a high-frequency interaction unit and a low-frequency interaction unit, and the multi-head allocation ratio of the high-frequency interaction unit and the low-frequency interaction unit is α∈(0,1); The high-frequency interaction unit adopts a self-attention mechanism, and the low-frequency interaction unit adopts a cross-attention mechanism.

4. The multi-scale remote sensing image target detection method based on frequency fusion according to claim 3 is characterized in that: The input of the high-frequency interaction unit of the cross-layer channel frequency encoder and the Q input of the low-frequency interaction unit are both sequences of channel sub-token sequences that are linearly transformed after splicing, and the KV input of the low-frequency interaction unit is a sequence of channel sub-token sequences that are linearly transformed and average pooled after splicing; The input of the high-frequency interaction unit of the cross-layer spatial frequency encoder and the Q input of the low-frequency interaction unit are both spatial sub-token sequences obtained by linearly transforming the two-dimensional features of different scales after splicing them in the spatial direction. The KV input of the low-frequency interaction unit is a spatial sub-token sequence obtained by averaging and merging the two-dimensional features of different scales and splicing them in the spatial direction.

5. The multi-scale remote sensing image target detection method based on frequency fusion according to claim 1, characterized in that: The first feature map, the second feature map, the third feature map and the fourth feature map are feature maps obtained by downsampling the remote sensing image data to be measured by 4 times, 8 times, 16 times and 32 times respectively and then adjusting the channel direction.

6. The multi-scale remote sensing image target detection method based on frequency fusion according to claim 1, characterized in that: In the strip sliding window operation, the window sizes of the token sequences corresponding to the second fused feature map, the third feature map, and the fourth feature map are 1, 3, and 5, respectively, and the overlapping window steps are 1, 2, and 4, respectively; In the rectangular sliding window operation, the window sizes of the two-dimensional features corresponding to the second fused feature map, the third feature map, and the fourth feature map are 1x1, 5x5, and 17x17, respectively, and the overlapping window steps are 1, 4, and 16.

7. The multi-scale remote sensing image target detection method based on frequency fusion according to claim 1, characterized in that: The detection head network consists of a Transformer decoder and a feedforward neural network.

8. A multi-scale remote sensing image target detection system based on frequency fusion, characterized in that: The system implements the multi-scale remote sensing image target detection method based on frequency fusion as described in any one of claims 1 to 7 to realize target detection in multi-scale remote sensing images.