Frequency domain adaptive feature enhancement and attention guidance target detection method and system

Through the cross-scale feature interaction module of frequency domain adaptive feature enhancement and attention-guided, the problem of difficulty in object detection in remote sensing images is solved, efficient feature enhancement and cross-scale information fusion are achieved, and the accuracy of object detection is significantly improved.

CN120147657APending Publication Date: 2025-06-13NANJING UNIV OF INFORMATION SCI & TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510230672.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

There are difficulties in object detection in remote sensing images, mainly due to the low resolution of the target object, insufficient feature information, poor background contrast, and complex remote sensing image scenes and diverse target sizes.

Method used

The frequency domain adaptive feature enhancement module and the attention-guided cross-scale feature interaction module are adopted to enhance high and low frequency features through frequency domain processing, and dynamically adjust feature weights through attention mechanisms to achieve cross-scale fusion of features and redundant information suppression.

Benefits of technology

It significantly improves the accuracy of object detection and feature expression capabilities in remote sensing images, enhances the detection capabilities of small objects, and reduces the dependence on computing overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147657A_ABST
    Figure CN120147657A_ABST
Patent Text Reader

Abstract

The invention discloses a frequency domain adaptive feature enhancement and attention guidance target detection method and system, and relates to the technical field of computer vision and remote sensing image detection and classification, and the method comprises the steps: receiving extracted shallow features and deep features, inputting the shallow features into a pre-established frequency domain adaptive feature enhancement module FAFEM, and carrying out the recognition of the deep features; outputting to obtain a processed feature layer; inputting the processed feature layer and the deep feature into a pre-established input attention guidance cross-scale feature interaction module, and outputting to obtain a fusion feature; and inputting the fusion feature into a detection head, and outputting to obtain a target detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision, remote sensing image detection and classification, specifically a target detection method and system based on frequency-domain adaptive feature enhancement and attention guidance. Background Art

[0002] The classification and detection of remote sensing images aim to identify the categories and location distributions of specific objects in remote sensing images. As one of the remote sensing image analysis tasks, it plays a crucial role in fields such as military reconnaissance, environmental protection, agricultural monitoring, and traffic management.

[0003] However, compared with conventional scenes, the complex remote sensing image scenes and diverse target sizes make remote sensing target detection still challenging. Remote sensing images are generally acquired from a large field of view, where the target objects occupy only a very small proportion of pixels compared to the broad and complex background. Additionally, problems such as insufficient feature information due to low resolution of target objects and poor contrast with the background lead to difficulties in target detection in remote sensing images. At the same time, in terms of optimizing the performance of feature extraction, the unique features of remote sensing images have not been fully explored. Generally, the challenges faced by remote sensing image detection mainly include: 1) How to extract effective features in limited images, establish long-range dependencies, improve feature representation, and reduce the feature loss rate. 2) How to prevent small-scale remote sensing targets from being overwhelmed by background and interference information. Summary of the Invention

[0004] To solve the deficiencies mentioned in the above background art, the purpose of the present invention is to provide a target detection method and system based on frequency-domain adaptive feature enhancement and attention guidance.

[0005] In a first aspect, the purpose of the present invention can be achieved through the following technical solutions: A target detection method based on frequency-domain adaptive feature enhancement and attention guidance, the method comprising the following steps:

[0006] Receive the extracted shallow features and deep features, input the shallow features into a pre-established frequency-domain adaptive feature enhancement module (FAFEM), and output a processed feature layer;

[0007] Input the processed feature layer and the deep features into a pre-established input attention-guided cross-scale feature interaction module, and output a fused feature; input the fused feature into a detection head, and output a target detection result.

[0008] In combination with the first aspect, in some implementation manners of the first aspect, the method further comprises: The process of inputting the shallow features into the pre-established frequency-domain adaptive feature enhancement module (FAFEM) is as follows:

[0009] The low-frequency feature information is obtained through average pooling, and the high-frequency feature information is obtained by subtracting the low-frequency feature information from the shallow features of the original input;

[0010] The low-frequency feature information is subjected to multi-scale edge enhancement to obtain low-frequency features;

[0011] The high-frequency feature information is subjected to high-frequency information enhancement to obtain enhanced high-frequency features;

[0012] The low-frequency features and the enhanced high-frequency features are fused to obtain fused features;

[0013] The fused features are input into the channel attention to suppress the interference of redundant channels, and then added and fused with the shallow features of the original input to finally obtain the processed feature layer.

[0014] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the calculation process of obtaining the low-frequency feature information through average pooling and subtracting the low-frequency feature information from the shallow features of the original input to obtain the high-frequency feature information:

[0015] LFI = H Avgpool (Input)

[0016] HFI = Input - H bilinear (LFI)

[0017] where Input represents the shallow feature input, H Avgpool represents the average pooling operation, LFI represents the obtained low-frequency feature information, H bilinear represents the bilinear interpolation upsampling operation, and HFI represents the high-frequency feature information.

[0018] The calculation process of performing multi-scale edge enhancement on the obtained low-frequency feature information to obtain low-frequency features:

[0019] Lowfeat = H MEEM (LFI).

[0020] where H MEEM represents the multi-scale edge enhancement operation, and Lowfeat represents the obtained low-frequency feature map.

[0021] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the calculation process of performing high-frequency information enhancement on the high-frequency feature information to obtain enhanced high-frequency features:

[0022] Highfeat = H HFIE (HFI)

[0023] where H HFIERepresents the high-frequency information enhancement operation, and Highfeat represents the obtained high-frequency feature map.

[0024] The calculation process of fusing the low-frequency features with the enhanced high-frequency features to obtain the fused features:

[0025] Fusefeat = H concat (Lowfeat, Highfeat).

[0026] Among them, H concat Represents the feature stacking operation, and Fusefeat represents the obtained high-low frequency fused features.

[0027] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the calculation process of inputting the fused features into the channel attention to suppress the interference of redundant channels and then adding and fusing them with the shallow features of the original input:

[0028] Output = Input + H CA (Fusefeat)

[0029] Among them, H CA Represents the channel attention module, and Output represents the final output feature map.

[0030] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the process of inputting the processed feature layer and the deep features into the pre-established input attention-guided cross-scale feature interaction module is as follows:

[0031] Upsample the deep features and fuse them with the shallow features to obtain cross-scale fused features;

[0032] Obtain the feature layer after symmetric interaction by passing the cross-scale fused features through feature symmetric interaction, where a gating mechanism is introduced during the feature symmetric interaction process;

[0033] Perform parallel CBR operations with three different convolutional kernels on the feature layer after symmetric interaction to obtain three parallel branch features;

[0034] Couple the shallow features and the deep features with the three parallel branch features through skip connections to obtain coupled features;

[0035] Dynamically adjust the weights of different channels and spatial positions in the coupled feature map to finally obtain the fused features.

[0036] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the calculation process of upsampling the deep features and fusing them with the shallow features:

[0037] F fuse = Fsf +H bilinear (F df )

[0038] Among them, F sf represents shallow features, F df represents deep features, H bilinear represents the bilinear interpolation upsampling operation, F fuse represents the feature output after the fusion of deep and shallow features.

[0039] The calculation process of obtaining the feature layer after symmetric interaction by performing symmetric interaction on cross-scale fusion features:

[0040] F' fuse = H FSIM (F fuse )

[0041] Among them, H FSIM represents the feature symmetric interaction operation, F' fuse represents the fused features after passing through the feature symmetric interaction module.

[0042] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: The calculation process of obtaining three parallel branch features by performing parallel CBR operations with three different convolution kernels on the feature layer after symmetric interaction:

[0043] F 1 = H 1×1CBR (F' fuse )

[0044] F 2 = H 1×1CBR (F' fuse )

[0045] F 3 = H 1×1CBR (F' fuse )

[0046] Among them, H 1×1CBR represents a 1×1 convolution operation with BatchNomal and ReLu activation functions, F 1 , F 2 , F 3 represents the feature maps obtained through three parallel H 1×1CBR branch operations.

[0047] The calculation process of coupling shallow features and deep features with three parallel branch features through skip connections to obtain coupled features:

[0048] F Add = H Add (F 1 , F2 , F 3 , F sf , H bilinear (F df ))

[0049] Among them, H Add represents the addition feature fusion operation, and F Add represents the fused feature map.

[0050] Combined with the first aspect, in some implementations of the first aspect, the method further includes: the calculation process of dynamically adjusting the weights of different channels and spatial positions in the coupled feature map:

[0051] F out = H Attention (F Add )

[0052] Among them, H Attention represents the attention guidance operation, and F out represents the output feature after attention-guided feature fusion.

[0053] In a second aspect, to achieve the above object, the present invention discloses a frequency-domain adaptive feature enhancement and attention-guided object detection system, including:

[0054] A feature processing module, configured to receive the extracted shallow features and deep features, input the shallow features into a pre-established frequency-domain adaptive feature enhancement module FAFEM, and output a processed feature layer;

[0055] An object detection module, configured to input the processed feature layer and the deep features into a pre-established input attention-guided cross-scale feature interaction module, and output a fused feature; input the fused feature into a detection head, and output an object detection result.

[0056] Advantages of the present invention:

[0057] In terms of feature enhancement, the present invention designs a frequency-domain adaptive feature enhancement module. It respectively performs enhancement operations on the high-frequency and low-frequency information in the feature layer, explicitly models the edge information, realizes multi-scale enhanced edge perception, highlights the key boundaries of objects, and complements each other with high-frequency detail information while capturing low-frequency context.

[0058] In terms of cross-scale feature interaction and fusion, the present invention designs an attention-guided cross-scale feature interaction module. Through effective interaction between symmetric views, it realizes the fusion of feature information between different views. At the same time, cross-scale feature fusion can combine low-level detail and high-level semantic information, enhance the expression ability of features, assist in establishing long-range dependencies of features, and improve the accuracy of feature expression.

[0059] In terms of improving the effective expression and synergy of features and suppressing interference information, a gating mechanism is introduced to dynamically adjust the degree of participation of feature information in the model, coordinate the synergy between features, and enhance the expression ability of fused features. At the same time, the attention module enables the model to focus more on the feature information beneficial to the detection task. Working in coordination with the gating mechanism, it can effectively retain important features selectively and filter redundant features, significantly reducing the dependence on computational overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0061] Figure 1 is a schematic flowchart of the method of the present invention;

[0062] Figure 2 is the overall framework diagram of the model proposed by the present invention;

[0063] Figure 3 is the frequency-domain adaptive feature enhancement module proposed by the present invention;

[0064] Figure 4 is the multi-scale edge enhancement module and high-frequency information enhancement module (components of the frequency-domain adaptive feature enhancement module) proposed by the present invention;

[0065] Figure 5 is the attention-guided cross-scale feature information interaction module proposed by the present invention;

[0066] Figure 6 is the feature symmetry interaction module (components of the attention-guided cross-scale feature information interaction module) proposed by the present invention;

[0067] Figure 7 is the system structure diagram of the present invention;

[0068] Figure 8 is the actual detection effect diagram of the present invention;

[0069] Figure 9 is the quantization diagram of the relevant evaluation indexes of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0071] Embodiment 1:

[0072] As Figure 1 shown, the frequency-domain adaptive feature enhancement and attention-guided object detection method includes the following steps:

[0073] Receive the extracted shallow features and deep features, input the shallow features into the pre-established frequency-domain adaptive feature enhancement module FAFEM, and output the processed feature layer;

[0074] The specific process is to obtain low-frequency feature information through average pooling, and subtract the low-frequency feature information from the original input to obtain high-frequency feature information. The formula is as follows:

[0075] LFI = H Avgpool (Input)

[0076] HFI = Input - H bilinear (LFI)

[0077] where Input represents the shallow feature input, H Avgpool represents the average pooling operation, LFI represents the obtained low-frequency feature information, H bilinear represents the bilinear interpolation upsampling operation, and HFI represents the high-frequency feature information.

[0078] Perform multi-scale edge enhancement on the obtained low-frequency feature information to further obtain low-frequency features, solve the problem of lack of detailed edge features, achieve edge enhancement perception at different scales, highlight the key boundaries of objects, and improve the fineness of low-frequency features. The formula is as follows:

[0079] Lowfeat = H MEEM (LFI)

[0080] where H MEEM represents the multi-scale edge enhancement operation, and Lowfeat represents the obtained low-frequency feature map.

[0081] After obtaining the high-frequency feature information, the reverse bottleneck structure and residual connection can not only improve the feature extraction efficiency but also enhance the high-frequency detail features, and are more prominent when obtaining context information. The formula is as follows:

[0082] Highfeat = H HFIE(HFI)

[0083] Among them, H HFIE represents the high-frequency information enhancement operation, and Highfeat represents the obtained high-frequency feature map.

[0084] Fusing high and low frequency features can not only obtain more extensive spatial context information, but also flexibly adapt to changes in the target scale. The formula is as follows:

[0085] Fusefeat = H concat (Lowfeat, Highfeat)

[0086] Among them, H concat represents the feature stacking operation, and Fusefeat represents the obtained high and low frequency fused features.

[0087] After adjusting the number of feature channels through 1×1 convolution to reduce the computational amount, the fused features are input into the channel attention to further suppress the interference of redundant channels and reduce the computational overhead. Subsequently, they are added and fused with the input features to reduce feature loss.

[0088] Output = Input + H CA (Fusefeat)

[0089] Among them, H CA represents the channel attention module, and Output represents the final output feature map.

[0090] S102: Input the processed feature layer and deep features into the pre-established input attention-guided cross-scale feature interaction module, and output the fused features; input the fused features into the detection head to output the target detection results.

[0091] The specific process is to separately take the shallow feature F sf and the deep feature F df . After the deep feature is upsampled, it is fused with the shallow feature. The formula is as follows:

[0092] F fuse = F sf + H bilinear (F df )

[0093] Among them, F sf represents the shallow feature, F df represents the deep feature, H bilinear represents the bilinear interpolation upsampling operation, and F fuse represents the feature output after fusing the deep and shallow features.

[0094] After obtaining the cross-scale fusion features, cross-view information fusion is achieved through symmetric view interaction, combining the features of the left and right views to improve the accuracy of feature expression. At the same time, a gating mechanism (in FSIM) is introduced to dynamically adjust the degree of participation of feature information in the model, coordinate the synergy between features, and enhance the expression ability of the fusion features. At the same time, the formula is as follows:

[0095] F' fuse = H FSIM (F fuse )

[0096] Among them, H FSIM represents the feature symmetric interaction operation, and F' fuse represents the fusion feature after passing through the feature symmetric interaction module.

[0097] Considering the problems of limited receptive field and long-range dependence, the feature layer after symmetric interaction is subjected to parallel CBR operations with three different convolutional kernels, and the receptive field is expanded to help the model better utilize the context information and establish long-range dependence to capture global information, thereby improving the ability to recognize and understand the target. The formula is as follows:

[0098] F 1 = H 1×1CBR (F' fuse )

[0099] F 2 = H 1×1CBR (F' fuse )

[0100] F 3 = H 1×1CBR (F' fuse )

[0101] Among them, H 1×1CBR represents the 1×1 convolutional operation with BatchNomal and ReLu activation functions, and F 1 , F 2 , F 3 represent the feature maps obtained through three parallel H 1×1CBR branch operations.

[0102] During the cross-scale feature fusion process, feature loss is inevitable. To alleviate the phenomenon of feature loss, the initial shallow features and deep features are coupled with the three parallel branches in step 203 through skip connections to generate richer and more discriminative feature representations. The formula is as follows:

[0103] F Add = H Add (F 1 , F 2 , F 3 , Fsf , H bilinear (F df ))

[0104] Among them, H Add represents the addition feature fusion operation, and F Add represents the feature map after fusion.

[0105] The attention module can dynamically adjust the weights of different channels and spatial positions in the feature map, enabling the model to focus more on the feature information beneficial to the detection task. Working in coordination with the gating mechanism can effectively selectively retain important features and filter redundant features, significantly reducing the dependence on computational overhead. The formula is as follows:

[0106] F out = H Attention (F Add )

[0107] Among them, H Attention represents the attention guidance operation, and F out represents the output feature after attention-guided feature fusion.

[0108] Embodiment 2: Second aspect, as Figure 7 shown, to achieve the above object, the present invention discloses a frequency-domain adaptive feature enhancement and attention-guided object detection system, including:

[0109] A feature processing module 11, configured to receive the extracted shallow features and deep features, input the shallow features into a pre-established frequency-domain adaptive feature enhancement module FAFEM, and output a processed feature layer;

[0110] An object detection module 12, configured to input the processed feature layer and the deep features into a pre-established input attention-guided cross-scale feature interaction module, and output a fused feature; input the fused feature into a detection head, and output an object detection result.

[0111] As Figure 8 shown, the present invention discloses a frequency-domain adaptive feature enhancement and attention-guided object detection system, including:

[0112] Visual qualitative experiment: By performing visual experiments on different types of objects in different scenarios in the dataset, the actual detection effect of the frequency-domain adaptive feature enhancement and attention-guided object detection system disclosed by the present invention is verified.

[0113] As Figure 9 shown, the present invention discloses a frequency-domain adaptive feature enhancement and attention-guided object detection system, including:

[0114] Quantification of evaluation metrics: Intuitively verify the reliability of the frequency-domain adaptive feature enhancement and attention-guided object detection system disclosed in the present invention during the actual detection process through data.

[0115] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically used to load and execute one or more instructions in the computer storage medium to implement the above method.

[0116] It should be further noted that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored, and the computer program, when run by a processor, executes the above method. The storage medium may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or combined with an instruction execution system, apparatus, or device.

[0117] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0118] The foregoing has shown and described the basic principles, main features and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements fall within the scope of the present disclosure claimed.

Claims

1. Frequency domain adaptive feature enhancement and attention-guided target detection method, characterized in that: The method comprises the following steps: Receive the extracted shallow features and deep features, input the shallow features into a pre-established frequency domain adaptive feature enhancement module FAFEM, and output the processed feature layer; The processed feature layer and deep features are input into the pre-established input attention-guided cross-scale feature interaction module, and the fused features are output; The fused features are input into the detection head and the target detection results are output.

2. The target detection method with frequency domain adaptive feature enhancement and attention guidance according to claim 1, characterized in that: The process of inputting the shallow features into the pre-established frequency domain adaptive feature enhancement module FAFEM is as follows: The low-frequency feature information is obtained by average pooling, and the high-frequency feature information is obtained by subtracting the shallow features of the original input from the low-frequency feature information; Perform multi-scale edge enhancement on low-frequency feature information to obtain low-frequency features; Performing high-frequency information enhancement on high-frequency feature information to obtain enhanced high-frequency features; Fuse the low-frequency features with the enhanced high-frequency features to obtain fused features; The fused feature input channel attention is used to suppress the interference of redundant channels, and then it is added and fused with the shallow features of the original input to finally obtain the processed feature layer.

3. The target detection method with frequency domain adaptive feature enhancement and attention guidance according to claim 2, characterized in that: The calculation process of obtaining low-frequency feature information by average pooling and subtracting the low-frequency feature information from the shallow features of the original input to obtain high-frequency feature information is as follows: LFI=H Avgpool (Input) HFI / Input-H bilinear (LFI) Among them, Input represents the shallow feature input, H Avgpool represents the average pooling operation, LFI represents the acquired low-frequency feature information, and H bilinear represents the bilinear interpolation upsampling operation, and HFI represents high-frequency feature information; The calculation process of obtaining low-frequency features by performing multi-scale edge enhancement on the acquired low-frequency feature information is as follows: Low feat=H MEEM (LFI) Among them, H MEEM represents the multi-scale edge enhancement operation, and Lowfeat represents the obtained low-frequency feature map.

4. The target detection method with frequency domain adaptive feature enhancement and attention guidance according to claim 3, characterized in that: The high-frequency feature information is enhanced to obtain the calculation process of the enhanced high-frequency feature: Highfeat=H HFIE (HFI) Among them, H HFIE Represents the high-frequency information enhancement operation, and Highfeat represents the obtained high-frequency feature map; The calculation process of fusing low-frequency features with enhanced high-frequency features to obtain the fused features: Fusefeat=H concat (Lowfeat,Highfeat) Among them, H concat Represents the feature stacking operation, and Fusefeat represents the obtained high- and low-frequency fusion features.

5. The target detection method with frequency domain adaptive feature enhancement and attention guidance according to claim 4, characterized in that: The calculation process of inputting the fused features into the channel attention to suppress the interference of redundant channels, and then adding and fusing them with the shallow features of the original input: Output=Input+H CA (Fusefeat) Among them, H CA represents the channel attention module, and Output represents the final output feature map.

6. The target detection method with frequency domain adaptive feature enhancement and attention guidance according to claim 1, characterized in that: The process of inputting the processed feature layer and deep features into the pre-established input attention guided cross-scale feature interaction module is as follows: The deep features are upsampled and fused with the shallow features to obtain cross-scale fusion features; The cross-scale fusion features are symmetrically interacted to obtain the feature layer after symmetrical interaction, wherein a gating mechanism is introduced in the process of symmetrical interaction of features; The feature layer after symmetrical interaction is subjected to parallel CBR operations of three different convolution kernels to obtain three parallel branch features; The shallow features and deep features are coupled with three parallel branch features through skip connections to obtain coupled features; Dynamically adjust the weights of different channels and spatial positions in the coupled feature map to finally obtain the fused features.

7. The target detection method with frequency domain adaptive feature enhancement and attention guidance according to claim 6, characterized in that: The calculation process of fusing the deep features with the shallow features after upsampling is as follows: F fuse =F sf +H bilinear (F df ) Among them, F sf represents shallow features, F df represents the deep features, H bilinear represents the bilinear interpolation upsampling operation, F fuse Represents the feature output after deep and shallow fusion; The calculation process of obtaining the feature layer after symmetric interaction by symmetric interaction of cross-scale fusion features: F’ fuse =H FSIM (F fuse ) Among them, H FSIM represents the feature symmetric interaction operation, F f ' use Represents the fused features after the feature symmetry interaction module.

8. The target detection method with frequency domain adaptive feature enhancement and attention guidance according to claim 7, characterized in that: The feature layer after symmetrical interaction is subjected to parallel CBR operation of three different convolution kernels to obtain the calculation process of three parallel branch features: F1=H 1×1CBR (F’ fuse ) F2=H 1×1CBR (F’ fuse ) <h2 style=";text-align:left;direction:ltr">F3=H<h2 style=";text-align:left;direction:ltr"> 1×1CBR <h2 style=";text-align:left;direction:ltr"> (F'<h2 style=";text-align:left;direction:ltr"> fuse <h2 style=";text-align:left;direction:ltr"> ) Among them, H 1×1CBR represents a 1×1 convolution operation with BatchNomal and ReLu activation function, and F1, F2, and F3 represent the convolution operation through three parallel H 1×1CBR Feature map obtained by branch operation; The shallow features and deep features are coupled with three parallel branch features through skip connections to obtain the calculation process of the coupled features: F Add =H Add (F1,F2,F3,F sf ,H bilinear (F df )) Among them, H Add represents the addition feature fusion operation, F Add Represents the fused feature map.

9. The target detection method with frequency domain adaptive feature enhancement and attention guidance according to claim 8, characterized in that: The calculation process of dynamically adjusting the weights of different channels and spatial positions in the coupling feature map: F out =H Attention (F Add ) Among them, H Attention represents the attention-guiding operation, F out Represents the output features after attention-guided feature fusion.

10. Frequency domain adaptive feature enhancement and attention guided target detection system, characterized in that: include: A feature processing module is used to receive the extracted shallow features and deep features, input the shallow features into a pre-established frequency domain adaptive feature enhancement module FAFEM, and output the processed feature layer; The object detection module is used to input the processed feature layer and deep features into the pre-established input attention-guided cross-scale feature interaction module, and output the fused features; The fused features are input into the detection head and the target detection results are output.

Citation Information

Cited By

  • Air-ground cooperative target detection and feature analysis system and method

    CN120318501A

  • An air-ground cooperative target detection and feature analysis system and method

    CN120318501B

  • Aerial photography small target detection method for complex background and dynamic interference

    CN120510538A

  • A small target detection method for aerial photography with complex background and dynamic interference

    CN120510538B