Small target detection method based on feature level fusion and frequency domain enhancement

By introducing a multi-scale feature aggregation module, a frequency domain dual attention mechanism and a P2-level detection layer in the small object detection method, the problems of insufficient feature representation and large background interference in the existing small object detection method are solved, and the detection accuracy and robustness are significantly improved.

CN120125809APending Publication Date: 2025-06-10SOUTHEAST UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510287780.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing small object detection methods perform poorly in terms of detection accuracy and robustness, mainly due to insufficient feature representation, large background interference and low detection accuracy.

Method used

A small object detection method based on feature hierarchy fusion and frequency domain enhancement is designed. The target features of different scales are captured through the multi-scale feature aggregation module, a dual attention mechanism for frequency domain is introduced to process high-frequency and low-frequency information, and a P2-level detection layer is specially designed to enhance the sensitivity to small objects.

Benefits of technology

It significantly improves the detection accuracy and robustness of small targets, especially when dealing with small-size, high-density and unevenly distributed target detection tasks in scenarios such as drone aerial photography.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125809A_ABST
    Figure CN120125809A_ABST
Patent Text Reader

Abstract

The invention discloses a small target detection method based on feature level fusion and frequency domain enhancement, and belongs to the field of target detection. According to the method, in the aspect of feature extraction, a multi-scale feature aggregation module is designed, and the module effectively captures target features of different scales through multi-scale partial convolution; a frequency domain double attention mechanism is provided, comparative processing of high-frequency information and low-frequency information is achieved through frequency domain decomposition, and a double attention mode is formed; and a detection layer specially aiming at a small target is introduced, so that the sensitivity of the model to the small target is enhanced. According to the end-to-end small target detection framework constructed by the invention, the calculation efficiency is maintained, and the detection precision and robustness of the small target are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target recognition, and particularly relates to a small target detection method based on feature hierarchical fusion and frequency domain enhancement. Background Art

[0002] Target detection is a core task in the field of computer vision and is widely used in application scenarios such as target tracking, instance segmentation, and scene understanding. In recent years, with the rapid development of deep convolutional neural networks (CNNs), significant progress has been made in target detection technology. Deep learning-based target detection techniques are mainly divided into two categories: two-stage models and one-stage models. Two-stage models such as Faster R-CNN, SPP-Net, and Feature Pyramid Network (FPN) models generate regions of interest (ROIs) in the first stage and then fine-tune the ROIs in the second stage to classify the targets and accurately locate the bounding boxes; one-stage models such as YOLO, SSD, and some anchor-free models directly classify and locate the targets from the feature map without the ROI stage, with faster detection speeds.

[0003] Small object detection (SOD), as an important branch of target detection, has received increasing attention in recent years, and its application scenarios include medical image analysis, surveillance systems, and drone vision. Small objects are usually defined according to their relative size (the proportion of the target bounding box in the image is lower than a threshold) or absolute size (e.g., less than 32×32 pixels in the COCO dataset).

[0004] Although target detection methods have made significant progress on large datasets such as MS COCO, ImageNet, and PASCAL VOC, their performance in small object detection is still not ideal. Taking the latest Co-DETR method as an example, its mean average precision (mAP) for small objects on the COCO dataset is only 48.4%, significantly lower than 67.1% and 77.3% for medium and large-sized objects. The main reasons for the poor performance of small object detection include: 1) low resolution of small objects, occupying fewer pixels than larger objects; 2) loss of spatial location information caused by downsampling and pooling operations in the convolutional network, making it more difficult for the detection head to locate small objects; 3) a large lack of small object datasets, and existing datasets mainly focus on specific scenarios such as faces, pedestrians, and traffic scenes, restricting the development of general small object detection models. Summary of the Invention

[0005] The present invention provides a small target detection method based on feature hierarchical fusion and frequency domain enhancement. In terms of feature extraction, a multi-scale feature aggregation module is designed, which effectively captures target features of different scales through multi-scale partial convolution; a frequency domain dual attention mechanism is proposed, which realizes contrastive processing of high-frequency and low-frequency information through frequency domain decomposition to form a dual attention pattern; a detection layer specifically for small targets is introduced to enhance the sensitivity of the model to tiny targets. The end-to-end small target detection framework constructed by the present invention significantly improves the detection accuracy and robustness of small targets while maintaining computational efficiency.

[0006] An embodiment of the present invention provides a small target detection method based on feature hierarchical fusion and frequency domain enhancement, including the following steps: obtaining a small target image dataset to be detected; inputting the small target image dataset into a pre-trained small target detection network for target detection; wherein, the small target detection network includes a backbone network, an efficient hybrid encoder, a query selection module, a decoder and a detection head. The backbone network includes a multi-scale feature aggregation module that captures target features of different scales through multi-scale partial convolution. The efficient hybrid encoder includes a frequency domain dual attention mechanism that decomposes target features into high-frequency information and low-frequency information through frequency domain decomposition and applies different attention processing to form a dual attention pattern. The detection head includes a P2-level detection layer. Input the four-layer feature maps with different resolutions output by passing the picture through the backbone network into the encoder for feature fusion, and then perform target detection through the query selection module, the decoder and the detection head to output the category and position of small targets in the image.

[0007] Optionally, in an embodiment of the present invention, the multi-scale feature aggregation module includes:

[0008] Multi-scale partial convolution, which is used to enhance the ability to capture features of different scales;

[0009] Cross-stage partial feature aggregation, which is used to split the input features into two parts, process them separately and then fuse them to aggregate and fuse features on different processing paths;

[0010] Channel grouping processing strategy, which is used to adopt a strategy combining grouped convolution and channel splitting in convolution operations of different scales to enhance the feature representation ability while maintaining computational efficiency.

[0011] Optionally, in an embodiment of the present invention, the FLOPs calculation formula of multi-scale partial convolution is:

[0012]

[0013] wherein, H is the height of the feature map, W is the width of the feature map, C p is the input channel number of partial convolution, k1 , k 2 , k 3 respectively represent the convolutional kernel sizes of three different scales.

[0014] Optionally, in an embodiment of the present invention, the frequency-domain dual attention mechanism decomposes the input features into frequency-domain representations through Haar wavelet transform, separates foreground objects from background information, so that the high-frequency components contain the edges and detail information of foreground small objects, and the low-frequency components retain the semantic structure information of the background.

[0015] Optionally, in an embodiment of the present invention, the mathematical expression of the frequency-domain dual attention mechanism is:

[0016]

[0017] where i and j are spatial position coordinates respectively, fg is the foreground, bg is the background, is the attention weight of the foreground feature, is the attention weight of the background feature, is the eigenvalue vector at position (i, j), is the convolution operation.

[0018] Optionally, in an embodiment of the present invention, the P2-level detection layer is constructed as follows:

[0019] F P2 = Conv 1×1 (F CMSP2 ) + Upsample(F P3 )

[0020] where F CMSP2 represents the output feature of the second stage of the backbone network, F P3 is the feature of the original P3 layer, Upsample(F P3 ) is the upsampling result of the P3 layer feature, and Conv 1×1 (F CMSP2 ) is the result of 1×1 convolution processing on the output feature of the second stage of the backbone network.

[0021] Optionally, in an embodiment of the present invention, the query selection module is used to extract key queries from the features generated by the efficient hybrid encoder, and preferentially select query points containing target information through an adaptive selection mechanism.

[0022] Optionally, in an embodiment of the present invention, the decoder includes multiple Transformer decoder modules, decodes the output of the efficient hybrid encoder, and extracts target candidate boxes and class features.

[0023] Optionally, in an embodiment of the present invention, the detection head obtains category information through a fully connected network, obtains candidate box information through a feedforward neural network, and converts the candidate box and category features of the target to be detected into a standard category confidence and candidate box format.

[0024] Optionally, in an embodiment of the present invention, the GIOU loss function is used as the bounding box regression loss during the training of the small target detection network.

[0025] The small target detection method based on feature hierarchical fusion and frequency domain enhancement in the embodiment of the present invention effectively solves the problems of insufficient feature representation, large background interference, and low detection accuracy in the existing small target detection methods through the design of a multi-scale feature aggregation module, a frequency domain dual attention mechanism, and a dedicated small target detection layer. While maintaining the computational efficiency, it significantly improves the detection accuracy and robustness of small targets.

[0026] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings

[0027] The above-mentioned and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:

[0028] Figure 1 FIG. is a flowchart of a small target detection method based on feature hierarchical fusion and frequency domain enhancement according to an embodiment of the present invention;

[0029] Figure 2 FIG. is a schematic diagram of the small target detection network architecture according to an embodiment of the present invention;

[0030] Figure 3 FIG. is a schematic diagram of the structure of the multi-scale aggregation module according to an embodiment of the present invention;

[0031] Figure 4 FIG. is a schematic diagram of the partial convolution principle according to an embodiment of the present invention;

[0032] Figure 5 FIG. is a schematic diagram of the frequency domain dual attention mechanism according to an embodiment of the present invention;

[0033] Figure 6 FIG. is a schematic diagram of the frequency domain decomposition of the two-dimensional Haar wavelet transform according to an embodiment of the present invention;

[0034] Figure 7 FIG. is a comparison chart of the mAP metrics of the method according to an embodiment of the present invention and the original method. Detailed Embodiments

[0035] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.

[0036] As Figure 1 and Figure 2 shown, the small target detection method based on feature hierarchical fusion and frequency domain enhancement includes the following steps:

[0037] Step S101, obtain a small target image data set to be detected.

[0038] Prepare a small target image data set. The data set uses the VisDrone data set and is divided into a training set, a validation set, and a test set according to a ratio of 7:2:1. The small target image data set is a large-scale unmanned aerial vehicle (UAV) aerial photography vision data set created by the Chinese UAV Vision Team, covering 10 categories of objects, and is particularly suitable for small target detection tasks. This data set contains a total of 10,155 images. The objects in the images are small in scale, high in density, and unevenly distributed, presenting high challenges.

[0039] Step S102, input the small target image data set into a pre-trained small target detection network for target detection; wherein, the small target detection network includes a backbone network, an efficient hybrid encoder, a query selection module, a decoder, and a detection head. The backbone network includes a multi-scale feature aggregation module that captures target features at different scales through multi-scale partial convolution. The efficient hybrid encoder includes a frequency domain dual attention mechanism that decomposes the target features into high-frequency information and low-frequency information through frequency domain decomposition and applies different attention processing to form a dual attention pattern. The detection head includes a P2-level detection layer. Input the four-layer feature maps with different resolutions output by the backbone network for the picture into the encoder for feature fusion, and then perform target detection through the query selection module, the decoder, and the detection head to output the categories and positions of small targets in the image.

[0040] In the embodiments of the present invention, a small target detection network is constructed, and the above small target detection network is trained using the training set and the validation set. The test set is input into the trained network to detect the targets in the pictures.

[0041] The small object detection network of the embodiments of the present invention mainly includes a backbone network, an efficient hybrid encoder, a query selection module, a decoder, and a detection head. The backbone network includes a multi-scale feature aggregation module (Cross-Stage Partial Multi-Scale Feature Aggregation, CMSP), which captures target features of different scales through multi-scale partial convolution. The efficient hybrid encoder includes a frequency-domain dual attention mechanism (Frequency-Domain Dual Attention Mechanism, FDAM), which decomposes the target features into high-frequency information and low-frequency information through frequency-domain decomposition and applies different attention processing to form a dual attention pattern. The detection head includes a P2-level detection layer.

[0042] The image is first processed by the backbone network to output four feature maps with different resolutions, and then undergoes feature fusion processing by the encoder. The backbone network of the present invention includes the multi-scale feature aggregation module CMSP, and the neck network includes the frequency-domain dual attention mechanism FDAM. Then, via the query selection module, it helps the model reduce uncertainty during object detection. The decoder processes the queries and image features generated by the encoder, and finally performs object detection through the "head" to output the category and location of the object.

[0043] In an embodiment of the present invention, the backbone network includes a multi-scale feature aggregation module, as Figure 3 shown. This module effectively captures target features of different scales through multi-scale partial convolution. Figure 3 The upper part shows the overall architecture of the CMSP module, and the lower part details the internal structure of the MSP-Block. This module effectively captures target features of different scales through multi-scale partial convolution.

[0044] The main innovations of the CMSP module include: 1) multi-scale partial convolution, which extends the single 3×3 PConv to a hierarchical partial convolution structure including 3×3, 5×5, and 7×7 kernel sizes, enhancing the model's ability to capture features of different scales; 2) cross-stage partial feature aggregation, adopting the CSP (Cross-Stage Partial) design concept to more effectively aggregate and fuse features on different processing paths; 3) channel grouping processing strategy, adopting a strategy combining grouped convolution and channel splitting in convolution operations of different scales to enhance the feature representation ability while maintaining computational efficiency.

[0045] Figure 4Shows the principle of partial convolution (PConv). The mathematical expression of PConv is shown in the lower left corner, and the structural details of the FasterNet Block are shown in the lower right corner. Different from traditional convolution, PConv applies the convolution operation only on some channels of the input feature map while keeping the remaining channels unchanged, thus significantly reducing the computational complexity.

[0046] Considering the FLOPs calculation of standard convolution and partial convolution, for standard convolution, the FLOPs can be expressed as:

[0047] FLOPs 标准 = H × W × k 2 × C 2

[0048] where H is the height of the feature map, W is the width of the feature map, C is the number of input channels of the convolution, and k is the convolution kernel size.

[0049] For the partial convolution strategy, the FLOPs is:

[0050]

[0051] where C p is the number of input channels of the partial convolution, k 1 = 3, k 2 = 5, k 3 = 7, respectively represent the convolution kernel sizes of three different scales.

[0052] This design significantly reduces the computational complexity. Theoretically, while maintaining the expressive ability, the computational amount is reduced by about 75%.

[0053] The above neck network contains the frequency-domain dual attention mechanism FDAM, as Figure 5 shown. This mechanism decomposes the features into two parts (high frequency and low frequency) through frequency-domain decomposition, and then applies different attention processing respectively, achieving efficient enhancement of small target features. The FDAM module decomposes the input features into frequency-domain representations through the Haar wavelet transform, effectively separating the foreground target and background information. This decomposition makes the high-frequency components mainly contain the edge and detail information of the foreground small targets, while the low-frequency components retain the semantic structure information of the background. Based on this decomposition, the dual effects of foreground enhancement and background suppression are achieved.

[0054] Specifically, as Figure 6 shown, the FDAM module decomposes the input features into a low-frequency subband and a high-frequency subband through the Haar wavelet transform. Given the input feature map X ∈ R C×H×W , the two-dimensional Haar wavelet transform decomposes it into four subbands:

[0055] {X LL ,XLH , X HL , X HH} = HWT(X)

[0056] Among them, is the low-frequency subband (approximate component), X LH , X HL and X HH are high-frequency subbands (corresponding to horizontal, vertical, and diagonal details respectively). In FDAM, X LL mainly contains background semantic information, while {X LH , X HL , X HH} contains key details such as the edges and textures of foreground objects.

[0057] The mathematical expression of the entire FDAM module can be summarized as:

[0058]

[0059] Among them, i and j are spatial position coordinates, fg is the foreground, bg is the background, is the attention weight of foreground features, is the attention weight of background features, V Δi,j is the eigenvalue vector at position (i, j), is the convolution operation.

[0060] The above network introduces a P2-level detection layer dedicated to small object detection. The traditional RT-DETR model uses three detection layers, P3, P4, and P5. The highest-resolution P3 layer still has an 8-fold downsampling rate, which is not fine enough for small object detection. The feature map size of the P2 detection layer of the present invention reaches 160×160×64, and the downsampling ratio is only 4 times, significantly improving the spatial resolution of the feature map, enabling small objects to have sufficient pixel representation on the feature map. The spatial resolution is greatly improved, and the detection ability for tiny objects is significantly enhanced. This layer is constructed in the following way:

[0061] F P2 = Conv 1×1 (F CMSP2 ) + Upsample(F P3 )

[0062] Among them, F CMSP2 represents the output feature of the second stage of the backbone network, F P3 is the feature of the original P3 layer, Upsample(F P3 ) is the upsampling result of the P3 layer feature, and Conv 1×1 (F CMSP2 ) is the result of 1×1 convolution processing of the output feature of the second stage of the backbone network.

[0063] The query selection module of the embodiment of the present invention is used to extract key queries from the features generated by the encoder. Through an adaptive selection mechanism, this module preferentially selects query points containing target information, reduces the uncertainty processed by the decoder, and improves the accuracy of small target detection.

[0064] The decoder network of the embodiment of the present invention includes multiple Transformer decoder modules, which decode the output of the encoder to extract target candidate boxes and class features.

[0065] The detection head of the embodiment of the present invention obtains class information through a fully connected network and candidate box information through a feed-forward neural network, and converts the target candidate boxes and class features to be detected into standard class confidence and candidate box formats.

[0066] When training the small target detection network of the embodiment of the present invention, the GIOU loss function is used as the bounding box regression loss. This loss function considers the area ratio of the overlapping area and the closure area between the predicted box and the ground truth box. Compared with the traditional IOU loss function, it can better handle non-overlapping situations and improve the bounding box localization accuracy. In addition, the cross-entropy loss is used as the classification loss, and the L1 loss is used as the auxiliary loss for bounding box regression to ensure the stability and convergence of training.

[0067] During training, the trained small target detection network is used to verify the accuracy of the small target detection network using the divided test set.

[0068] As Figure 7 shown, compared with the baseline method RT-DETR-r18, the present invention has achieved significant performance improvement in small target detection. Under the same number of training rounds, the mAP@0.5 value of the present invention is significantly higher than that of the baseline method, and as the training progresses, the performance advantage becomes more obvious. When finally converging, the method in the present invention has increased by 5%.

[0069] The small target detection method based on feature hierarchical fusion and frequency domain enhancement proposed according to the embodiments of the present invention integrates core components such as a backbone network, an efficient hybrid encoder, a query selection module, a decoder, and a detection head. Compared with the original method, in terms of feature extraction, this method designs a multi-scale feature aggregation module, which effectively captures target features of different scales through multi-scale partial convolution; at the same time, a frequency domain dual attention mechanism is proposed, which realizes the contrastive processing of high-frequency and low-frequency information through frequency domain decomposition to form a dual attention mode; in addition, the present invention specifically introduces a P2-level detection layer for small target detection, which significantly enhances the detection sensitivity of the model to tiny targets. Finally, the present invention constitutes an end-to-end small target detection framework, which significantly improves the detection accuracy and robustness of small targets while maintaining computational efficiency, especially performing excellently in dealing with small-size, high-density, and unevenly distributed target detection tasks in scenarios such as drone aerial photography. Experimental results show that this method significantly improves the detection accuracy of small targets while maintaining computational efficiency on public datasets.

[0070] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or N embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0071] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0072] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or N executable instructions for implementing a customized logical function or process, and the scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art of the embodiments of the present invention.

Claims

1. A small target detection method based on feature level fusion and frequency domain enhancement, characterized in that: The following steps are involved: Obtain a dataset of small target images to be detected; The small target image data set is input into a pre-trained small target detection network for target detection; wherein, the small target detection network includes a backbone network, an efficient hybrid encoder, a query selection module, a decoder and a detection head, the backbone network includes a multi-scale feature aggregation module, which captures target features of different scales through multi-scale partial convolution, the efficient hybrid encoder includes a frequency domain dual attention mechanism, which decomposes the target features into high-frequency information and low-frequency information through frequency domain decomposition, and applies different attention processing to form a dual attention mode, and the detection head includes a P2-level detection layer; the four-layer feature maps of different resolutions obtained by the output of the backbone network are input into the encoder for feature fusion, and then target detection is performed through the query selection module, the decoder and the detection head, and the category and position of the small target in the image are output.

2. The method according to claim 1, characterized in that The multi-scale feature aggregation module includes: Multi-scale partial convolution is used to enhance the ability to capture features of different scales; Cross-stage partial feature aggregation is used to split the input features into two parts and process them separately and then fuse them to aggregate and fuse features on different processing paths; The channel grouping processing strategy is used to adopt a strategy combining grouped convolution and channel segmentation in convolution operations of different scales to maintain computational efficiency while enhancing feature representation capabilities.

3. The method according to claim 2, characterized in that The FLOPs calculation formula for multi-scale partial convolution is: Among them, H is the height of the feature map, W is the width of the feature map, and C p is the number of input channels of the partial convolution, k1, k2, and k3 represent the convolution kernel sizes of three different scales.

4. The method according to claim 1, characterized in that The frequency domain dual attention mechanism decomposes the input features into frequency domain representation through Haar wavelet transform, separates the foreground target and the background information, so that the high-frequency components contain the edge and detail information of the small foreground target, and the low-frequency components retain the semantic structure information of the background.

5. The method according to claim 4, characterized in that The mathematical expression of the frequency domain dual attention mechanism is: Among them, i and j are spatial position coordinates, fg is foreground, bg is background, is the attention weight of the foreground feature, is the attention weight of the background feature, is the eigenvalue vector at position (i, j), is the convolution operation.

6. The method according to claim 1, characterized in that The P2 level detection layer is constructed as follows: F P2 =Conv 1×1 (F CMSP2 )+Upsample(F P3 ) Among them, F CMSP2 represents the output features of the second stage of the backbone network, F P3 is the original P3 layer feature, Upsample(F P3 ) is the upsampling result of the P3 layer feature, Conv 1×1 (F CMSP2 ) is the result of 1×1 convolution processing on the output features of the second stage of the backbone network.

7. The method according to claim 1, characterized in that The query selection module is used to extract key queries from the features generated by the efficient hybrid encoder, and preferentially select query points containing target information through an adaptive selection mechanism.

8. The method according to claim 1, characterized in that The decoder includes multiple Transformer decoder modules, which decode the output of the efficient hybrid encoder to extract target candidate boxes and category features.

9. The method according to claim 1, characterized in that: The detection head obtains category information through a fully connected network, obtains candidate box information through a feedforward neural network, and converts the candidate box and category features of the target to be detected into a standard category confidence and candidate box format.

10. The method according to claim 1, characterized in that The GIOU loss function is used as the bounding box regression loss during the training of the small object detection network.

Citation Information

Cited By

  • Unmanned aerial vehicle target detection method based on cross-spatial frequency domain and electronic equipment

    CN120318499A

  • Unmanned aerial vehicle lightweight real-time small target recognition device and method based on SOD-DETR

    CN120747791A