CoMFE-based small target detection method, apparatus and device, and medium

By introducing coordinate attention modules and multi-branch feature enhancement modules into the lightweight backbone network, the problem of low detection accuracy of small targets from the perspective of the drone is solved, and efficient small target detection effect is achieved.

CN120495629APending Publication Date: 2025-08-15NANJING QUANSHI TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510574689.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the detection of small targets from the perspective of drone, the prior art has problems such as low detection accuracy and missed detection and false detection, especially in complex environments, it is difficult to effectively utilize the spatial details of small targets.

Method used

The coordinate attention module and multi-branch feature enhancement module are introduced in the lightweight backbone network. Through multi-feature extraction and fusion, the spatial details and context information of small targets are enhanced, and the redundant prediction box is removed in combination with a non-maximum suppression algorithm.

Benefits of technology

It significantly improves the accuracy and robustness of small object detection, reduces the number of model parameters, is suitable for different versions of YOLO networks, and has good deployment flexibility and engineering application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495629A_ABST
    Figure CN120495629A_ABST
Patent Text Reader

Abstract

The invention discloses a CoMFE-based small target detection method, apparatus and device, and a medium. The method comprises the steps of obtaining an original image; an original image is input into a pre-trained small target detection model, a detection result is output, in the small target detection model, a backbone network is a backbone network based on a coordinate multi-branch feature enhancement module, namely a CoMFE module, a coordinate attention module and the multi-branch feature enhancement module are introduced into the backbone network, and the detection result is output; the plurality of coordinate multi-branch feature enhancement modules are used for performing multi-feature extraction output on the original image, the output of each coordinate multi-branch feature enhancement module and the final output of the backbone network are connected to a neck network, the neck network is used for fusing different scale features output by the backbone network and then outputting fused features, and the neck network is connected with a plurality of detection heads. Each detection head is used for performing target detection on the fusion features and outputting a detection result; and removing a redundant prediction frame from a detection result by adopting a non-maximum suppression algorithm to obtain a final small target detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and deep learning technology, and in particular to an improved target detection method for small target detection tasks, specifically a small target detection method, device, equipment and medium based on CoMFE. Background Art

[0002] With the widespread application of drone technology in security monitoring, intelligent transportation, and agricultural inspections, small target detection from a drone's perspective is facing increasingly higher demands for precision and real-time performance. However, due to their small size, low pixel ratio in the image, and high density distribution from the drone's perspective, small targets are easily lost during downsampling or feature fusion. Furthermore, small targets in real-world scenarios may be subject to complex environmental factors such as mutual occlusion, changes in the sampling device's posture, and uneven lighting. Traditional target detection methods often suffer from missed and false detections, seriously impacting the reliability and security of the system.

[0003] To improve the detection model's performance for small objects, patent publication number CN117523267A discloses a small object detection system and method based on an improved YOLOv5. This method utilizes a coordinate attention mechanism for small object detection. However, simple channel or spatial attention mechanisms cannot account for the fine-grained position information of small objects in the spatial dimension, resulting in some fine-grained features still not being effectively utilized. Therefore, how to improve the detection model's performance for small objects while minimizing the number of model parameters remains an urgent problem. Summary of the Invention

[0004] Technical Objective: To address these needs and challenges, this paper proposes a small target detection method, apparatus, device, and medium based on CoMFE. By introducing a coordinate attention module and a multi-branch feature enhancement module into a lightweight backbone network, this method effectively preserves and enhances the spatial details and contextual information of small targets without significantly increasing the number of parameters, thereby significantly improving overall detection accuracy and robustness. This method can be used as a module for a wide range of target detection network frameworks.

[0005] Technical solution: To achieve the above technical objectives, the present invention proposes a small target detection method based on CoMFE, which includes:

[0006] Get the original image containing multiple small targets;

[0007] The original image is input into the pre-trained small target detection model, and the detection result is output, which includes the category, confidence and boundary information of the prediction box of each detected target; the small target detection model includes a backbone network, a neck network and a detection head connected in sequence, wherein the backbone network is a backbone network based on a coordinate multi-branch feature enhancement module, i.e., a CoMFE module. The coordinate attention module and the multi-branch feature enhancement module are introduced into the backbone network to perform multi-feature extraction and output on the original image. The output of each coordinate multi-branch feature enhancement module and the final output of the backbone network are connected to the neck network. The neck network is used to fuse the different scale features output by the backbone network and output the fused features. The neck network is connected to several detection heads, each detection head is used to perform target detection on the fused features and output the detection results.

[0008] The non-maximum suppression algorithm is used to remove redundant prediction boxes from the detection results to obtain the final small target detection results.

[0009] On the other hand, the present application also discloses a small target detection device based on CoMFE, comprising:

[0010] An image acquisition module is used to acquire an original image containing multiple small targets;

[0011] The model processing module is connected to the image acquisition module and is used to input the original image into the pre-trained small target detection model and output the detection results, which include the category, confidence and boundary information of the prediction box of each detected target; the small target detection model includes a backbone network, a neck network and a detection head connected in sequence, wherein the backbone network is a backbone network based on the coordinate multi-branch feature enhancement module, i.e., the CoMFE module. The coordinate attention module and the multi-branch feature enhancement module are introduced into the backbone network to perform multi-feature extraction and output on the original image. The output of each coordinate multi-branch feature enhancement module and the final output of the backbone network are connected to the neck network. The neck network is used to fuse the different scale features output by the backbone network and output the fused features. The neck network is connected to several detection heads, each detection head is used to perform target detection on the fused features and output the detection results.

[0012] The result output module is connected to the model processing module and is used to remove redundant prediction boxes from the detection results using the non-maximum suppression algorithm to obtain the final small target detection results.

[0013] On the other hand, the present application also discloses a computer device, which includes a memory and a processor connected to the memory, wherein the memory stores a computer program running on the processor, and when the processor executes the computer program, the above-mentioned small target detection method based on CoMFE is implemented.

[0014] On the other hand, the present application further discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned small target detection method based on CoMFE is implemented.

[0015] Beneficial effects: Compared with the prior art, the present invention has the following beneficial effects:

[0016] (1) The coordinate multi-branch feature enhancement module proposed in the small target detection model effectively integrates the coordinate attention mechanism and the multi-branch feature fusion structure. While enhancing the spatial information expression capability of small targets, it introduces multi-scale semantic information, effectively improving the detection accuracy and recall rate of small targets.

[0017] (2) The MFE module proposed in this invention is reasonably designed. The three branches are responsible for lightweight feature extraction, multi-scale context modeling, and basic information retention respectively. The fusion of the three makes the enhanced features have stronger discrimination ability and performs better in complex backgrounds and target-dense scenes.

[0018] (3) The design of the present invention has strong compatibility and versatility, can be flexibly embedded in different versions of YOLO networks, and has good deployment flexibility and engineering application value.

[0019] (4) The method of the present invention has a simple structure and does not significantly increase the number of model parameters while ensuring the improvement of small target detection performance. It has good prospects for deployment and application in practical scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings used in the embodiments.

[0021] Figure 1 This is the overall network structure diagram of the original YOLOv5;

[0022] Figure 2 This is a diagram of the overall network structure of the present invention applied to YOLOv5;

[0023] Figure 3 This is a network structure diagram of the present invention applied to YOLOv8;

[0024] Figure 4 This is a network structure diagram of the present invention applied to YOLOv7;

[0025] Figure 5 This is a schematic diagram of the structure of the coordinate multi-branch feature enhancement module used in the present invention;

[0026] Figure 6 The structure diagram of the coordinate attention mechanism introduced in this invention;

[0027] Figure 7 Multi-branch feature enhancement structure diagram designed for the present invention;

[0028] Figure 8 This is a comparison chart of experimental indicators applied to YOLOv5 in the present invention;

[0029] Figure 9 This is a comparison chart of the test results of the present invention applied to YOLOv5;

[0030] Figure 10 is a flow chart of the method of the present invention;

[0031] Figure 11 Schematic diagram of the device structure of the present invention. DETAILED DESCRIPTION

[0032] To more clearly illustrate the technical solution of the present invention, the following, in conjunction with the accompanying drawings and examples, further describes a small target detection method, apparatus, device, and medium based on CoMFE. The following examples will help those skilled in the art further understand the present invention but do not limit the present invention in any way. It should be noted that those skilled in the art may make various variations and improvements without departing from the scope of the present invention. These variations and improvements are all within the scope of protection of the present invention.

[0033] As attached Figure 10 As shown, a small target detection method based on CoMFE of the present invention includes the following steps:

[0034] S1. Obtain an original image containing multiple small targets;

[0035] The original image also includes data augmentation operations on the data set before entering S2 processing. The core of data augmentation is to expand the training data set, thereby improving the robustness of the model to images from different environments. In this application, two types of data augmentation strategies, photometric and geometric transformation, are adopted at the same time. For photometric enhancement, the hue, saturation and brightness of the image are adjusted to simulate the image performance under various lighting conditions. For geometric enhancement, random scaling, center cropping, flipping, shearing and rotation are used to reproduce changes in different camera perspectives and object postures.

[0036] In addition to the above-mentioned global pixel-level enhancement methods, this application also introduces data enhancement technology for multi-image combination. For example, the MixUp method generates a mixed image by randomly selecting two training samples and performing weighted summation, while fusing the labels accordingly; the CutMix method uses an area of another image to replace part of the current image. This method avoids using a full-zero "black cloth" to block the image; the Mosaic method splices four images together, greatly enriching the background information of the detected object. In addition, batch normalization calculates activation statistics based on multiple images at each layer, which also helps to improve the generalization ability of the model;

[0037] S2. Input the original image into the pre-trained small target detection model and output the detection results, which include the category, confidence level and boundary information of the prediction box of each detected target;

[0038] The small target detection model includes a backbone network (backbone), a neck network (neck) and a detection head (prediction head) connected in sequence; wherein, the backbone network is a backbone network based on a coordinate multi-branch feature enhancement module, that is, a number of coordinate multi-branch feature enhancement modules (CoMFE modules) are introduced into the backbone network including multiple layers of convolutional layers (Conv), batch normalization layers (BN), activation function layers (SiLU), maximum pooling layers (Maxpool), and splicing layers (Concat), that is, a coordinate attention module and a multi-branch feature enhancement module are introduced into the backbone network for performing multi-feature extraction and output on the original image, the output of each coordinate multi-branch feature enhancement module and the final output of the backbone network are connected to the neck network, the neck network is used to fuse the different scale features output by the backbone network and output the fused features, the neck network is connected to a number of detection heads, each detection head is used to perform target detection on the fused features and output the detection results;

[0039] The number of coordinate multi-branch feature enhancement modules is the number of detection heads minus 1. In other embodiments of the present application, the number of coordinate multi-branch feature enhancement modules can be arbitrarily set according to actual conditions. Figure 5 As shown in the figure, the coordinate multi-branch feature enhancement module includes a coordinate attention module (CA) and a multi-branch feature enhancement module (MFE) connected in sequence, which is used to retain and enhance spatial details and context information in small target feature extraction and output an enhanced feature map;

[0040] The coordinate attention module (CA) is used to generate direction-sensitive spatial attention features by performing global pooling along the vertical and horizontal directions respectively, so as to further enhance the feature map’s ability to perceive spatial information; the coordinate attention module outputs the enhanced feature map Y.

[0041] As attached Figure 6As shown in the figure, in the coordinate attention module (CA), the input feature map (input) is processed by the residual layer (Residual), and then global average pooling operations are performed along the horizontal and vertical directions of the input feature map, that is, after being processed by the horizontal average pooling layer (Vertical Avg Pool) and the vertical average pooling layer (Horizontal Avg Pool), one-dimensional feature vectors of the two directions are obtained respectively; then the one-dimensional feature vectors of the two directions are spliced by the splicing layer (Concat) according to the number of channels, and processed by the 1×1 two-dimensional image convolution layer (Conv2d), the batch normalization layer (BatchNorm) and the nonlinear activation function layer (Non-linear) to obtain the fused direction-sensitive features; finally, the fused direction-sensitive features are respectively passed through two independent 1×1 two-dimensional image convolution layers and the activation function layer (Sigmoid) to generate attention weights in the horizontal and vertical directions, and the attention weights are re-weighted to the original input feature map along the corresponding directions through the reweighting layer (Re-weight) to obtain the enhanced feature map output.

[0042] Specifically, given an input X of size (C, W, H), where C, W, and H are the number of channels, width, and height, respectively, two pooling kernels with different spatial ranges (H, 1) and (1, W) are used in the horizontal average pooling layer and the vertical average pooling layer to encode each channel along the horizontal and vertical coordinate directions, respectively. In this way,

[0043] The output of the cth channel at height h can be expressed as:

[0044]

[0045] Similarly, the output of the cth channel at width w can be expressed as:

[0046]

[0047] in, and Represents the feature map x of the c-th channel height h respectively c The average output feature map in the width i direction and the feature map x of the c-th channel width w c The average output feature map in the height j direction; then, the feature maps in the two directions are spliced and sent to the convolution function F to obtain f in different directions h and f w After normalization and activation, they are fed into the convolution function F h and F wFinally, the attention weights in two directions are obtained through the sigmoid activation function, and the attention weights are re-weighted to the original input features along the corresponding directions. Figure X In , the enhanced feature map Y is obtained.

[0048] In addition, the multi-branch feature enhancement module includes GB branch (GhostBoost branch), MD branch (Multi-Dilation branch) and BSC branch (BaseConv branch). The enhanced feature map Y output by the coordinate attention module is processed by GB branch, MD branch and BSC branch respectively, and then information is fused through the splicing layer (Cat) and convolution layer (Conv) to form a final enhanced feature map with richer semantics and stronger information expression ability.

[0049] Specifically, the GhostBoost branch uses lightweight Ghost convolutional layers to quickly extract rich small object features;

[0050] As attached Figure 7 As shown, the feature map Y enhanced by the coordinate attention module is input into the GhostBoost branch. Assuming that the size of the feature map Y is (C Y , H Y , W Y ), since the stride of this method is 1, depth-wise separable convolution is not used, and two Ghost convolutions (GhostConv) are used to obtain the Ghost convolution output y g , the size is (C Y , H Y , W Y ), and finally adopt the residual connection method to connect the input Y and y g After adding (Add) and processing through the 1*1 convolution layer (Conv), the enhanced feature map y of the GhostBoost branch is obtained GB , the size is (C Y / 2,H Y , W Y ).

[0051] The Multi-Dilation branch uses dilated convolution to construct a multi-scale receptive field and extract multi-scale context information;

[0052] The feature map Y enhanced by the coordinate attention module is input into the Multi-Dilation branch. Assuming that the feature map size is (C Y , H Y , W Y ), firstly, the first output y1 is obtained through the dilation rate 1 void convolution (Conv) and point-by-point convolution (CBR), the size is (C Y / 2,HY , W Y ), then the second output y2 is obtained through the dilation rate 1 and point-by-point convolution, the size remains unchanged, and finally the final enhanced feature map y enhanced by the Multi-Dilation branch is obtained through the dilation rate 1 and point-by-point convolution. MD , the size is (C Y / 2,H Y , W Y ); Point-by-point convolution (CBR) includes sequentially connected convolution layers, batch normalization layers, and activation function layers (ReLU);

[0053] The BaseConv branch uses standard convolution to reduce the number of feature map channels and retain basic small object information;

[0054] The feature map Y enhanced by the coordinate attention module is input into the BaseConv branch, assuming that the feature map size is (C Y , H Y , W Y ), reduce the channel dimension of the feature map through the standard 3×3 convolution, highlight the basic feature information, and obtain the enhanced feature map y of the BaseConv branch BSC , the size is (C Y / 2,H Y , W Y );

[0055] Finally, the feature maps obtained by the above three branches are spliced along the channel dimension, and 1×1 convolution is used for information fusion to form the final enhanced feature map y with richer semantics and stronger information expression ability. CoMFE .

[0056] The pre-training process of the small object detection model includes: using Visdrone-2019 as the training dataset, preprocessing the dataset during training, namely the aforementioned data augmentation, then inputting the improved model for training, and finally using the Ciou loss function to measure the difference between the predicted value and the true value.

[0057] S3. Use the non-maximum suppression algorithm to remove redundant prediction boxes from the detection results to obtain the final small target detection results. In this application, the small target magnitude is defined by the absolute scale, that is, the target pixel size is less than 32*32 is considered a small target;

[0058] Example 1

[0059] In this embodiment, the method described in this application is applied to YOLOv5. Figure 1 The overall network structure diagram of the original YOLOv5 is given in Figure 2This is a diagram of the overall network structure of the present invention applied to YOLOv5, that is, the overall architecture of the small target detection model is the YOLOv5 network structure, the backbone network is CSPDarknet53, and the coordinate attention module and the multi-branch feature enhancement module are introduced into the backbone network to perform multi-feature extraction and output on the original image; in addition, in order to avoid the loss of small target feature information, the subsequent structure of the backbone network is improved in this embodiment, and the two convolution modules at the end of the original backbone network are removed to better retain the detailed feature information of the small target; the input image is sent to the backbone network for layer-by-layer feature extraction to obtain initial feature maps of different scales;

[0060] In this example, we removed the two layers before the SPPF in the YOLOv5 backbone network. By optimizing the YOLO backbone network structure and removing the terminal convolutional and residual modules, this method further preserves shallow detail feature information. While maintaining moderate computational complexity, it also improves the model's ability to perceive small targets, making it particularly suitable for small target detection tasks from the perspective of drones.

[0061] In particular, the method of this embodiment was trained on the VisDrone-2019 dataset and tested and verified. Figure 8 As shown in the figure, compared with the original model (YOLOv5s (baseline)), the average precision (Map@.5) of this method (CoMFE-YOLOv5) increased by 7.6%, and the indicators of Map@[.5:.95], precision P, and recall rate R also increased simultaneously, and the model parameter Params decreased by 39.35%.

[0062] Therefore, after integrating the advantages of each module, the final improved version of this application demonstrated excellent performance, successfully retained more small target information, and performed targeted information fusion and enhancement at key spatial positions, thereby significantly improving the detection accuracy of the model.

[0063] like Figure 9 As shown in the figure, a comparison chart of the small target detection performance of this method and the original model is given. The left is the small target detection output result of the original model, and the right is the small target detection output result of this method. The comparison shows that this method can have a better detection effect on dense small target crowds.

[0064] Example 2

[0065] In this embodiment, the method described in this application is applied to YOLOv8, as shown in the attached Figure 3 As shown in the figure, this method is applied to the overall network structure of YOLOv8. The last two layers of convolution modules can also be deleted. In terms of performance, this method is also better than YOLOv8.

[0066] Example 3

[0067] In this embodiment, the method described in this application is applied to YOLOv7, as shown in the attached Figure 4 Figure 2 shows the overall network structure of this method applied to YOLOv7. CoMFE can be used directly as a plug-and-play module. This method also achieves better performance than YOLOv7.

[0068] Attachment Figure 11 This is a structural diagram of a small target detection device based on CoMFE of the present invention. Figure 11 As shown, the device includes an image acquisition module, a model processing module, and a result output module.

[0069] An image acquisition module is used to acquire an original image containing multiple small targets;

[0070] The model processing module is connected to the image acquisition module and is used to input the original image into the pre-trained small target detection model and output the detection results, which include the category, confidence and boundary information of the prediction box of each detected target; the small target detection model includes a backbone network, a neck network and a detection head connected in sequence, wherein the backbone network is a backbone network based on a coordinate multi-branch feature enhancement module, which is used to extract and output multiple features of the original image, and the output of each coordinate multi-branch feature enhancement module and the final output of the backbone network are connected to the neck network. The neck network is used to fuse the different scale features output by the backbone network and output the fused features. The neck network is connected to several detection heads, each of which is used to perform target detection on the fused features and output the detection results;

[0071] The result output module is connected to the model processing module and is used to remove redundant prediction boxes from the detection results using the non-maximum suppression algorithm to obtain the final small target detection results.

[0072] The present invention also discloses a computer device, which includes a processor and a memory coupled to the processor.

[0073] The memory stores program instructions for implementing the above-mentioned small target detection method based on CoMFE. The processor is used to execute the program instructions stored in the memory to detect small targets.

[0074] The processor may also be referred to as a CPU (Central Processing Unit). A processor may be an integrated circuit chip with signal processing capabilities. The processor may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. A general-purpose processor may be a microprocessor or any conventional processor.

[0075] The present invention also discloses a computer-readable storage medium storing a program file capable of implementing all of the above methods, wherein the program file can be stored in the above-mentioned computer-readable storage medium in the form of a software product, including a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned computer-readable storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program code, or a terminal device such as a computer, server, mobile phone, or tablet.

[0076] In the description of the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0077] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0078] The above are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A small target detection method based on CoMFE, characterized in that: Methods include: Get the original image containing multiple small targets; The original image is input into the pre-trained small target detection model, and the detection result is output, which includes the category, confidence and boundary information of the prediction box of each detected target; the small target detection model includes a backbone network, a neck network and a detection head connected in sequence, wherein the backbone network is a backbone network based on a coordinate multi-branch feature enhancement module, i.e., a CoMFE module. The coordinate attention module and the multi-branch feature enhancement module are introduced into the backbone network to perform multi-feature extraction and output on the original image. The output of each coordinate multi-branch feature enhancement module and the final output of the backbone network are connected to the neck network. The neck network is used to fuse the different scale features output by the backbone network and output the fused features. The neck network is connected to several detection heads, each detection head is used to perform target detection on the fused features and output the detection results. The non-maximum suppression algorithm is used to remove redundant prediction boxes from the detection results to obtain the final small target detection results.

2. The small target detection method based on CoMFE according to claim 1, characterized in that: The number of coordinate multi-branch feature enhancement modules is the number of detection heads minus 1.

3. The small target detection method based on CoMFE according to claim 1, characterized in that: The coordinate multi-branch feature enhancement module includes a coordinate attention module and a multi-branch feature enhancement module connected in sequence, which is used to retain and enhance spatial details and contextual information in small object feature extraction and output an enhanced feature map; The coordinate attention module is used to generate direction-sensitive spatial attention features by performing global pooling in the vertical and horizontal directions respectively, so as to further enhance the feature map's ability to perceive spatial information; the coordinate attention module outputs the enhanced feature map; The multi-branch feature enhancement module includes GB branch, MD branch and BSC branch. The enhanced feature Y output by the coordinate attention module is processed by GB branch, MD branch and BSC branch respectively, and the final enhanced feature map is output.

4. The small target detection method based on CoMFE according to claim 3, characterized in that: In the coordinate attention module, after the input feature map is processed by the residual layer, global average pooling operations are performed along the horizontal and vertical directions of the input feature map to obtain one-dimensional feature vectors in each direction. Then, the one-dimensional feature vectors in the two directions are spliced through the splicing layer according to the number of channels, and processed using the two-dimensional image convolution layer, batch normalization layer and nonlinear activation function layer to obtain the fused direction-sensitive features. Finally, the fused direction-sensitive features are passed through two independent two-dimensional image convolution layers and activation function layers to generate attention weights in the horizontal and vertical directions. The attention weights are re-weighted to the original input feature map along the corresponding directions through Re-weight to obtain the enhanced feature map output.

5. The small target detection method based on CoMFE according to claim 3, characterized in that: The GhostBoost branch uses a lightweight Ghost convolution layer to quickly extract rich small target features; the feature map Y enhanced by the coordinate attention module is input into the GhostBoost branch, and two Ghost convolutions are performed to obtain the Ghost convolution output y g Finally, the residual connection method is adopted to connect the input Y and y g After adding and processing through the convolution layer, the enhanced feature map of the GhostBoost branch is obtained.

6. The small target detection method based on CoMFE according to claim 3, characterized in that: The Multi-Dilation branch uses dilated convolution to construct a multi-scale receptive field and extract multi-scale context information; The feature map Y enhanced by the coordinate attention module is input into the Multi-Dilation branch. First, the first output y1 is obtained by the dilation convolution and point-by-point convolution with a dilation rate of 1. Then, the second output y2 is obtained by the dilation convolution and point-by-point convolution with a dilation rate of 1. The size remains unchanged. Finally, the final enhanced feature map enhanced by the Multi-Dilation branch is obtained by the dilation convolution and point-by-point convolution with a dilation rate of 1.

7. The small target detection method based on CoMFE according to claim 3, characterized in that: The BaseConv branch uses standard convolution to reduce the number of feature map channels and retain basic small target information; The feature map Y enhanced by the coordinate attention module is input into the BaseConv branch, and the channel dimension of the feature map is reduced by standard convolution to highlight the basic feature information, thereby obtaining the enhanced feature map of the BaseConv branch.

8. A small target detection device based on CoMFE, characterized in that: include: An image acquisition module is used to acquire an original image containing multiple small targets; The model processing module is connected to the image acquisition module and is used to input the original image into the pre-trained small target detection model and output the detection results, which include the category, confidence and boundary information of the prediction box of each detected target; the small target detection model includes a backbone network, a neck network and a detection head connected in sequence, wherein the backbone network is a backbone network based on the coordinate multi-branch feature enhancement module, i.e., the CoMFE module. The coordinate attention module and the multi-branch feature enhancement module are introduced into the backbone network to perform multi-feature extraction and output on the original image. The output of each coordinate multi-branch feature enhancement module and the final output of the backbone network are connected to the neck network. The neck network is used to fuse the different scale features output by the backbone network and output the fused features. The neck network is connected to several detection heads, each detection head is used to perform target detection on the fused features and output the detection results. The result output module is connected to the model processing module and is used to remove redundant prediction boxes from the detection results using the non-maximum suppression algorithm to obtain the final small target detection results.

9. A computer device comprising a memory and a processor connected to the memory, wherein the memory stores a computer program running on the processor, wherein: When the processor executes the computer program, the small target detection method based on CoMFE according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the small target detection method based on CoMFE is implemented as claimed in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Small target detection system and method based on improved YOLOv5

    CN117523267A

  • Unmanned aerial vehicle image detection method based on multi-scale feature fusion and context enhancement

    CN117037004A

  • Ultra-wide metal surface flaw detection and identification method

    CN117392116A

  • Traffic sign detection method based on lightweight convolutional neural network

    CN118587676A

  • Efficient refinement neural network for real-time generic object-detection systems and methods

    US20220019843A1