Small target detection method based on CAFM-ADown-EMA fusion
By integrating the small object detection method of CAFM, ADown and EMA modules in the YOLOv8 framework, the problems of large amount of model parameters and high computational complexity in the prior art are solved, and the small object detection effect with high precision and low complexity is achieved.
Patent Information
- Application Number
- CN202510059184.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-30
AI Technical Summary
The existing small-objective detection technology based on deep learning has problems such as large number of model parameters, high computational complexity and strict hardware configuration requirements in remote sensing images, and it is difficult to reduce the computational complexity and parameter quantity while ensuring detection accuracy.
Using a small object detection method based on CAFM-ADown-EMA fusion, the CAFM module, ADown module and EMA module are introduced into the backbone and neck network of the YOLOv8 framework, the C2f module and the Conv module are optimized, and the detection head is improved in combination with the self-attention mechanism to reduce the parameter amount of the model and the calculation complexity.
The detection accuracy of small object detection is improved, the calculation complexity and parameter quantity of the model are reduced, the feature extraction capability in complex environments is improved, and the small object detection with high accuracy and low complexity is realized.
Smart Images

Figure CN120070848A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to object detection, and specifically to a small object detection method based on the fusion of CAFM-ADown-EMA. Background Art
[0002] Remote sensing image detection plays an important role in many fields, such as environmental monitoring, urban planning, disaster assessment, etc. Especially in the military and security fields, the ability to autonomously and accurately detect and identify key targets in remote sensing images can provide key information for intelligence collection, strategic planning, disaster response, etc. Due to its shooting characteristics, remote sensing images often face challenges such as small target sizes, unclear features, and low signal-to-noise ratios, which all have a negative impact on the performance of detection algorithms.
[0003] Small object detection technology plays a key role in remote sensing image detection. It is mainly used to detect targets with fewer pixels in remote sensing images. Such targets have less feature information, low signal-to-noise ratio, strong background interference, and certain occlusion of the targets. Currently, the commonly used object detection technology based on deep learning can greatly improve the detection accuracy. However, on the premise of ensuring the accuracy, it will expose the problems of large model parameter quantity, high computational complexity, and strict hardware configuration requirements. Summary of the Invention
[0004] Object of the Invention: Aiming at the problems in the background art, the present invention provides a small object detection method based on the fusion of CAFM-ADown-EMA, which has high detection accuracy, low computational complexity, and small parameter quantity.
[0005] Technical Solution: The present invention discloses a small object detection method based on the fusion of CAFM-ADOWN-EMA, including the following steps:
[0006] Step (1): Obtain relevant small object class images, and preprocess the images to form a data set;
[0007] Step (2): Build a small object detection model based on the YOLOv8 framework. In the backbone network and neck network of the small object detection model, the C2f module is improved using the CAFM module, and some Conv modules are replaced with the ADown module; the EMA module is additionally added to the neck network; in the detection head part, the Conv module is modified using the self-attention mechanism to form the Detect-SAT detection head;
[0008] Step (3): Use the data set to train the small object detection model, and use the trained model to perform small object detection.
[0009] Further, in the step (1), the sizes of all the collected images are adjusted to 640×640 pixels and then annotated, and the ratio of the training set to the test set is 8:2.
[0010] Furthermore, the improved part of the backbone network includes three C2f-CAFM modules and one ADown module. Among them, the three C2f-CAFM modules are respectively placed after P3, P4, and P5, and one ADown module replaces the Conv module of P5;
[0011] The improved part of the neck network includes two C2f-CAFM modules, two ADown modules and three EMA modules. Among them, the two C2f-CAFM modules are respectively placed after the third and fourth Concat operations, the two ADown modules are respectively placed before the third and fourth Concat operations, and the three EMA modules are respectively placed in the part connected to the detection head.
[0012] Furthermore, when using the CAFM module to improve the C2f module in step (2), the Bottleneck in the C2f module is replaced by the CAFM module to form C2f-CAFM;
[0013] The CAFM module includes a local branch and a global branch. The local branch includes Channel shuffle and Conv modules, and the global branch includes an attention map, Dconv module and Conv module;
[0014] When the input feature map enters the local branch, the number of channels is adjusted by a 1*1 Conv and local features are captured, and then the feature diversity is increased by Channel shuffle; when the input feature map enters the global branch, the number of channels is adjusted by a 1*1 Conv and the attention map is formed by Dconv to extract global features, and then the features are output through residual connection.
[0015] Furthermore, the processing process of the CAFM module is as follows:
[0016] F c = W 3×3×3 (CS(W 1×1 (Y))) (1)
[0017]
[0018] F out-CAFM = F c + F a (3)
[0019] Among them, F c represents the output of the local branch, W 1×1 represents the 1×1 Conv, W 3×3×3 represents the 3×3×3 Conv, CS represents channel shuffle, Y represents the input feature, F aRepresents the output of the global branch, Calculate the required tensors respectively. α represents the control parameter of matrix multiplication, F out-CAFM Represents the output features after residual connection.
[0020] Furthermore, the input feature map passes through the C2f-CAFM module and then enters the ADown module for downsampling to optimize the spatial dimension.
[0021] Furthermore, the ADown module performs average pooling, splitting, max pooling, parallel convolution, and concatenation operations on the input feature map. The formula is as follows:
[0022] F in = Chunk(AvgPool(X)) (4)
[0023] F out-ADown = ηW(F in1 )+(1 - η)W(MaxPool(F in2 )) (5)
[0024] Among them, X represents the input feature, AvgPool represents average pooling, Chunk represents the splitting operation, F in represents the preprocessing of the input feature, W represents the Conv operation, MaxPool represents max pooling, η represents the weight value between 0 and 1, F out-ADown represents the output of the ADown module.
[0025] Furthermore, the three output branches of the neck network are:
[0026] The first output branch upsamples the feature map P5 and then uses the Concat module to splice it with the feature map P4. After that, it enters the C2f module for feature fusion, then upsamples again and uses the Concat module to splice it with the feature map P3. After feature enhancement through the C2f module and the EMA module, the first output branch is input into the detection head;
[0027] The second output branch is downsampled through the ADown module and then spliced using the Concat module. After that, it enters the C2f-CAFM module for feature fusion and then is input into the detection head after feature enhancement through the EMA module;
[0028] The third output branch is downsampled through the ADown module and then spliced with the feature map P5 using the Concat module. After that, it enters the C2f-CAFM module for feature fusion and then is input into the detection head after feature enhancement through the EMA module.
[0029] Furthermore, after the EMA module groups the channels of the input feature map, it enhances feature extraction through global pooling and a 3×3 Conv, then performs feature fusion through the Sigmoid activation function, and finally calculates the adaptive weights and concatenates the outputs.
[0030] Furthermore, the Detect-SAT detection head first adjusts the number of channels of the input feature map through a Conv module, then captures context information through the SAT module for feature enhancement, then calculates the bounding box regression loss and classification loss through the Conv2d layer to optimize the network parameters, and finally outputs the detection results; the SAT module uses MHSA to enhance global feature extraction and adaptively changes the feature dimension through two groups of 1×1 Convs.
[0031] Beneficial effects:
[0032] (1) The present invention proposes a small target detection model integrating CAFM-ADown-EMA based on the YOLO framework. By introducing the CAFM module into the backbone network and the neck network to optimize the C2f module, the CAFM module can effectively model the global and local features of the image, extract deep features from multiple perspectives, thereby enhancing the model's recognition ability for small targets; the ADown module replaces part of the Conv module, and through advanced downsampling technology, the ADown module effectively retains key information, thereby improving the accuracy of small target detection; a new EMA module is added to the neck network, which can strengthen feature extraction, reduce the interference of complex backgrounds on small target detection, and thus improve the precision of small target detection. Integrating the CAFM module, ADown module, and EMA module into the YOLO framework can improve the downsampling of the original model and feature extraction in complex environments, thereby improving the detection accuracy of the model for small targets.
[0033] (2) The present invention uses the self-attention mechanism to improve the Conv module in the detection head part. Compared with the Conv module in the original detection head that focuses on retaining local information, this self-attention mechanism module can efficiently integrate feature information in different regions, reduce redundant calculations, thereby reducing the number of model parameters and computational complexity; by applying the CAFM and ADown modules in the backbone network and the neck network, the present invention can, without increasing the number of parameters, refine the complementary information between different features through iterative optimization of feature fusion, effectively controlling the growth of the number of parameters. Description of the drawings
[0034] Figure 1 is a flowchart of the small target detection method based on the CAFM-ADown-EMA fusion;
[0035] Figure 2It is the structural diagram of a small target detection network based on the CAFM-ADown-EMA fusion;
[0036] Figure 3 It is the structural diagram of the CAFM module;
[0037] Figure 4 It is the structural diagram of the C2f-CAFM module;
[0038] Figure 5 It is the structural diagram of the ADown module;
[0039] Figure 6 It is the structural diagram of the Detect-SAT module;
[0040] Figure 7 It is the structural diagram of the SAT module;
[0041] Figure 8 It is the training process curve diagram of the small target detection model in the present invention;
[0042] Figure 9 It is the effect schematic diagram of the small target detection model in the present invention. Detailed implementation manners
[0043] Next, the present invention will be elaborated in detail in conjunction with the accompanying drawings.
[0044] As Figure 1 shown, the present invention provides a small target detection method based on the CAFM-ADown-EMA fusion, which specifically includes the following steps:
[0045] Step (1) Collect relevant small target class pictures and preprocess the images;
[0046] Adjust the sizes of all the collected images to 640×640 pixels and then label them, where the ratio of the training set to the test set is 8:2. The folder for storing data is sods under the datasets directory, which includes sub-folders images and labels. There are two sub-folders train and val in images and labels respectively. The training pictures and related labels are put into train, and the test pictures and related labels are put into val.
[0047] Step (2) Build a small target detection model based on the YOLO framework. In its backbone network and neck network, use the CAFM module to improve the C2f module, and use the ADown module to replace some Conv modules; additionally add the EMA module in the neck network; use the self-attention mechanism to modify the Conv module in its detection head part.
[0048] As Figure 2As shown in the figure, the improved part of the backbone network includes three C2f-CAFM modules and one ADown module. Among them, the three C2f-CAFM modules are respectively placed after P3, P4, and P5, and one ADown module replaces the Conv module of P5. The improved part of the neck network includes two C2f-CAFM modules, two ADown modules and three EMA modules. Among them, the two C2f-CAFM modules are respectively placed after the third and fourth Concat operations, the two ADown modules are respectively placed before the third and fourth Concat operations, and the three EMA modules are respectively placed in the part connected to the detection head. The improved part of the detection head is to use three Detect-SAT to replace the original detection head.
[0049] The input feature map first undergoes channel expansion, feature map size adjustment, and feature aggregation through the Conv module and C2f module of the original network, and then the output feature is generated. Then the output feature enters the C2f-CAFM and ADown modules for feature enhancement to optimize the spatial dimension and parameters of the feature map.
[0050] As Figure 3 shown in the figure, the CAFM module contains a local branch and a global branch. The local branch contains a Channelshuffle and a Conv module, and the global branch contains an attention map, a Dconv module, and a Conv module. When the input feature map enters the local branch, it first adjusts the number of channels through a 1×1 Conv, then further mixes the channel information through Channel shuffle to increase feature diversity, and finally captures local features through a 3×3×3 Conv; when the input feature map enters the global branch, it first adjusts the number of channels through a 1×1 Conv, then generates Q, K, and V through a 3×3 Dconv to produce three tensors of shape H×W×C, and then extracts global features through reshaping and the attention map, and finally outputs the feature through a residual connection. The formula is as follows:
[0051] F c = W 3×3×3 (CS(W 1×1 (Y))) (1)
[0052]
[0053] F out-CAFM = F c + F a (3)
[0054] Among them, F c represents the output of the local branch, W 1×1 represents the 1×1 Conv, W 3×3×3 represents the 3×3×3 Conv, CS represents channel shuffle, Y represents the input feature, Fa Represents the output of the global branch, Calculate the required tensors separately. α represents the control parameter for matrix multiplication, F out-CAFM Represents the output features after residual connection.
[0055] As Figure 4 shown, use the CAFM module to improve the Bottleneck in C2f to form C2f-CAFM. After the input feature map enters the C2f-CAFM module, first use a Conv with k = 1, stride s = 1, and parameter p = 0 in the convolutional layer to adjust the parameters. Then, the feature map is split into two parts through the Split module, with the number of channels in each part halved. Then, use the n-layer CAFM module to extract features from the feature map. The feature map output from the Split module is copied multiple times to match the output quantity of the CAFM module. These copied feature maps will be concatenated with the output of the CAFM module in the Concat module. Finally, parameter adjustment is performed through a Conv with k = 1, stride s = 1, and parameter p = 0 to keep the spatial dimension of the feature map unchanged.
[0056] The input feature map is input to the ADown module for downsampling after passing through C2f-CAFM to optimize the spatial dimension. As Figure 5 shown, the ADown module first performs average pooling on the input feature map to compress its spatial dimension while retaining the number of channels. Then, the feature map is split into two parts in a certain ratio through the Chunk layer. Then, one part is used for feature extraction through the Conv module, and the other part is compressed in spatial dimension through the max pooling layer and then used for feature extraction through the Conv module. Finally, it is output through concatenation in the Concat module. The formula is as follows:
[0057] F in = Chunk(AvgPool(X)) (4)
[0058] F out-ADown = ηW(F in1 )+(1 - η)W(MaxPool(F in2 )) (5)
[0059] Among them, X represents the input feature, AvgPool represents average pooling, Chunk represents the splitting operation, F in represents the preprocessing of the input feature, W represents the Conv operation, MaxPool represents max pooling, η represents the weight value between 0 and 1, and F out-ADown represents the output of the ADown module.
[0060] The first output branch of the neck network first upsamples the feature map P5 to reduce its spatial dimension to 0.5 times the original, then uses the Concat module to concatenate it with the feature map P4 with the same spatial dimension and number of channels, then enters the C2f module for feature fusion, and then upsamples to reduce both the spatial dimension and the number of channels of the output features and uses the Concat module to concatenate them with the feature map P3, and then passes through the C2f module and the EMA module for feature enhancement and inputs the first output branch into the detection head. The second output branch downsamples the output of the first output branch through the ADown module, then uses the Concat module for concatenation, then enters the C2f-CAFM module for feature fusion, and then passes through the EMA module for feature enhancement and inputs it into the detection head. The third output branch downsamples the output of the second output branch through the ADown module and concatenates it with the feature map P5 using the Concat module, then enters the C2f-CAFM module for feature fusion, and then passes through the EMA module for feature enhancement and inputs it into the detection head.
[0061] The EMA module adopts a parallel sub-network structure, which includes a parallel sub-network that processes a 1×1 Conv and a 3×3 Conv. This structure helps to effectively capture cross-dimensional interactions and establish dependencies between different dimensions, thereby improving the ability of feature representation. EMA divides the input feature map X into G sub-feature groups, and each group learns different semantics, where G is much smaller than the number of channels C. Such a grouping method can strengthen the feature learning of semantic regions and compress noise. Adding the EMA module after the outputs of C2f and C2f-CAFM in the neck network can learn effective channel descriptions without reducing the channel dimension and generate better pixel-level attention for the high-level feature map, thereby reducing the model computational complexity and improving the detection accuracy.
[0062] As Figure 6 shown, the Conv module in the detection head part is improved using the self-attention mechanism to form Detect-SAT. The input feature map first adjusts the number of channels through an original Conv module, then captures more extensive context information through the SAT module for feature enhancement, then calculates the bounding box regression loss and classification loss through the Conv2d layer to optimize the network parameters, and finally outputs the detection results.
[0063] As Figure 7 shown, after the input feature map enters the SAT module, it first reduces the feature dimension through a 1×1 Conv, then captures the long-range dependencies between features through MHSA, then restores the feature dimension through a 1×1 Conv, and at the same time retains the uncompressed residual information through an additional bypass 1×1 Conv, and finally outputs the enhanced features through the residual connection.
[0064] In step (3), the small target detection model is trained using the dataset, and the trained model is used for small target detection.
[0065] The relevant training set images and label data in the dataset are input into the small target detection model, and training is performed using the GPU. The small target dataset includes airplanes and oil tanks in a remote sensing background. When the model training is completed, the optimal model parameters during the training process are automatically saved and named best.pt.
[0066] As Figure 8 shown, it shows the evolution curves of various key performance indicators of the small target detection model within 300 training epochs. After verification after the model training is completed, the mAP50 reaches 98.4%. Among them, mAP is used to evaluate the overall performance of the multi-class target detection model, and the larger this indicator, the better the performance of the model network.
[0067] Compared with the original YOLOv8n, the mAP50 of the present invention has increased by 1.4%, the number of parameters has decreased by 18.3%, and the computational complexity has decreased by 16%. The comparison of model indicators is shown in Table 1, which proves the feasibility of the small target detection model.
[0068] Table 1 Comparison of experimental results of different models
[0069] Model Class Map50(%) Params(M) Glops(M) YOLOv8n PlaneandOiltank 0.970 3.00 8.2 Ours PlaneandOiltank 0.984 2.45 6.8
[0070] As Figure 9 shown, it shows the detection effect of the present invention. As can be seen from Table 1, the present invention improves the accuracy of small target detection, reduces the computational complexity, and reduces the amount of calculation.
[0071] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it cannot be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A small target detection method based on CAFM-ADOWN-EMA fusion, characterized in that: The following steps are involved: Step (1) obtaining relevant small target class images and preprocessing the images to form a data set; Step (2) building a small target detection model based on the YOLOv8 framework, using the CAFM module to improve the C2f module in the backbone network and the neck network of the small target detection model, and using the ADown module to replace part of the Conv module; An additional EMA module is added to the neck network; in the detection head part, the Conv module is modified using the self-attention mechanism to form the Detect-SAT detection head; Step (3) uses the data set to train the small target detection model, and uses the trained model to perform small target detection.
2. A small target detection method based on CAFM-ADown-EMA fusion according to claim 1, characterized in that: In the step (1), all collected images are resized to 640×640 pixels and then annotated, wherein the ratio of the training set to the test set is 8:
2.
3. A small target detection method based on CAFM-ADown-EMA fusion according to claim 1, characterized in that: The improved part of the backbone network includes three C2f-CAFM modules and one ADown module, wherein the three C2f-CAFM modules are placed after P3, P4, and P5 respectively, and one ADown module replaces the Conv module of P5; The improved parts of the neck network include two C2f-CAFM modules, two ADown modules and three EMA modules, among which the two C2f-CAFM modules are placed after the third and fourth Concat operations respectively, the two ADown modules are placed before the third and fourth Concat operations respectively, and the three EMA modules are placed in the parts connected to the detection head.
4. A small target detection method based on CAFM-ADown-EMA fusion according to claim 3, characterized in that: When the C2f module is improved by using the CAFM module in step (2), the CAFM module is used to replace the Bottleneck in the C2f module to form a C2f-CAFM; The CAFM module contains local branches and global branches. The local branch contains Channel shuffle and Conv modules, and the global branch contains attention map, Dconv module and Conv module. When the input feature map enters the local branch, a 1*1 Conv is used to adjust the number of channels and capture local features, and then Channel shuffle is used to increase feature diversity; when the input feature map enters the global branch, a 1*1 Conv is used to adjust the number of channels and Dconv is used to form an attention map to extract global features, and then the features are output through residual connections.
5. A small target detection method based on CAFM-ADown-EMA fusion according to claim 4, characterized in that: The processing process of the CAFM module is: F c =W 3×3×3 (CS(W 1×1 (Y))) (1) F out-CAFM =F c +F a (3) Among them, F c represents the output of the local branch, W 1×1 Represents 1×1 Conv, W 3×3×3 represents 3×3×3 Conv, CS represents channel shuffle, Y represents input features, F a represents the output of the global branch, Calculate the required tensors respectively, α represents the control parameter of matrix multiplication, F out-CAFM Represents the output features after residual connection.
6. A small target detection method based on CAFM-ADown-EMA fusion according to claim 3, characterized in that: The input feature map passes through the C2f-CAFM module and then enters the ADown module for downsampling to optimize the spatial dimension.
7. A small target detection method based on CAFM-ADown-EMA fusion according to claim 6, characterized in that: The ADown module performs average pooling, segmentation, maximum pooling, parallel convolution, and concatenation operations on the input feature map. The formula is as follows: F in =Chunk(AvgPool(X)) (4) F out-ADown =ηW(F in1 )+(1-η)W(MaxPool(F in2 )) (5) Among them, X represents the input feature, AvgPool represents the average pooling, Chunk represents the segmentation operation, and F in represents the preprocessing of input features, W represents the Conv operation, MaxPool represents the maximum pooling, η represents the weight between 0 and 1, and F out-ADown Represents the output of the ADown module.
8. The small target detection method based on CAFM-ADown-EMA fusion according to claim 1 is characterized in that: The three output branches of the neck network are: The first output branch upsamples the feature map P5 and then concatenates it with the feature map P4 using the Concat module. Then, it enters the C2f module for feature fusion. After upsampling again, it concatenates it with the feature map P3 using the Concat module. After feature enhancement through the C2f module and the EMA module, the first output branch is input into the detection head. The second output branch is downsampled by the ADown module and then concatenated by the Concat module. It then enters the C2f-CAFM module for feature fusion and is then enhanced by the EMA module before being input into the detection head. The third output branch is downsampled by the ADown module and concatenated with the feature map P5 using the Concat module. It then enters the C2f-CAFM module for feature fusion and is then enhanced by the EMA module before being input into the detection head.
9. The small target detection method based on CAFM-ADown-EMA fusion according to claim 7, characterized in that: After the EMA module groups the input feature map channels, it uses global pooling and 3×3 Conv to enhance feature extraction, then performs feature fusion through the Sigmoid activation function, and finally calculates the adaptive weights and splices the output.
10. The small target detection method based on CAFM-ADown-EMA fusion according to claim 1, characterized in that: The Detect-SAT detection head first adjusts the number of channels of the input feature map through a Conv module, then captures context information through the SAT module for feature enhancement, and then calculates the bounding box regression loss and classification loss through the Conv2d layer to optimize the network parameters, and finally outputs the detection result; the SAT module uses MHSA to enhance global feature extraction and adaptively changes the feature dimension through two groups of 1×1 Conv.
Citation Information
Cited By
Improved YOLOv8-based low-altitude citrus pest and disease damage identification method, equipment and medium
CN121545151A
Low-altitude citrus pest and disease identification method and device based on improved yolov8 and medium
CN121545151B
Lightweight satellite cloud detection method, apparatus and device, and storage medium
CN122049708A