Image processing method based on small target detection
Multi-scale feature maps are generated through U-Netv2 network, and combined with MSCA and FFM modules, the problem of inaccurate small object detection in traditional methods is solved, and high-precision image analysis is achieved.
Patent Information
- Application Number
- CN202510448092.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional small object detection methods are difficult to accurately identify small objects in complex contexts, and information loss leads to inaccurate detection.
The U-Netv2 network is used to generate multi-scale feature maps, and feature extraction and fusion are performed through the multi-scale cross-axis attention module (MSCA) and feature fusion module (FFM), to enhance the model's perception of details and boundaries.
It significantly improves the accuracy and robustness of small object detection, effectively solves the problem of information loss, and improves the accuracy and reliability of detection.
Smart Images

Figure CN120298671A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to an image processing method based on small target detection. Background Art
[0002] In the field of computer vision, with the continuous progress of artificial intelligence technology, small target detection has gradually become a research hotspot. However, due to the small proportion of small targets in images, limited detail information, and frequent interference from complex backgrounds and noises, traditional detection methods often have difficulty in accurately identifying them. To break through this bottleneck, researchers have begun to explore new technical paths. The multi-scale feature extraction technology has emerged as the times require. By analyzing image features at different scales, it can simultaneously capture the local details and overall contours of small targets, effectively making up for the deficiencies of single-scale feature extraction. This technology not only enhances the model's perception ability of small targets but also provides a richer information basis for subsequent detection tasks.
[0003] At the same time, the attention mechanism in deep learning has brought new opportunities for small target detection. It can guide the model to focus on key information regions in the image, ignore background interference, so as to quickly locate small targets in complex scenes and improve the detection accuracy and efficiency. Combining multi-scale feature extraction with the attention mechanism effectively improves the accuracy and robustness of small target detection. In practical applications, such as in the fields of drone monitoring and satellite remote sensing image analysis, the application of this technology has greatly improved the reliability and practicality of small target detection.
[0004] The present invention proposes a new small target detection method. By integrating the advantages of multi-scale feature extraction and the attention mechanism, it can simultaneously capture the detailed features and global information of small targets, and dynamically focus on key regions, effectively solving the problem of inaccurate detection caused by information loss in traditional methods. Summary of the Invention
[0005] The purpose of the present invention is to provide an image processing method based on small target detection. Through a feature extraction network, hierarchical feature maps are generated and sequentially transmitted to a multi-scale cross-axis attention (MSCA) module and a feature fusion module (FFM) for refinement, better capturing and integrating global and multi-scale spatial information, and enhancing the model's perception ability for details and boundaries, thereby extracting new feature map information.
[0006] The present invention is realized through the following technical solutions:
[0007] Step 1: Use the U-Netv2 network to perform four-layer downsampling on the input image to generate four-level feature maps of different scales;
[0008] Step 2: Input the above feature map into the multi-scale cross-axis attention module (MSCA), and optimize the expression ability of multi-scale spatial features through cross-dimensional interaction and global perception;
[0009] Step 3: Use the feature fusion module (FFM) to fuse the processed feature maps pairwise, and combine the information integration in the channel dimension to enhance the model's sensitivity to detail and edge features;
[0010] Step 4: Reconstruct the fused features through the decoder and output the enhanced image detection results.
[0011] Furthermore, the specific content of Step 1 is as follows: Select the DRIVE dataset as the benchmark, divide it into a training set (60%) and a test set (40%) according to the ratio of 6:4, and use the U-Netv2 network architecture to perform four-layer downsampling on the input image, and extract feature maps of different scales layer by layer to enhance the model's detection ability for small targets.
[0012] Furthermore, the specific content of Step 2 is as follows: Input the four extracted feature maps into the MSCA module. This module uses a two-way attention mechanism on the horizontal and vertical axes, processes multi-scale features with strip convolution kernels of different sizes, and establishes cross-attention interaction between the two-axis features. Finally, fuse the features through 1×1 convolution and perform residual connection with the input features, significantly improving the detection and localization ability for small targets.
[0013] Furthermore, the multi-scale cross-axis attention (MSCA) module is specifically as follows: The MSCA module performs a series of processes on the four feature maps respectively. First, the MSCA module receives the four feature maps generated by the feature extraction network. For each layer of feature map, horizontal and vertical axis paths are established and processed with strip convolution kernels of different sizes to obtain F x and F y , and the specific calculation expressions are as follows:
[0014]
[0015] where and respectively represent 1D convolution along the x-axis and y-axis directions, Norm(·) represents layer normalization, F represents the input, and the sizes of the 1D convolution kernels are set to 1×7 (7×1), 1×11 (11×1), and 1×21 (21×1) respectively;
[0016] Then continue to calculate the cross-attention between F x and F y . Specifically, for the upper branch, regard F x as the key matrix and value matrix, while F yConsider it as the query matrix, and vice versa for the lower-branch role. Finally, we get and The specific calculation expressions are as follows:
[0017]
[0018] where MHCA y (·, ·, ·) and MHCA x (·, ·, ·) respectively represent the multi-head cross attention along the x-axis and y-axis;
[0019] Finally, apply 1×1 convolution kernels to the horizontal and vertical axis paths respectively to adjust the number of channels to obtain the final output feature map. The specific calculation expressions are as follows:
[0020]
[0021] where Conv1×1 represents 1×1 convolution, and F is the initial input feature.
[0022] Furthermore, the specific content of step three is as follows: By pairwise inputting the multi-scale feature maps processed by the MSCA module into the feature fusion module (FFM), use 3×3 and 5×5 double-branch convolutions to extract local detail and global structure features respectively. After channel attention weighted fusion, perform residual connection with the original features, which not only enhances the model's perception ability of image details and boundaries, but also effectively avoids the problem of feature information loss.
[0023] Furthermore, the specific content of the feature fusion module (FFM) is as follows: FFM processes the feature maps passing through the MSCA module in the following steps. First, splice the feature maps from two different layers along the channels, and apply 1×1 convolution to adjust the number of channels, so that the feature maps can retain important feature information while adapting to subsequent processing. Then, split the spliced feature maps into two parts. The specific formula is as follows:
[0024] F1,F2 = chunk(Conv1×1(cat(f1,f2)))
[0025] where Conv1×1 represents 1×1 convolution, cat represents the splicing operation, and chunk represents the splitting operation;
[0026] Next, a dual-branch parallel processing strategy is adopted to perform 3×3 and 5×5 depthwise separable convolution operations on the segmented feature maps respectively. The former focuses on local detailed feature extraction, and the latter is responsible for capturing large-scale global features. The local and global features are synergistically complemented through element-wise addition operations between the feature maps. Then, the GELU activation function is used to non-linearly enhance the dual-path features, improving the model's recognition ability for complex data patterns. Subsequently, a cross-scale feature fusion operation is performed, and the advantages of local and global features are complemented through element-wise multiplication. To ensure feature integrity, a residual connection mechanism is introduced, adding the original features to the processing results, which not only prevents information loss but also enhances the feature learning effect. Finally, the optimized feature representation is output. The specific formula is as follows:
[0027]
[0028] Where dwconv3 and dwconv5 represent the 3×3 and 5×5 depthwise separable convolutions respectively, and f1 and f2 represent the initial input features.
[0029] Furthermore, the specific content of Step 4 is as follows: The MSCA module is used to deeply extract and integrate cross-scale spatial information. This module uses a hierarchical processing mechanism to perform refined analysis on feature maps of different scales. Subsequently, the FFM module is used to optimize and fuse the features. It innovatively adopts a dual-path processing strategy, enhancing the detail perception ability while retaining the original features, significantly improving the detection accuracy of small targets under complex backgrounds, and finally forming a complete processing link from feature extraction to fusion enhancement, providing a reliable feature representation basis for high-precision image analysis.
[0030] The beneficial effects of the present invention are as follows:
[0031] By combining the multi-scale cross-axis attention (MSCA) module and the feature fusion module (FFM), the present invention can better capture and integrate global and multi-scale spatial information, strengthen the model's perception ability for details and boundaries, effectively improve the missed detection and false detection of the model for small target detection, and provide an innovative solution for the extraction of image feature information. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0033] Figure 1 is the flow diagram of the present invention;
[0034] Figure 2 Schematic diagram of the model structure of the present invention;
[0035] Figure 3 Schematic diagram of the structure of MSCA;
[0036] Figure 4 Schematic diagram of the structure of FFM; Specific implementation manners
[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0038] Please refer to Figures 1-4 As shown, the present invention is an image processing method based on small target detection, including the following steps:
[0039] S101: Use the U-Netv2 network to perform four-layer downsampling on the input image to generate four-level feature maps of different scales;
[0040] S102: Input the above feature maps into the multi-scale cross-axis attention module (MSCA), and optimize the expression ability of multi-scale spatial features through cross-dimensional interaction and global perception;
[0041] S103: Use the feature fusion module (FFM) to fuse the processed feature maps pairwise, and combine the information integration in the channel dimension to enhance the sensitivity of the model to detail and edge features;
[0042] S104: Reconstruct the fused features through the decoder and output the enhanced image detection results.
[0043] As an optimized solution of the above embodiment, the specific step one is: Select the DRIVE dataset as the benchmark, divide it into a training set (60%) and a test set (40%) according to a ratio of 6:4, and use the U-Netv2 network architecture to perform four-layer downsampling processing on the input image, and extract feature maps of different scales layer by layer to enhance the detection ability of the model for small targets.
[0044] As an optimized solution of the above embodiment, the specific step two is: Input the four-layer feature maps extracted into the MSCA module. This module uses the two-way attention mechanism of the horizontal axis and the vertical axis, processes multi-scale features with strip convolution kernels of different sizes, establishes cross-attention interaction between the two-axis features, and finally fuses the features through 1×1 convolution and connects them residually with the input features, significantly improving the detection and positioning ability for small targets.
[0045] As an optimized solution of the above embodiment, the multi-scale cross-axis attention (MSCA) module is specifically as follows: The MSCA module performs a series of processes on four layers of feature maps. First, the MSCA module receives four layers of feature maps generated by the feature extraction network. For each layer of feature map, horizontal and vertical axis paths are established and processed using strip convolution kernels of different sizes to obtain F x and F y , and the specific calculation expressions are as follows:
[0046]
[0047] where and respectively represent 1D convolutions along the x-axis and y-axis directions, Norm(·) represents layer normalization, F represents the input, and the sizes of the 1D convolution kernels are set to 1×7 (7×1), 1×11 (11×1), and 1×21 (21×1);
[0048] Then, continue to calculate the cross-attention between F x and F y . Specifically, for the upper branch, F x is regarded as the key matrix and value matrix, while F y is regarded as the query matrix. For the lower branch, the roles are reversed. Finally, and are obtained. The specific calculation expressions are as follows:
[0049]
[0050] where MHCA y (·,·,·) and MHCA x (·,·,·) respectively represent the multi-head cross-attention along the x-axis and y-axis;
[0051] Finally, 1×1 convolution kernels are applied to the horizontal and vertical axis paths respectively to adjust the number of channels to obtain the final output feature map. The specific calculation expressions are as follows:
[0052]
[0053] where Conv1×1 represents 1×1 convolution, and F is the initial input feature.
[0054] As an optimized solution of the above embodiments, step three is specifically as follows: The multi-scale feature maps processed by the MSCA module are pairwise input into the Feature Fusion Module (FFM). 3×3 and 5×5 dual-branch convolutions are used to extract local detail and global structure features respectively. After channel attention weighted fusion, they are connected with the original feature residuals, which not only enhances the model's perception ability of image details and boundaries, but also effectively avoids the problem of feature information loss.
[0055] As an optimized solution of the above embodiments, the Feature Fusion Module (FFM) is specifically as follows: The FFM processes the feature maps passing through the MSCA module in the following steps. First, the feature maps from two different layers are concatenated along the channels, and 1×1 convolution is applied to adjust the number of channels, so that the feature maps retain important feature information while adapting to subsequent processing. Then, the concatenated feature maps are split into two parts, and the specific formula is as follows:
[0056] F1,F2 = chunk(Conv1×1(cat(f1,f2)))
[0057] Where Conv1×1 represents 1×1 convolution, cat represents the concatenation operation, and chunk represents the splitting operation;
[0058] Next, a dual-branch parallel processing strategy is adopted to perform 3×3 and 5×5 depthwise separable convolution operations on the segmented feature maps respectively. The former focuses on local detail feature extraction, and the latter is responsible for large-range global feature capture. The local and global features are synergistically complemented through element-wise addition operations between the feature maps. Then, the GELU activation function is used to nonlinearly enhance the dual-path features, improving the model's recognition ability for complex data patterns. Subsequently, a cross-scale feature fusion operation is performed, and the advantages of local and global features are complemented through element-wise multiplication. To ensure feature integrity, a residual connection mechanism is introduced, and the original feature is added to the processing result, which not only prevents information loss but also strengthens the feature learning effect, and finally outputs the optimized feature representation. The specific formula is as follows:
[0059]
[0060] Where dwconv3 and dwconv5 represent 3×3 and 5×5 depthwise separable convolutions respectively, and f1 and f2 represent the initial input features.
[0061] As an optimized solution of the above embodiments, step four is specifically as follows: the MSCA module is used to achieve in-depth extraction and integration of cross-scale spatial information. This module uses a hierarchical processing mechanism to perform refined analysis on feature maps of different scales, and then the FFM module is used to optimize and fuse the features. It innovatively adopts a dual-path processing strategy, enhancing the detail perception ability while retaining the original features, significantly improving the detection accuracy of small targets under complex backgrounds, and finally forming a complete processing link from feature extraction to fusion enhancement, providing a reliable feature representation basis for high-precision image analysis.
[0062] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, they are not limitations on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts still fall within the protection scope of the present invention.
Claims
1. An image processing method based on small target detection, characterized in that The described image processing method based on small target detection includes the following steps: Step 1: Use the U-Netv2 network to perform four-layer downsampling on the input image to generate four-level feature maps of different scales; Step 2: Input the above feature maps into the multi-scale cross-axis attention module (MSCA), and through cross-dimensional interaction and global perception, optimize the expression ability of multi-scale spatial features; Step 3: Use the feature fusion module (FFM) to fuse the processed feature maps pairwise, and combine the information integration in the channel dimension to enhance the model's sensitivity to detail and edge features; Step 4: Reconstruct the fused features through the decoder and output the enhanced image detection results.
2. The method according to claim 1, characterized in that The specific content of Step 1 is as follows: Select the DRIVE dataset as the benchmark, divide it into a training set (60%) and a test set (40%) according to a ratio of 6:4, and use the U-Netv2 network architecture to perform four-layer downsampling processing on the input image, and extract feature maps of different scales layer by layer to enhance the model's detection ability for small targets.
3. The method according to claim 1, characterized in that The specific content of Step 2 is as follows: Input the four extracted feature maps into the MSCA module. This module uses a two-way attention mechanism on the horizontal and vertical axes, processes multi-scale features with strip convolution kernels of different sizes, establishes cross-attention interaction between the two-axis features, and finally fuses the features through 1×1 convolution and connects them residually with the input features, significantly improving the detection and localization ability for small targets.
4. The method according to claim 3, wherein The described multi-scale cross-axis attention (MSCA) module is specifically as follows: The MSCA module performs a series of processes on four layers of feature maps. First, the MSCA module receives four layers of feature maps generated by the feature extraction network. Each layer of feature map establishes horizontal and vertical axis paths and is processed using strip convolution kernels of different sizes to obtain F x and F y , and the specific calculation expressions are as follows: Among them and respectively represent 1D convolutions along the x-axis and y-axis directions, Norm(·) represents layer normalization, F represents the input, and the sizes of the 1D convolution kernels are set to 1×7 (7×1), 1×11 (11×1), and 1×21 (21×1); Then continue to calculate F x and F y Calculate the cross-attention between them. Specifically, for the upper branch, regard F x as the key matrix and value matrix, while regard F y as the query matrix. For the lower branch, the roles are reversed. Finally, obtain F x 1 and F y 1 . The specific calculation expressions are as follows: Among them, MHCA y (·,·,·) and MHCA x (·,·,·) respectively represent the multi-head cross attention along the x-axis and y-axis; Finally, apply 1×1 convolution kernels to the horizontal and vertical axis paths respectively to adjust the number of channels to obtain the final output feature map. The specific calculation expression is as follows: Where Conv1×1 represents 1×1 convolution, and F is the initial input feature.
5. The method according to claim 1, characterized in that, The specific content of Step 3 is as follows: Input the multi-scale feature maps processed by the MSCA module pairwise into the feature fusion module (FFM), use 3×3 and 5×5 dual-branch convolutions to extract local detail and global structure features respectively, and after channel attention weighted fusion, connect them residually with the original features, which not only enhances the model's perception ability of image details and boundaries, but also effectively avoids the problem of feature information loss.
6. The method according to claim 5, wherein The specific content of the feature fusion module (FFM) is as follows: FFM performs the following steps on the feature maps processed by the MSCA module in sequence. First, splice two feature maps from different levels by channel and adjust the number of channels through 1×1 convolution to adapt to the subsequent processing flow while retaining key feature information. Subsequently, split the spliced feature map into two parts. The specific formula is as follows: F1,F2 = chunk(Conv1×1(cat(f1,f2))) Where Conv1×1 represents 1×1 convolution, cat represents the splicing operation, and chunk represents the splitting operation; Next, a dual-branch parallel processing strategy is adopted to perform 3×3 and 5×5 depthwise separable convolution operations on the segmented feature maps respectively. The former focuses on local detail feature extraction, and the latter is responsible for large-range global feature capture. The local and global features are synergistically complemented through element-wise addition between the feature maps. Then, the GELU activation function is used to nonlinearly enhance the dual-path features, improving the model's recognition ability for complex data patterns. Subsequently, a cross-scale feature fusion operation is performed, and the advantages of local and global features are complemented through element-wise multiplication. To ensure feature integrity, a residual connection mechanism is introduced, adding the original features to the processing results, which not only prevents information loss but also enhances the feature learning effect. Finally, an optimized feature representation is output. The specific formula is as follows: Where dwconv3 and dwconv5 represent the 3×3 and 5×5 depthwise separable convolutions respectively, and f1 and f2 represent the initial input features.
7. The method according to claim 1, wherein Specifically, step four is as follows: The MSCA module is used to deeply extract and integrate cross-scale spatial information. This module uses a hierarchical processing mechanism to finely analyze feature maps of different scales. Subsequently, the FFM module is used to optimize and fuse the features. It innovatively adopts a dual-path processing strategy, enhancing the detail perception ability while retaining the original features, significantly improving the detection accuracy of small targets in complex backgrounds. Finally, a complete processing link from feature extraction to fusion enhancement is formed, providing a reliable feature representation basis for high-precision image analysis.