Small target detection method based on multi-scale cavity fusion

Through the multi-scale hollow fusion method and lightweight detection head strategy, the problem of excessive computing resource consumption of existing small object detection methods is solved, and more efficient and accurate small object detection is achieved, which is suitable for high-density small object environments.

CN119992390APending Publication Date: 2025-05-13DONGHUA UNIV

Patent Information

Application Number
CN202510166938.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When the existing small-objective detection method deals with high-resolution and dense target scenarios, the computing resources are consumed too much, resulting in low detection speed and efficiency.

Method used

The multi-scale hollow fusion method is adopted to expand the receptive field by setting different hollow rates, and combined with the lightweight detection head strategy, the detection head is redesigned to improve the efficiency and accuracy of small-object detection.

Benefits of technology

Without increasing the amount of calculation, expand the receptive field, improve the efficiency and accuracy of small-objective detection, adapt to the high-density small-objective environment of drone aerial photography, and provide more efficient and resource-friendly solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992390A_ABST
    Figure CN119992390A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision and deep learning, and particularly relates to a small target detection method based on multi-scale hole fusion. The method solves the technical problems that in an unmanned aerial vehicle aerial image target detection task, the proportion of a small target is usually the maximum, the number of effective pixel points contained in the small target is small, and the amount of information sufficient for correct positioning is difficult to obtain under the complex background condition. In traditional target detection, convolution operation enables an image to become smaller and image boundary information to be lost, so that edge pixel points of a small target cannot play a role, and meanwhile, many context features are ignored. The small target detection method based on multi-scale cavity fusion is introduced for the problem, the problem that the small target is difficult to detect can be solved, and the identifiability and the positioning accuracy of the small target in a complex background are improved by expanding a receptive field, enhancing the feature extraction capability of the small target, fusing multi-scale features and improving the utilization of context information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and deep learning technology, and specifically to a small target detection method based on multi-scale hole fusion. Background Art

[0002] Object detection is a core task in image processing and computer vision, which aims to identify and locate specific objects in images or videos. Object detection technology plays a vital role in fields such as autonomous driving, security monitoring, medical image analysis, and intelligent manufacturing. YOLO is an object detection algorithm based on deep learning. It has been widely used in the field of object detection due to its fast and efficient characteristics. Among them, the YOLOv8 algorithm further improves the detection speed and accuracy based on the optimization of YOLOv1 to YOLOv7, and achieves a more balanced performance. By introducing a new network structure and enhanced feature extraction capabilities, YOLOv8 has more advantages in real-time detection.

[0003] In the field of target detection, small targets usually refer to targets that occupy a small pixel area in the image, have low resolution, and have limited feature information. Specifically, small targets are defined as targets with an area less than 32×32 pixels or objects whose width-to-height ratio of the target bounding box to the width-to-height ratio of the image is less than 0.1. Small target detection is very challenging in practical applications because small targets have limited feature information and are easily affected by background interference and image noise, resulting in low detection accuracy, which has become a difficult problem to solve in the field.

[0004] Patent No. CN202410462830.X is a method for detecting small targets in drone aerial photography based on feature transformation and sample optimization. This method provides a method for detecting small targets in drone aerial photography based on feature transformation and sample optimization. An adaptive feature transformation mechanism is designed according to the characteristics of drone aerial photography data, focusing on target-dense areas, and improving the adaptability of the model in dense small target scenes. At the same time, a target-guided sample allocation strategy is adopted, and dynamic sample allocation from coarse to fine is achieved through coarse screening of position information and fine screening guided by target prediction information, effectively improving the detection performance of small targets and dense targets under complex backgrounds. This method excels in accuracy and real-time performance, reduces missed detections and false detections, and achieves more accurate and efficient detection of small targets in drones.

[0005] This computational overhead stems from the implementation of the adaptive feature transformation mechanism and the multi-stage sample screening strategy. Since the features need to be transformed and optimized at multiple levels, the system is prone to speed bottlenecks when processing a large number of images in real time. Although this method improves the adaptability and accuracy of small target detection in drone aerial photography, its main drawback is that the complexity of feature transformation and sample allocation strategy may lead to excessive consumption of computing resources, especially when dealing with high-resolution and dense target scenes, which may place high demands on device hardware. Summary of the invention

[0006] The present invention provides a small target detection method based on multi-scale hole fusion. By introducing a multi-scale hole convolution fusion method and combining it with a lightweight detection head strategy to redesign the detection head, the receptive field can be expanded without increasing the amount of calculation. By setting different hole rates, the system can capture image features at different spatial scales, adapt to the high-density small target environment of drone aerial photography, and provide a more efficient and resource-friendly solution.

[0007] A small target detection method based on multi-scale hole fusion, characterized by comprising the following steps:

[0008] Step S1: construct a small object dataset, using the existing public dataset and filtering out images containing small objects according to the standard of object size 32×32;

[0009] Step S2: Multi-scale hole feature extraction: The input image enters the ML-YOLO model, and multi-scale feature extraction is performed according to the preset hole rate group to obtain feature information of different scales, and feature combination is performed with the conventional convolution branch and the global branch;

[0010] Step S3: Global and local feature information enhancement: The image processed in step S1 is input into the feature information enhancement part; through channel attention and spatial attention in parallel, weights are applied on different dimensions of the features; finally, the weights are readjusted according to the image features processed in step S1 to enhance or suppress the relevant feature information;

[0011] Step S4: feature fusion: the high-dimensional and low-dimensional information output in step S3 are fused through upsampling and downsampling operations to obtain the final fused features;

[0012] Step S5: Classification and positioning of lightweight small target detection head: The final fusion feature in step S3 is input to the lightweight small target detection head for classification and positioning prediction, and the bounding box, category and confidence are output. The final detection result is obtained and output through NMS;

[0013] Step S6: Embed the model into the mobile terminal for deployment and perform small target detection in various scenarios.

[0014] The present invention relates to a small target detection method based on multi-scale hole fusion, which uses multi-scale hole fusion to obtain feature information of different scales, combines features with conventional convolution branches and global branches, and then enhances global and local feature information to improve the detection ability of small targets. Using a lightweight deep differential enhanced convolution detection head can perform effective edge detection on the basis of saving computational costs and memory consumption, and finally achieve the function of being able to detect small targets well.

[0015] The existing target detection algorithm structures such as Faster R-CNN and YOLO series are designed to adapt to the receptive field of medium and large targets, and are therefore widely used in processing target detection tasks. When the prior art CNN network performs downsampling, it gradually reduces the feature map through multi-layer convolution and pooling operations to complete the feature extraction. At the same time, the existing improved algorithm has a large amount of calculation and high memory usage, which puts a huge pressure on hardware resources and seriously affects the running speed and efficiency of small target detection. Starting from the idea of ​​creating small target detection with multi-scale void fusion and reducing computational overhead, the present invention proposes a small target detection method based on multi-scale void fusion to replace the existing target detection algorithm for small target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flow chart of the small target detection method based on multi-scale hole fusion of the present invention.

[0017] Figure 2 It is a structural schematic diagram of the MSFA multi-scale spatial fusion attention module of the present invention.

[0018] Figure 3 The present invention is a structural schematic diagram of a lightweight deep differential enhanced convolution detection head. DETAILED DESCRIPTION

[0019] In the target detection task of drone aerial images, small targets usually account for the largest proportion, and they contain a small number of effective pixels, making it difficult to obtain enough information for correct positioning under complex background conditions. In traditional target detection, the convolution operation will make the image smaller and the image boundary information will be lost, which will cause the edge pixels of small targets to be unable to function and ignore many contextual features. In order to solve this problem, the present invention introduces a small target detection method of multi-scale void fusion, which can solve the problem that small targets are difficult to detect. By expanding the receptive field, enhancing the feature extraction capability of small targets, fusing multi-scale features, and improving the utilization of contextual information, the recognizability and positioning accuracy of small targets in complex backgrounds are improved.

[0020] The technical solution of the embodiment of the present invention is further described below in conjunction with the accompanying drawings.

[0021] like Figure 1 As shown, a small target detection method based on multi-scale hole fusion includes the following steps:

[0022] Step S1: Construct a small object dataset, using existing public datasets to filter out images containing small objects according to the standard of object size 32×32.

[0023] Step S2: Multi-scale hole feature extraction: The input image enters the ML-YOLO model, and multi-scale feature extraction is performed according to the preset hole rate group to obtain feature information of different scales, and the features are combined with the conventional convolution branch and the global branch.

[0024] Specific content includes:

[0025] In the target detection task of drone aerial images, small targets usually account for the largest proportion, and they contain a small number of valid pixels. Ordinary convolution operations will make the image smaller and the image boundary information will be lost. Therefore, it is necessary to set multi-scale void ratios to expand the convolution receptive field so that the edge pixels of small targets can play a role.

[0026] Figure 2 The proposed MSFA multi-scale spatial fusion attention module is presented. This module can extract multi-scale hole features. It consists of five branches: a 1×1 convolution, three 3×3 convolutions with different hole rates, global flat pooling, channel attention, and spatial attention. In MSFA, the input image X and a convolution kernel K, the convolution with hole rate d, its operation can be expressed as:

[0027]

[0028] Among them, i and j are the pixel positions of the output feature map, m and n are the indices of the convolution kernel, and k and l are the radii of the convolution kernel in the horizontal and vertical directions, respectively. Furthermore, the calculation of the receptive field F takes into account the size and expansion rate of the convolution kernel. Specifically, the receptive field can be calculated by the following formula:

[0029]

[0030] Among them, RF l+1 is the receptive field size of the current feature map, RF l is the receptive field size of the feature map of the previous layer, f l+1 is the current convolution kernel size, Represents the stride product of the previous convolutional layer.

[0031] Furthermore, a global pooling branch is set up in MSFA in parallel with the dilated convolution, which will extract the global features of the entire image and help the model understand the overall scene structure.

[0032] Step S3: Global and local feature information enhancement: Input the processed image from step S1 to the feature information enhancement part; weight the features in different dimensions by using channel attention and spatial attention in parallel. Finally, readjust the weights according to the processed image features in step S1 to enhance or suppress the relevant feature information.

[0033] Specific content includes:

[0034] For channel attention, a given feature map Perform global average pooling on each channel to obtain the global feature X of each channel C c ,

[0035]

[0036] Next, a 1×1 convolution layer is applied to reduce the dimension of the input channel and output a single-channel feature map. The weight matrix of the 1×1 convolution is

[0037]

[0038] Subsequently, the ReLU activation function is applied to the convolution result to introduce nonlinearity.

[0039]

[0040] Then, the result after ReLU activation is adjusted to the same spatial dimension H×W as the input feature map X through interpolation operation, and the interpolation method is nearest neighbor interpolation. Finally, the upsampled weight Y upsampled It is expanded to the same dimension as the input X (i.e., B×C×H×W) through the broadcast mechanism and multiplied element-by-element with the input feature map to achieve weighting of each channel.

[0041]

[0042] It enhances the expression of features by evaluating the importance of different channels, strengthens useful feature channels, and suppresses the negative impact of redundant or unimportant channels, thereby ensuring the quality of feature representation.

[0043] For the spatial attention mechanism, we first use a 1×1 convolution to reduce the channel dimension of the input feature map X from C to 1 to generate a spatial attention map.

[0044]

[0045] Next, the Sigmoid function is applied to normalize the output of the 1×1 convolution to obtain a weight map between [0,1].

[0046]

[0047] Finally, the generated spatial attention weight Y sigmoid Expand to the channel dimension C of the input feature map, make it the same dimension as the input feature map X through the broadcast mechanism, and then multiply element by element.

[0048]

[0049] This process essentially weights the features at each position, with the weights determined by the spatial attention at that position, helping the model better focus on important spatial areas.

[0050] Step S4: Feature fusion: The high-dimensional and low-dimensional information output in step S3 are fused through upsampling and downsampling operations to obtain the final fused features.

[0051] Specific content includes:

[0052] In order to combine channel attention and spatial attention together, the merging module adopts the operation of taking the maximum value of each element, that is, the feature map U after the weighting of spatial attention and channel attention kongjian and U tongdao , take the maximum value element by element:

[0053] U out,b,c,i,j =max(U kongjian,b,c,i,j ,U tongdao,b,c,i,j )

[0054] After obtaining the output of the attention mechanism, it is necessary to effectively combine it with multi-scale features to achieve the goal of focusing on local features while taking into account global information, and enhancing the model's ability to detect small objects in the image. Finally, MSFA will use 1×1 convolution to reduce the dimension of features in order to reduce the amount of subsequent calculations and integrate information from various branches.

[0055] Step S5: Classification and positioning of lightweight small target detection head: The final fusion feature in S3 is input into the lightweight small target detection head for classification and positioning prediction, and the bounding box, category and confidence are output. The final detection result is obtained and output through NMS.

[0056] Specific content includes:

[0057] Lightweight depth difference enhanced convolution detection head LDEDH (Lightweight detail enhanced detection head) such as Figure 3 As shown in the figure, the original yolov8 detection head was redesigned by adding a small target detection head, DDEB, and replacing some ordinary convolutions with normalized convolutions.

[0058] Small target detection head:

[0059] There are three detection heads by default in the Yolov8 network structure, which can perform multi-scale target detection. The sizes are: the detection feature map size corresponding to P3 is 80pixel×80pixel, which is used to detect targets larger than 8pixel×8pixel; the detection feature map size corresponding to P4 is 40pixel×40pixel, which is used to detect targets larger than 16pixel×16pixel. The input image size of the model is 640pixel×640pixel. The three detection heads meet the basic detection requirements, but in actual use, the detection effect is still poor for small targets. To this end, this paper adds a 160pixel×160pixel small target detection head to the model, which concatenates the upsampled feature map with the P2 layer feature map of the backbone network. The P2 layer is usually a high-resolution feature map used to capture more spatial detail information. This feature fusion method can comprehensively utilize deep semantic information and shallow detail information, thereby enhancing the performance of small target detection.

[0060] Deep difference enhancement module DDEB:

[0061] The depthwise difference enhancement block (DDEB) enhances the feature extraction capability by introducing differential convolution (DC) and depthwise separable convolution (DSConv). DSConv is used to further optimize the feature extraction process. Depthwise convolution extracts features in the spatial dimension and operates each channel independently to reduce computational overhead. All features are then aggregated into global features through 1×1 convolution, so that features between different channels can interact with each other, improving the feature expression capability.

[0062] Differential convolution can more sensitively capture the edges and details of the target by calculating pixel differences. Combining them can enhance the representation and generalization capabilities of vanilla convolution. DDEB uses differential convolution (CDC, ADC, HDC, VDC) in multiple stages of feature extraction, where:

[0063] Centered Differential Convolution (CDC): CDC increases sensitivity to edges by calculating the difference between the center point and neighboring pixels. The calculation formula is as follows:

[0064]

[0065] where x(p0) is the value of the center pixel, x(p0+p n), is the value of the pixel in the neighborhood, w(p n ) is the convolution kernel weight, by calculating x(p0+p n )-x(p0), CDC directly captures the relative intensity difference with the central pixel at each pixel position, reflecting the local change of pixel value, which is particularly obvious in the edge area.

[0066] Angular differential convolution (ADC): ADC performs convolution on the difference between pixels in the diagonal direction to capture richer directional information, especially edge and texture information along the diagonal direction. The calculation formula is as follows,

[0067]

[0068] where x(p0+p n ) and x(p0-p n ) represent the adjacent pixels of the center point in the diagonal direction. ADC is more sensitive to oblique edges or structures because it can identify subtle changes along the diagonal direction and is suitable for detecting inclined edges or more complex texture information.

[0069] Horizontal Differential Convolution (HDC): HDC calculates the difference of pixel pairs in the horizontal direction to extract horizontal edge features, so that the convolution kernel can more effectively detect edge changes in the horizontal direction. The calculation formula is as follows,

[0070]

[0071] where p n ∈R horizontal Indicates taking neighboring pixels in the horizontal direction, x(p0+p n ) and x(p0-p n ) are the left and right adjacent pixel values ​​respectively.

[0072] Vertical differential convolution (VDC): VDC captures vertical gradient information by calculating the difference in the vertical direction, thereby enhancing the convolution layer's ability to perceive vertical edges. Its formula is similar to HDC, but the direction is vertical. The calculation formula is as follows:

[0073]

[0074] where p n ∈R vertical It means that only the neighboring pixels are taken in the vertical direction, x(p0+p n ) and x(p0-p n ) are the upper and lower adjacent pixel values ​​respectively.

[0075] Step S6: Embed the model into the mobile terminal for deployment and perform small target detection in various scenarios.

[0076] The small target detection method based on multi-scale hole fusion proposed in the present invention uses multi-scale hole fusion to obtain feature information of different scales, combines features with conventional convolution branches and global branches, and then enhances global and local feature information to improve the detection ability of small targets. The use of lightweight deep differential enhanced convolution detection head can perform effective edge detection on the basis of saving computational cost and memory consumption, and finally realize the function of being able to detect small targets well.

Claims

1. A small target detection method based on multi-scale hole fusion, characterized in that The steps include: Step S1: Construct a small object dataset, use the existing public dataset, and filter out images containing small objects according to the standard of object size 32×32; Step S2: Multi-scale hole feature extraction: The input image enters the ML-YOLO model, and multi-scale feature extraction is performed according to the preset hole rate group to obtain feature information of different scales, and feature combination is performed with the conventional convolution branch and the global branch; Step S3: Global and local feature information enhancement: The processed image in step S1 is input into the feature information enhancement part; through channel attention and spatial attention in parallel, weighting is performed on different dimensions of the feature; Finally, the weights are readjusted according to the image features processed in step S1 to strengthen or suppress relevant feature information; Step S4: feature fusion: the high-dimensional and low-dimensional information output in step S3 are fused through upsampling and downsampling operations to obtain the final fused features; Step S5: Classification and positioning of lightweight small target detection head: The final fusion feature in step S3 is input to the lightweight small target detection head for classification and positioning prediction, and the bounding box, category and confidence are output. The final detection result is obtained and output through NMS; Step S6: Embed the model into the mobile terminal for deployment and perform small target detection in various scenarios.

2. The small target detection method based on multi-scale hole fusion according to claim 1 is characterized in that The specific process of the above step S2 is: The MSFA multi-scale spatial fusion attention module is used for multi-scale hole feature extraction, which consists of a 1×1 convolution, three 3×3 convolutions with different hole rates, and five branches of global flat pooling, channel attention, and spatial attention. The input image X and a convolution kernel K and a hole rate d convolution in MSFA are expressed as: Among them, i and j are the pixel positions of the output feature map, m and n are the indices of the convolution kernel, k and l are the radii of the convolution kernel in the horizontal and vertical directions respectively; the calculation of the receptive field F takes into account the size and expansion rate of the convolution kernel; The receptive field is calculated using the following formula: Among them, RF l+1 is the receptive field size of the current feature map, RF l is the receptive field size of the feature map of the previous layer, fl +1 is the current convolution kernel size, Represents the stride product of the previous convolutional layer.

3. The small target detection method based on multi-scale hole fusion according to claim 2 is characterized in that The specific process of the above step S3 is: For channel attention, a given feature map Perform global average pooling on each channel to obtain the global feature X of each channel C c , Next, a 1×1 convolution layer is applied to reduce the dimension of the input channel and output a single-channel feature map. The weight matrix of the 1×1 convolution is Subsequently, the ReLU activation function is applied to the convolution result to introduce nonlinearity. Then, the result after ReLU activation is adjusted to the same spatial dimension H×W as the input feature map X through interpolation operation, and the interpolation method is nearest neighbor interpolation. Finally, the upsampled weight Y upsampled The dimension of the input X is expanded to the same as that of the input X, that is, B×C×H×W, through the broadcast mechanism, and multiplied element by element with the input feature map to achieve weighting of each channel; It enhances the expression of features and strengthens useful feature channels by evaluating the importance of different channels; For the spatial attention mechanism, we first use a 1×1 convolution to reduce the channel dimension of the input feature map X from C to 1 to generate a spatial attention map; Next, the Sigmoid function is applied to normalize the output of the 1×1 convolution to obtain a weight map between [0, 1]; Finally, the generated spatial attention weight Y sigmoid Expand to the channel dimension C of the input feature map, make it the same dimension as the input feature map X through the broadcast mechanism, and then multiply element by element; The features at each position are weighted by the spatial attention at that position, so as to better focus on important spatial regions.

4. The small target detection method based on multi-scale hole fusion according to claim 3 is characterized in that The specific process of the above step S4 is: In order to combine channel attention and spatial attention together, the merging module adopts the operation of taking the maximum value of each element, that is, the feature map U after the weighting of spatial attention and channel attention kongjian and U tongdao , take the maximum value element by element: U out,b,c,i,j =max(U kongjian,b,c,i,j ,U tongdao,b,c,i,j ) After obtaining the output result of the attention mechanism, it is effectively combined with the multi-scale features to achieve the goal of paying attention to local features while taking into account global information, enhancing the model's ability to detect small targets in the image; finally, MSFA will use 1×1 convolution to perform feature dimensionality reduction to reduce the amount of subsequent calculations and integrate information from various branches.

5. The small target detection method based on multi-scale hole fusion according to claim 4 is characterized in that The specific process of the above step S5 is: The original yolov8 detection head was redesigned, a small target detection head was added, and some ordinary convolutions were replaced with normalized convolutions; Small target detection head The Yolov8 network structure has three detection heads by default for multi-scale target detection. A 160 pixel × 160 pixel small target detection head is added, which concatenates the upsampled feature map with the P2 layer feature map of the backbone network. The P2 layer is usually a high-resolution feature map used to capture more spatial detail information. The Deep Differential Enhancement module uses differential convolutions at multiple stages of feature extraction, where: Center difference convolution: By calculating the difference between the center point and the neighboring pixels, the sensitivity to the edge is enhanced; the calculation formula is as follows, where x(p0) is the value of the center pixel, x(p0+p n ) is the value of the pixel in the neighborhood, w(p n ) is the convolution kernel weight, by calculating x(p0+p n )-x(p0), the central difference convolution directly captures the relative intensity difference with the central pixel at each pixel position, reflecting the local change of pixel value; Angular differential convolution: Convolution is performed by taking the difference between pixels in the diagonal direction to capture richer directional information, especially the edge and texture information along the diagonal direction. The calculation formula is as follows: where x(p0+p n ) and x(p0-p n ) represent the adjacent pixels of the center point in the diagonal direction respectively; Horizontal differential convolution: Calculate the horizontal pixel difference to extract horizontal edge features, so that the convolution kernel can more effectively detect horizontal edge changes; the calculation formula is as follows, where p n ∈R horizontal Indicates taking neighboring pixels in the horizontal direction, x(p0+p n ) and x(p0-p n ) are the left and right adjacent pixel values ​​respectively; Vertical differential convolution: By calculating the difference in the vertical direction, the vertical gradient information is captured, thereby enhancing the convolution layer's perception of vertical edges; the calculation formula is as follows: where p n ∈R vertical It means that only the neighboring pixels are taken in the vertical direction, x(p0+p n ) and x(p0-p n ) are the upper and lower adjacent pixel values ​​respectively.

Citation Information

Patent Citations

  • Unmanned aerial vehicle aerial photography small target detection method based on feature transformation and sample optimization

    CN118072205A

Cited By

  • Belt coal falling detection method for resource-constrained equipment

    CN120259636A

  • Image enhancement device based on bionic vision, underwater target detection system and method

    CN120355593A

  • Rural house outer wall crack detection method

    CN120655992A

  • Small target feature processing method and system based on multi-scale convolution

    CN120726440A

  • Scale perception progressive learning method for narrow segmentation of X-ray coronary angiography

    CN121353679A