Faster Rcnn-S-based small target detection method
By optimizing the Backbone, Neck and RPN modules of the FasterRcnn-S network, combining feature fusion and decoupling detection heads, the problems of fewer features, environmental interference and positioning accuracy in small target detection are solved, and the detection effect is improved.
Patent Information
- Application Number
- CN202510643192.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-26
AI Technical Summary
In the detection of small targets, existing algorithms have problems such as few features, susceptible to environmental interference, high positioning accuracy requirements, unbalanced sample and inappropriate network structures to the characteristics of small targets, resulting in poor detection results.
FasterRcnn-S network is adopted, combined with MDLA multi-scale decoupling large-core convolution hybrid attention module, FEAFPN feature enhanced alignment dual-stage feature pyramid network and CRPN multi-level regional suggestion network, fine-grained decoupling detection head is designed, Backbone, Neck and RPN modules are optimized, and feature fusion and positioning capabilities are enhanced.
It improves the accuracy and accuracy of small object detection, reduces feature aberration and deviation, enhances the model's detailed feature acquisition ability at various levels, and optimizes the quality of candidate areas.
Smart Images

Figure CN120544004A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of image processing technology and computer vision, and particularly relates to a small target detection method based on FasterRcnn-S. Background Art
[0002] Small target detection is widely used and crucial in the real world. In autonomous driving, it helps vehicles accurately capture tiny obstacles that could cause traffic accidents in high-resolution driving scenes, significantly improving driving safety. In smart healthcare, small target detection technology can be used to identify tiny lesions or abnormalities in medical images, providing doctors with more accurate diagnostic evidence. Furthermore, in industrial automation, this technology can effectively locate subtle, imperceptible defects on product surfaces, ensuring product quality. Small target detection is also crucial for satellite remote sensing image analysis, helping government agencies accurately identify tiny objects in images, such as small boats or vehicles. This is crucial for combating illegal activities and maintaining public safety.
[0003] Small target detection has the basic characteristics of a small pixel ratio, a small coverage area, and little information, and faces the following challenges: there are few available features. Low-resolution small targets have little visual information, making it difficult to extract discriminative features, and are easily interfered with by environmental factors, which makes it difficult for detection models to accurately locate and identify small targets. High positioning accuracy is required. During the prediction process, a one-pixel offset in the predicted bounding box has a much greater impact on small targets than on large / medium-scale targets. There is also the problem of sample imbalance. To locate the position of the target in the image, most existing methods pre-generate a series of anchor boxes at each position in the image. During training, a fixed threshold is set to determine whether the anchor box belongs to a positive or negative sample. This approach leads to an imbalance in positive samples for targets of different sizes during model training. Due to network structure reasons, in the field of target detection, the design of existing algorithms often focuses on the detection performance of large / medium-scale targets. There are few optimization designs for the characteristics of small targets. Coupled with the difficulties brought by the inherent characteristics of small targets, existing algorithms generally perform poorly in small target detection. Summary of the Invention
[0004] In order to solve the above problems, the present invention provides a small target detection method based on FasterRcnn-S.
[0005] To achieve the above-mentioned purpose, the present invention is implemented through the following technical solutions: The present invention provides a small target detection method based on FasterRcnn-S, comprising the following steps: S1. Obtain the TinyPerson small object detection dataset, perform data augmentation on the dataset, and divide it into a training set and a test set; S2. Construct a small target detection model based on FasterRcnn-S. The model uses the FasterRcnn network as the basic network, including the BackBone module, the Neck module, the RPN module, and the Head module. The BackBone module adopts the RseNet-50 neural network structure, and adds the MDLA multi-scale decoupled large kernel convolution hybrid attention module after each stage from Stage 1 to Stage 4. The Neck module uses the FEAFPN feature enhancement and alignment two-level feature pyramid network for feature fusion. The RPN module introduces the CRPN multi-level region proposal network structure. The Head module designs a fine-grained decoupled detection head FGDHead to replace the ordinary target detection head. S3. Use the training set to train the small target detection model based on FasterRcnn-S to obtain a trained model; S4. Input the images in the test set into the trained model to obtain the target detection results.
[0006] Furthermore, the data enhancement in step S1 includes cropping, stacking, scaling, and transformation.
[0007] Furthermore, in step S2, the Backbone module specifically includes: Add an MDLA multi-scale decoupled large kernel convolution hybrid attention module after each Stage block of the Resnet-50 residual structure in the Backbone module; the MDLA multi-scale decoupled large kernel convolution hybrid attention module includes the MDLConvs multi-scale decoupled large kernel convolution module and the AFM attention feedback module; The MDLConvs multi-scale decoupled large kernel convolution module includes a first regularization layer, a first convolution layer, four parallel decoupled large kernel convolution layers, a second convolution layer and a GELU activation function; the first convolution layer includes a 1×1 convolution kernel and a 5×5 convolution kernel; the four parallel decoupled large kernel convolution layers are DLKConv5 convolution layer, DLKConv11 convolution layer, DLKConv17 convolution layer and DLKConv23 convolution layer from top to bottom; the The second convolutional layer is a 1×1 convolution kernel. The DLKConv5 convolution layer stacks two (3,1) depth-expanded convolutions. The DLKConv11 convolution layer stacks (3,1) and (5,2) depth-expanded convolutions. The DLKConv17 convolution layer stacks (5,1) and (7,2) depth-expanded convolutions. The DLKConv23 convolution layer stacks (5,1) and (7,3) depth-expanded convolutions. The first parameter in the brackets indicates the size of the convolution kernel. , the second parameter represents the expansion rate of the convolution ; The calculation formula for the stacked depth dilated convolution receptive field is: , in, represents the bottom receptive field of the stack, Represents the receptive field of the previous layer of the stack, represents the stride of convolution, represents the size of the convolution kernel, represents the subtraction operation, represents the multiplication operation, Represents addition operation; The output feature x of the Stage block is processed by the first regularization layer to obtain the first regularized feature ; The first regularization feature After the first convolution layer, the first convolution feature is obtained ; The first convolution feature Split the features into 4 groups along the channel dimension to get the DLKConv5 convolution layer input features. , DLKConv11 convolutional layer input features , DLKConv17 convolutional layer Input features and DLKConv23 convolutional layer input features , After processing by four parallel decoupled large kernel convolution layers, the processed features are spliced to obtain the spliced features ; Splicing features The features output by the second convolutional layer and the GELU activation function are added to the output features x of the Stage block to obtain the output features of the MDLA multi-scale decoupled large kernel convolution hybrid attention module. ; The formula for the above process is as follows: , , , , , in, represents the operation of the regularization layer, represents a convolution operation with a convolution kernel size of 5×5. Represents a convolution operation with a convolution kernel size of 1×1, Indicates that the features are split into 4 groups evenly along the channel dimension. Indicates the number of channels, Indicates splitting features along the channel dimension, Represents the operation of the DLKConv5 convolutional layer, Represents the operation of the DLKConv11 convolutional layer, Represents the operation of the DLKConv17 convolutional layer, Represents the operation of the DLKConv23 convolutional layer, Concat represents the concatenation operation, and GELU represents the GELU activation function; The AFM attention feedback module includes a second regularization layer, a local attention module, and a global attention module; the local attention module includes a first branch and a second branch; the first branch includes a 1×1 convolution kernel and a 3×3 convolution kernel, and the second branch includes a 1×1 convolution kernel and a Sigmoid activation function; the global attention module includes a third convolution layer, a Softmax function, and a fourth convolution layer; the third convolution layer includes a 1×1 convolution kernel and a 3×3 depth convolution kernel, and the fourth convolution layer is a 1×1 convolution kernel; Output features of MDLA multi-scale decoupled large kernel convolution hybrid attention module After processing by the second regularization layer, the second regularization feature is obtained ; The first regularization feature After the first branch, the output of the first branch is obtained, and the second regularization feature The output of the second branch is obtained through the second branch, and the output of the first branch is multiplied by the output of the second branch to obtain the local attention feature. ; The second regularization feature After the third convolution layer, query Q, key K and value V are generated through 1×1 convolution and 3×3 depth convolution, and query Q, key K and value V are shaped into 、 as well as ,according to 、 and Calculate the attention map , attention map After the fourth convolution layer and the second regularization feature Add up to get the global attention feature ; The formula is as follows: , , , in, Represents a depthwise convolution operation with a convolution kernel size of 3×3; Represent the feature matrices after query Q, key K and value V are shaped, Represents a shaping operation, represents the Softmax activation function, represents the dot product operation, Represents the division operation, Represents the control parameter; local attention features and global attention features The concatenation is performed and processed in sequence through 1×1 convolution kernel, GELU activation function, 1×1 convolution kernel, and then the output features of the MDLA multi-scale decoupled large kernel convolution mixed attention module are obtained. Add up to get the output features of the AFM attention feedback module , the formula is: .
[0008] Furthermore, the Neck module in step S2 is specifically: The FEAFPN feature-enhanced aligned two-level feature pyramid network includes a top-down primary fusion path and a bottom-up secondary fusion path; In the first fusion path from top to bottom, receive the output features of Backbone , the calculation formula of the fusion path from top to bottom is as follows: in, represents the i-th level feature after fusion from bottom to top, Represents the i-1th level feature, which is the multi-level feature obtained by processing the input image through the Backbone module ={ ,i=2,3,4,5}, It is the feature generated after the output from stage1 to stage4 passes through the above-mentioned MDLA multi-scale decoupling large kernel convolution mixed attention module; FEA represents the FEA feature enhancement alignment module. Represents the i-1th level feature after fusion; the fusion process from top to bottom is: , and so on, we get the top-down fusion feature { ,i=2,3,4,5}; In the bottom-up fusion path, the features fused from top to bottom are received { ,i=2,3,4,5}, the top-down fusion path calculation formula is as follows: , in, Represents the i-th level feature after fusion from bottom to top, Represents the i-1th level features after fusion from bottom to top, Represents the APF adaptive perception fusion module; the fusion process from bottom to top is as follows: , , and so on, we get the bottom-up fusion features { ,i=2,3,4,5}.
[0009] Furthermore, the FEA feature enhancement and alignment module includes the FRE feature refinement and enhancement module and the upsampling module. The specific operations are as follows: The fused i-th level features After being processed by the FRE feature refinement enhancement module, the output features of the FRE feature refinement enhancement module are obtained. ;right Perform upsampling matching The feature size is then learned, and the offset size between the two levels of features is learned. Then, deformable convolution is performed according to the offset size to obtain the aligned up-sampled features. The calculation formula is as follows: in, represents the upsampled features of the i-th level features after the FRE module during fusion, Represents the output of the i-1th layer of the Backbone module, Represents the learned alignment offset, which is a standard convolution with a kernel size of 3×3. Indicates the offset between the features of two layers. represents the deformable convolution operation, Represents the aligned upsampled features, i.e., the output features of the FEA feature enhanced alignment module.
[0010] Furthermore, the specific operations of the APF adaptive perception fusion module are: The i-1th level features after bottom-up fusion The kernel size is 3 The strided convolution operation of 3, the result of the strided convolution is the same as Splicing is performed by dimension, and the spliced features are processed with a kernel size of 1. The convolution of 1 adjusts the channel dimension to 2, and performs a Softmax operation to obtain the second spatial attention map. , then use The features output by the strided convolution are multiplied by the second spatial attention map. Multiply it by the second spatial attention map, and then add the two features multiplied by the second spatial attention map to obtain the output feature of the APF adaptive perception fusion module. The calculation formula is as follows: , , , in Represents the i-1th level feature after fusion from bottom to top, represents the strided convolution operation, express The output obtained by strided convolution is Indicates taking out the spatial attention map of the first channel, Indicates taking out the spatial attention map of the second channel, It represents the i-th level feature after fusion from bottom to top, that is, the output feature of the APF adaptive perception fusion module.
[0011] Furthermore, the specific process of the first stage of the CRPN multi-level region proposal network structure in the RPN part in step S2 is as follows: the output of the Neck part After the convolution operation of the RPN network, the first regression offset of the anchor frame is obtained ; By first regression offset Calculate the anchor frame position and get the corrected anchor frame , thus obtaining the deviation o between the corrected anchor box and the initial anchor box; correcting the features through adaptive convolution operation Align it with the corrected anchor frame to obtain the features corrected by adaptive convolution .
[0012] Furthermore, the specific process of the second stage of the CRPN multi-level region proposal network structure in the RPN part in step S2 is as follows: the features corrected by adaptive convolution are subjected to convolution operation to obtain the classification score cls and the second regression offset , using the second regression offset Correct the anchor frame again to get the final output fine anchor frame , the final output fine anchor box Mapping back to the original input image to obtain the final generated image object candidate box .
[0013] Furthermore, the Head module in step S2 specifically includes: The fine-grained decoupling detection head FGDHead includes a classification branch and a regression branch; the classification branch includes two fully connected layers and one classification fully connected layer; the regression branch includes a convolution layer with a convolution kernel size of 3×3, an MDC multi-scale detail capture module and a regression fully connected layer; The final generated image object candidate box After the RoiPooling pooling layer extracts the fixed-size feature map, the pooling feature is obtained ; Pooling features The classification output of the detection head is obtained through the classification branch ; Pooling features The bounding box output of the detection head regression is obtained through the regression branch .
[0014] Furthermore, the MDC multi-scale detail capture module of the regression branch in the Head module in step S2 is specifically implemented as follows: The pooling features The intermediate features obtained by the 3×3 convolution layer Divide into four groups along the channel dimension and get the first feature , the second feature , the third characteristic And the fourth characteristic ; The first processed features are obtained by processing with depthwise separable convolutions of kernel sizes of 3×3, 13×1, and 1×13. , the second processing feature , the third processing feature , Do nothing; 、 、 as well as Perform splicing to obtain the spliced features ; Features after splicing After a normalization layer and MLP layer, the intermediate features Add up to get the output of the MDC multi-scale detail capture module .
[0015] The advantages of the present invention are: The present invention adds a designed MDLA multi-scale decoupled large kernel convolution mixed attention module behind each Stage block of the Resnet-50 residual structure network in the Backbone module, which can expand the receptive field and better capture the contextual information of small targets; the FEAFPN feature enhanced aligned two-level feature pyramid network designed in the Neck part can better realize the multi-scale fusion of small target information, reduce deviation and feature misalignment, and the precise positioning signal of the low layer can be better transmitted upward, enhancing the positioning expression ability of the overall feature pyramid. The model can obtain accurate detail features at all levels, thereby improving the effect of detecting small target objects; the multi-level CRPN structure is introduced in the RPN part to improve the quality of the candidate area to optimize the performance of small target detection; the designed fine-grained decoupled detection head FGDHead is used in the Head module to replace the ordinary detection head. By separating the classification and regression tasks into two different heads, the mutual interference between tasks is reduced. This design enables the model to focus more on its respective tasks, thereby improving the accuracy of small target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0017] Figure 1 is a flow chart of the steps of the method of the present invention; Figure 2 The following is a comparison of detection images between the method of the present invention and the existing method. DETAILED DESCRIPTION
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments derived by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0019] Example 1 In this embodiment, Figure 1 As shown, the present invention provides a small target detection method based on FasterRcnn-S, which specifically includes the following steps: S1. Obtain the TinyPerson small object detection dataset, perform data augmentation on the dataset, and divide it into a training set and a test set; Specifically, the data enhancement includes cropping, stacking, scaling, and transformation; each image is cropped into a patch of 640×512 pixels with 30 pixels overlapping; the image is moved by 6.25% of the image size in the horizontal and vertical directions and rotated 45 degrees, bilinear interpolation is used for image transformation, and a 50% probability of applying this transformation is set for the image; the brightness of the image is randomly adjusted in the range of -10% to 30%, the contrast of the image is randomly adjusted in the range of 10% to 30%, and a 40% probability of applying this transformation is set; the values of the R, G, and B channels of the image are randomly adjusted in the range of ±10, and the hue, saturation, and brightness values of the image are randomly adjusted in the range of The ranges are between ±20, ±30 and ±20 respectively. One of the two transformations will be randomly selected and applied, and a 10% probability of applying this transformation is specified. The lower limit of the image compression quality is 85, the upper limit of the image compression quality is 95, and a 20% probability of applying this transformation is specified. The R, G, and B channels of the image are randomly shuffled, and a 10% probability of applying this transformation is specified. Gaussian blur is applied to the image, median blur is applied to the image, and motion blur is applied to the image, with a maximum blur degree of 3. One of the three blur transformations will be randomly selected and applied, and a 10% probability of being applied is specified. The adjusted image dataset is divided into a training set and a test set in a ratio of 7:3.
[0020] S2. Construct a small target detection model based on FasterRcnn-S. The model uses the FasterRcnn network as the basic network, including the BackBone module, the Neck module, the RPN module, and the Head module. The BackBone module adopts the RseNet-50 neural network structure, and adds the MDLA multi-scale decoupled large kernel convolution hybrid attention module after each stage from Stage 1 to Stage 4. The Neck module uses the FEAFPN feature enhancement and alignment two-level feature pyramid network for feature fusion. The RPN module introduces the CRPN multi-level region proposal network structure. The Head module designs a fine-grained decoupled detection head FGDHead to replace the ordinary target detection head. Specifically in the Backbone module: Add an MDLA multi-scale decoupled large kernel convolution hybrid attention module after each Stage block of the Resnet-50 residual structure in the Backbone module; the MDLA multi-scale decoupled large kernel convolution hybrid attention module includes the MDLConvs multi-scale decoupled large kernel convolution module and the AFM attention feedback module; The MDLConvs multi-scale decoupled large kernel convolution module includes a first regularization layer, a first convolution layer, four parallel decoupled large kernel convolution layers, a second convolution layer and a GELU activation function; the first convolution layer includes a 1×1 convolution kernel and a 5×5 convolution kernel; the four parallel decoupled large kernel convolution layers are DLKConv5 convolution layer, DLKConv11 convolution layer, DLKConv17 convolution layer and DLKConv23 convolution layer from top to bottom; the The second convolutional layer is a 1×1 convolution kernel. The DLKConv5 convolution layer stacks two (3,1) depth-expanded convolutions. The DLKConv11 convolution layer stacks (3,1) and (5,2) depth-expanded convolutions. The DLKConv17 convolution layer stacks (5,1) and (7,2) depth-expanded convolutions. The DLKConv23 convolution layer stacks (5,1) and (7,3) depth-expanded convolutions. The first parameter in the brackets indicates the size of the convolution kernel. , the second parameter represents the expansion rate of the convolution ; The calculation formula for the stacked depth dilated convolution receptive field is: , in, represents the bottom receptive field of the stack, Represents the receptive field of the previous layer of the stack, represents the stride of convolution, represents the size of the convolution kernel, represents the subtraction operation, represents the multiplication operation, Represents the addition operation; for example: the designed DLKConv is a stack of two depth-wise dilated convolutions. To calculate the actual receptive field of DLKConv, we need to work from back to front. For example, the DLKConv11 described in this article is obtained by stacking depth-wise dilated convolutions (3,1) and (5,2). Then, based on the receptive field of the depth-wise dilated convolution, we first calculate RF2, which is the receptive field of (5,2). 5+(5-1)×(2-1)=9, which is calculated to be 9. Then, we calculate the receptive field RF1 of the bottom layer, which is the receptive field of the DLKConv proposed in this invention. According to the formula for calculating the receptive field of stacked depth-wise dilated convolutions, the stride and ksize are 1 and 3 respectively, corresponding to (3,1). (9-1)×1+3=11, which is calculated to be 11, which is the receptive field of DLKConv11. In other words, we first calculate the receptive field of the last convolution according to the receptive field calculation formula of the depth-wise dilated convolution, and then use the stacked receptive field calculation formula to calculate the overall receptive field from back to front. Simply put, in DLKConv, the receptive field of the second convolution is calculated using the depth-expanded convolution receptive field calculation formula, and the receptive field of the first convolution is calculated using the stacked depth-expanded convolution receptive field calculation formula. The result is the overall receptive field.
[0021] in, Indicates the actual receptive field size of the deep dilated convolution, represents the convolution kernel size, represents the convolution expansion rate, represents the addition operation, Represents a multiplication operation.
[0022] The output feature x of the Stage block is processed by the first regularization layer to obtain the first regularized feature ; The first regularization feature After the first convolution layer, the first convolution feature is obtained ; The first convolution feature Split the features into 4 groups along the channel dimension to get the DLKConv5 convolution layer input features. , DLKConv11 convolutional layer input features , DLKConv17 convolutional layer Input features and DLKConv23 convolutional layer input features , After processing by four parallel decoupled large kernel convolution layers, the processed features are spliced to obtain the spliced features ; Splicing features The features output by the second convolutional layer and the GELU activation function are added to the output features x of the Stage block to obtain the output features of the MDLA multi-scale decoupled large kernel convolution hybrid attention module. ; The formula for the above process is as follows: , , , , , in, represents the operation of the regularization layer, represents a convolution operation with a convolution kernel size of 5×5. Represents a convolution operation with a convolution kernel size of 1×1, Indicates that the features are split into 4 groups evenly along the channel dimension. Indicates the number of channels, Indicates splitting features along the channel dimension, Represents the operation of the DLKConv5 convolutional layer, Represents the operation of the DLKConv11 convolutional layer, Represents the operation of the DLKConv17 convolutional layer, Represents the operation of the DLKConv23 convolutional layer, Concat represents the concatenation operation, and GELU represents the GELU activation function; The AFM attention feedback module includes a second regularization layer, a local attention module, and a global attention module; the local attention module includes a first branch and a second branch; the first branch includes a 1×1 convolution kernel and a 3×3 convolution kernel, and the second branch includes a 1×1 convolution kernel and a Sigmoid activation function; the global attention module includes a third convolution layer, a Softmax function, and a fourth convolution layer; the third convolution layer includes a 1×1 convolution kernel and a 3×3 depth convolution kernel, and the fourth convolution layer is a 1×1 convolution kernel; Output features of MDLA multi-scale decoupled large kernel convolution mixed attention module After processing by the second regularization layer, the second regularization feature is obtained ; The first regularization feature After the first branch, the output of the first branch is obtained, and the second regularization feature The output of the second branch is obtained through the second branch, and the output of the first branch is multiplied by the output of the second branch to obtain the local attention feature. ; The formula is as follows: , , in, represents a convolution operation with a convolution kernel size of 3×3. Represents the Sigmoid activation function; the second regularization feature After the third convolution layer, query Q, key K and value V are generated through 1×1 convolution and 3×3 depth convolution, and query Q, key K and value V are shaped into 、 as well as ,according to 、 and Calculate the attention map , attention map After the fourth convolution layer and the second regularization feature Add up to get the global attention feature ; The formula is as follows: , , , in, Represents a depthwise convolution operation with a convolution kernel size of 3×3; Represent the feature matrices after query Q, key K and value V are shaped, Represents a shaping operation, represents the Softmax activation function, represents the dot product operation, Represents the division operation, Represents the control parameter; local attention features and global attention features The concatenation is performed and processed in sequence through 1×1 convolution kernel, GELU activation function, 1×1 convolution kernel, and then the output features of the MDLA multi-scale decoupled large kernel convolution mixed attention module are obtained. Add up to get the output features of the AFM attention feedback module , the formula is: .
[0023] The Neck module is specifically: The FEAFPN feature-enhanced aligned two-level feature pyramid network includes a top-down primary fusion path and a bottom-up secondary fusion path. The FEAFPN feature-enhanced aligned two-level feature pyramid network designed in the Neck part can better realize the multi-scale fusion of small target information and reduce deviation and feature misalignment.
[0024] In the first fusion path from top to bottom, receive the output features of Backbone , the calculation formula of the fusion path from top to bottom is as follows: in, represents the i-th level feature after fusion from bottom to top, Represents the i-1th level feature, which is the multi-level feature obtained by processing the input image through the Backbone module ={ ,i=2,3,4,5}, It is the feature generated after the output from stage1 to stage4 passes through the above-mentioned MDLA multi-scale decoupling large kernel convolution mixed attention module; FEA represents the FEA feature enhancement alignment module. Represents the i-1th level feature after fusion; the fusion process from top to bottom is: , and so on, we get the top-down fusion feature { ,i=2,3,4,5}.
[0025] Specifically, the FEA feature enhancement alignment module includes the FRE feature refinement enhancement module and the upsampling module; the processing steps are: the fused i-th level feature After being processed by the FRE feature refinement enhancement module, the output features of the FRE feature refinement enhancement module are obtained. ;right Perform upsampling matching The feature size is then learned, and the offset size between the two levels of features is learned. Then, deformable convolution is performed according to the offset size to obtain the aligned up-sampled features. The calculation formula is as follows: in, represents the upsampled features of the i-th level features after the FRE module during fusion, Represents the output of the i-1th layer of the Backbone module, Represents the learned alignment offset, which is a standard convolution with a kernel size of 3×3. Indicates the offset between the features of two layers. represents the deformable convolution operation, Represents the aligned upsampled features, i.e., the output features of the FEA feature enhanced alignment module.
[0026] Specifically, the FRE feature refinement enhancement module is calculated as follows: Input , average pooling is performed along the height and width to generate two 1D sequence features, where B, C, H, and W represent the batch size, channel size, width, and height of the 1D sequence features respectively. In order to learn different spatial distributions and contextual relationships, the two 1D sequence features are divided into four groups of independent sub-features of the same size and represented as and , the number of channels of each sub-feature is , the default value k = 4. The decomposition formula is as follows: in, represents the i-th level feature during fusion, and Respectively represent average pooling along the height and width, represents width 1D sequence features, Represents highly 1D sequence features; Indicates the operation of dividing sub-features. The default value of k is 4, which means that the sub-features are divided into 4 groups. The value of g is ; and They represent the independent sub-features of the width 1D sequence features and the height 1D sequence features respectively; in order to enrich the semantic information, enhance the semantic consistency, and minimize the semantic gap, the four sub-features are parallelized through the deep one-dimensional convolution with kernel sizes of 3, 5, 7 and 9. The dependency between the two dimensions is implicitly modeled by learning the consistent features in the two dimensions. Then, different semantic sub-features are aggregated by Concat and normalized using group normalization GN with K groups. A simple Sigmoid activation function is used to generate the first spatial attention map. The calculated first spatial attention map is multiplied and added with the original feature. Finally, the output of the FRE module is obtained through the GELU activation function. The calculation formula is as follows: in Indicates that the sub-features of the group are subjected to a deep 1D convolution. c is the convolution kernel size and the default value is {3,5,7,9}. g is , and They represent the independent sub-features of the height 1D sequence feature and the width 1D sequence feature after the depth 1D convolution, respectively. Specifically go through get , go through get , and so on, we get 4 sets of features, Concat means splicing features according to dimensions, and Respectively represent the group normalization operation of sequence features in height and width directions, represents the Sigmoid activation function, and Representing the first spatial attention map in height and width directions respectively, Represents the output of FRE.
[0027] In the bottom-up fusion path, the features fused from top to bottom are received { ,i=2,3,4,5}, the top-down fusion path calculation formula is as follows: , in, Represents the i-th level feature after fusion from bottom to top, Represents the i-1th level features after fusion from bottom to top, Represents the APF adaptive perception fusion module; the fusion process from bottom to top is as follows: , , and so on, we get the bottom-up fusion features { ,i=2,3,4,5}.
[0028] Specifically, in order to further enhance the information flow in the feature pyramid network, a bottom-up path is adopted for enhancement, so that the precise positioning signals of the lower layers can be better transmitted upward, and the positioning expression ability of the overall feature pyramid is enhanced, so that the target detection model can obtain accurate detail features at all levels, thereby improving the effect of detecting small target objects. In the APF adaptive perception fusion module, the i-1th level feature after bottom-up fusion is Perform strided convolution with a kernel size of 3x3. The result of strided convolution is the same as Splice by dimension, then pass through the kernel with size 1 The convolution of 1 adjusts the dimension of the spliced feature channel to 2, and performs a Softmax operation to obtain the second spatial attention map , then use The features output by the strided convolution are multiplied by the second spatial attention map. Multiply it by the second spatial attention map, and then add the two features multiplied by the second spatial attention map to obtain the output feature of the APF adaptive perception fusion module. The calculation formula is as follows: , , , in Represents the i-1th level feature after fusion from bottom to top, represents the strided convolution operation, express The output obtained by strided convolution is Indicates taking out the spatial attention map of the first channel, Indicates taking out the spatial attention map of the second channel, It represents the i-th level feature after fusion from bottom to top, that is, the output feature of the APF adaptive perception fusion module.
[0029] Specifically, the CRPN multi-level region proposal network structure in the RPN part includes the first stage and the second stage: In the first stage, the output of the Neck part After the convolution operation of the RPN network, the first regression offset of the anchor frame is obtained ; By first regression offset Calculate the anchor frame position and get the corrected anchor frame , thus obtaining the deviation o between the corrected anchor box and the initial anchor box; correcting the features through adaptive convolution operation Align it with the corrected anchor frame to obtain the features corrected by adaptive convolution ; The formula for the above process is as follows: , , , , in, Represents the output features of the Neck part, represents the convolution operation, Represents the operation of recalculating the anchor box position based on the regression offset, represents the initial anchor box, Represents the adaptive convolution operation; In the second stage, the features corrected by adaptive convolution are convolved to obtain the classification score cls and the second regression offset , using the second regression offset Correct the anchor frame again to get the final output fine anchor frame , the final output fine anchor box Mapping back to the original input image to obtain the final generated image object candidate box ; The formula for the above process is as follows: , , , in, Represents a map operation.
[0030] Specifically, in the Head module, the fine-grained decoupling detection head FGDHead includes a classification branch and a regression branch; the classification branch includes two fully connected layers and one classification fully connected layer; the regression branch includes a convolution layer with a convolution kernel size of 3×3, an MDC multi-scale detail capture module, and a regression fully connected layer; The final generated image object candidate box After the RoiPooling pooling layer extracts the fixed-size feature map, the pooling feature is obtained ; Pooling features The classification output of the detection head is obtained through the classification branch ; Pooling features The bounding box output of the detection head regression is obtained through the regression branch ; The formula is as follows: , , , in, represents the fully connected layer operation, Represents the classification fully connected layer operation, Represents the operation of the MDC multi-scale detail capture module, Represents the operation of the regression fully connected layer.
[0031] Specifically, the implementation process of the MDC multi-scale detail capture module of the regression branch is as follows: The pooling features The intermediate features obtained by the 3×3 convolution layer Divide into four groups along the channel dimension and get the first feature , the second feature , the third characteristic And the fourth characteristic , the formula is as follows: , in, Respectively represent the first feature, the second feature, the third feature and the fourth feature split evenly along the channel dimension; The first processed features are obtained by processing with depthwise separable convolutions of kernel sizes of 3×3, 13×1, and 1×13. , the second processing feature , the third processing feature , Do nothing; 、 、 as well as Perform splicing to obtain the spliced features ; Features after splicing After a normalization layer and MLP layer, the intermediate features Add up to get the output of the MDC multi-scale detail capture module ; The formula for the above process is as follows: , , , , , , in, represents the first processing feature, represents the second processing feature, represents the third processing feature, 、 and Respectively indicate that the kernel size is 3 3, 13 1 and 1 13 depth-wise separable convolution operations, Represents the operation of the linear change perception layer.
[0032] S3. Use the training set to train the small target detection model based on FasterRcnn-S to obtain a trained model; S4. Input the images in the test set into the trained model to obtain the target detection results.
[0033] Example 2 The method of the present invention was trained for 12 epochs during the experiment, using stochastic gradient descent (SGD) as the optimizer, with a learning rate of 0.005, a momentum of 0.9, and a weight decay of 0.0001. The decay was performed at the 8th and 11th epochs, and the mean average precision (mAP) was used as the evaluation metric. The size targets in the range of [2, 20] pixels were further divided into three sub-intervals: tiny1 [2, 8] size, tiny2 [8, 12] size, and tiny3 [12, 20] size. Evaluation Metrics , , , , They represent the average precision of [2, 20] pixel size targets, [2, 8] size targets, [8, 12] size targets, [12, 20] size targets, and [2, 32] size targets when the IOU threshold is greater than 0.5. and They represent the average precision of objects with a size of [2, 20] pixels when the IOU threshold is greater than 0.25 and 0.75, respectively. As can be seen from Table 1, based on the FasterRcnn model, whether it is data enhancement or optimization and improvement of each module of the FasterRcnn model, the average precision of the original FasterRcnn model can be improved to a certain extent. The method of the present invention adopts the RseNet-50 neural network structure in the BackBone module, adds the MDLA multi-scale decoupled large kernel convolution hybrid attention module after each stage from Stage 1 to Stage 4, adopts the FEAFPN feature enhancement and alignment two-level feature pyramid network for feature fusion in the Neck module, introduces the CRPN multi-level region proposal network structure in the RPN module, and designs a fine-grained decoupled detection head FGDHead in the Head module to replace the ordinary target detection head. Finally, the target detection model is obtained. The average precision of this model is higher than that of the single module optimized separately, which proves the effectiveness of the method of the present invention.
[0034] Table 1 Comparison of ablation experimental results of the method of the present invention like Figure 2 As shown, Figure 2 (a) is the effect of applying the FasterRcnn model for detection. Figure 2 (b) is the detection effect of FasterRcnn-S of the present invention. It can be seen from the figure that the detection effect of the method of the present invention is significantly better than that of the FasterRcnn model in the upper left and middle dense areas of the detection image. The FasterRcnn model has a large number of missed detection phenomena in the detection image. The effectiveness of the present invention for small target detection is proved by comparing the detection results.
[0035] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A small target detection method based on FasterRcnn-S, characterized in that: The following steps are involved: S1. Obtain the TinyPerson small object detection dataset, perform data augmentation on the dataset, and divide it into a training set and a test set; S2. Construct a small target detection model based on FasterRcnn-S. The model uses the FasterRcnn network as the basic network, including the BackBone module, the Neck module, the RPN module, and the Head module. The BackBone module adopts the RseNet-50 neural network structure, and adds the MDLA multi-scale decoupled large kernel convolution hybrid attention module after each stage from Stage 1 to Stage 4. The Neck module uses the FEAFPN feature enhancement and alignment two-level feature pyramid network for feature fusion. The RPN module introduces the CRPN multi-level region proposal network structure. The Head module designs a fine-grained decoupled detection head FGDHead to replace the ordinary target detection head. S3. Use the training set to train the small target detection model based on FasterRcnn-S to obtain a trained model; S4. Input the images in the test set into the trained model to obtain the target detection results.
2. The small target detection method based on FasterRcnn-S according to claim 1, characterized in that: The data enhancement in step S1 includes cropping, stacking, scaling, and transformation.
3. The small target detection method based on FasterRcnn-S according to claim 2, characterized in that: In step S2, the Backbone module is specifically as follows: Add an MDLA multi-scale decoupled large kernel convolution hybrid attention module after each Stage block of the Resnet-50 residual structure in the Backbone module; the MDLA multi-scale decoupled large kernel convolution hybrid attention module includes the MDLConvs multi-scale decoupled large kernel convolution module and the AFM attention feedback module; The MDLConvs multi-scale decoupled large kernel convolution module includes the first regularization layer, the first convolution layer, four parallel decoupled large kernel convolution layers, the second convolution layer and the GELU activation function in sequence; the first convolution layer includes an 11 convolution kernel and a 5×5 convolution kernel; the four parallel decoupled large kernel convolution layers are DLKConv5 convolution layer, DLKConv11 convolution layer, DLKConv17 convolution layer and DLKConv23 convolution layer from top to bottom ... The second convolutional layer is a 1×1 convolution kernel; the DLKConv5 convolution layer stacks two (3,1) depth dilated convolutions, the DLKConv11 convolution layer stacks (3,1) and (5,2) depth dilated convolutions in sequence, the DLKConv17 convolution layer stacks (5,1) and (7,2) depth dilated convolutions in sequence, and the DLKConv23 convolution layer stacks (5,1) and (7,3) depth dilated convolutions in sequence. The first parameter in the brackets indicates the size of the convolution kernel. , the second parameter represents the expansion rate of the convolution ; The calculation formula for the receptive field of stacked depth dilated convolution is: , in, represents the bottom receptive field of the stack, Represents the receptive field of the previous layer of the stack, represents the stride of convolution, represents the size of the convolution kernel, represents the subtraction operation, represents the multiplication operation, Represents addition operation; The output feature x of the Stage block is processed by the first regularization layer to obtain the first regularized feature ; The first regularization feature After the first convolution layer, the first convolution feature is obtained ; The first convolution feature Split the features into 4 groups along the channel dimension to get the DLKConv5 convolution layer input features. , DLKConv11 convolutional layer input features , DLKConv17 convolutional layer Input features and DLKConv23 convolutional layer input features , After processing by four parallel decoupled large kernel convolution layers, the processed features are spliced to obtain the spliced features ; Splicing features The features output by the second convolutional layer and the GELU activation function are added to the output features x of the Stage block to obtain the output features of the MDLA multi-scale decoupled large kernel convolution hybrid attention module. ; The formula for the above process is as follows: , , , , , in, represents the operation of the regularization layer, represents a convolution operation with a convolution kernel size of 5×5. Represents a convolution operation with a convolution kernel size of 1×1, Indicates that the features are split into 4 groups evenly along the channel dimension. Indicates the number of channels, Indicates splitting features along the channel dimension, Represents the operation of the DLKConv5 convolutional layer, Represents the operation of the DLKConv11 convolutional layer, Represents the operation of the DLKConv17 convolutional layer, Represents the operation of the DLKConv23 convolutional layer, Concat represents the concatenation operation, and GELU represents the GELU activation function; The AFM attention feedback module includes a second regularization layer, a local attention module, and a global attention module; the local attention module includes a first branch and a second branch; the first branch includes a 1×1 convolution kernel and a 3×3 convolution kernel, and the second branch includes a 1×1 convolution kernel and a Sigmoid activation function; the global attention module includes a third convolution layer, a Softmax function, and a fourth convolution layer; the third convolution layer includes a 1×1 convolution kernel and a 3×3 depth convolution kernel, and the fourth convolution layer is a 1×1 convolution kernel; Output features of MDLA multi-scale decoupled large kernel convolution hybrid attention module After processing by the second regularization layer, the second regularization feature is obtained ; The first regularization feature After the first branch, the output of the first branch is obtained, and the second regularization feature The output of the second branch is obtained through the second branch, and the output of the first branch is multiplied by the output of the second branch to obtain the local attention feature. ; The second regularization feature After the third convolution layer, query Q, key K and value V are generated through 1×1 convolution and 3×3 depth convolution, and query Q, key K and value V are shaped into 、 as well as ,according to 、 and Calculate the attention map , attention map After the fourth convolution layer and the second regularization feature Add up to get the global attention feature ; The formula is as follows: , , , in, Represents a depthwise convolution operation with a convolution kernel size of 3×3; Represent the feature matrices after query Q, key K and value V are shaped, Represents a shaping operation, represents the Softmax activation function, represents the dot product operation, Represents the division operation, Represents the control parameter; local attention features and global attention features The concatenation is performed and processed in sequence through 1×1 convolution kernel, GELU activation function, 1×1 convolution kernel, and then the output features of the MDLA multi-scale decoupled large kernel convolution mixed attention module are obtained. Add up to get the output features of the AFM attention feedback module , the formula is: .
4. The small target detection method based on FasterRcnn-S according to claim 3, characterized in that The Neck module in step S2 is specifically: The FEAFPN feature-enhanced aligned two-level feature pyramid network includes a top-down primary fusion path and a bottom-up secondary fusion path; In the first fusion path from top to bottom, receive the output features of Backbone , the calculation formula of the fusion path from top to bottom is as follows: in, represents the i-th level feature after fusion from bottom to top, Represents the i-1th level feature, which is the multi-level feature obtained by processing the input image through the Backbone module ={ ,i=2,3,4,5}, It is the feature generated after the output from stage1 to stage4 passes through the above-mentioned MDLA multi-scale decoupling large kernel convolution mixed attention module; FEA represents the FEA feature enhancement alignment module. Represents the i-1th level feature after fusion; the fusion process from top to bottom is: , and so on, we get the top-down fusion feature { ,i=2,3,4,5}; In the bottom-up fusion path, the features fused from top to bottom are received { ,i=2,3,4,5}, the top-down fusion path calculation formula is as follows: , in, Represents the i-th level feature after fusion from bottom to top, Represents the i-1th level features after fusion from bottom to top, Represents the APF adaptive perception fusion module; the fusion process from bottom to top is as follows: , , and so on, we get the bottom-up fusion features { ,i=2,3,4,5}.
5. The small target detection method based on FasterRcnn-S according to claim 4, characterized in that: The FEA feature enhancement and alignment module includes the FRE feature refinement and enhancement module and the upsampling module. The specific operations are as follows: The fused i-th level features After being processed by the FRE feature refinement enhancement module, the output features of the FRE feature refinement enhancement module are obtained. ;right Perform upsampling matching The feature size is then learned, and the offset size between the two levels of features is learned. Then, deformable convolution is performed according to the offset size to obtain the aligned up-sampled features. The calculation formula is as follows: in, represents the upsampled features of the i-th level features after the FRE module during fusion, Represents the output of the i-1th layer of the Backbone module, Represents the learned alignment offset, which is a standard convolution with a kernel size of 3×3. Indicates the offset between the features of two layers. represents the deformable convolution operation, Represents the aligned upsampled features, i.e., the output features of the FEA feature enhanced alignment module.
6. The small target detection method based on FasterRcnn-S according to claim 5, characterized in that: The specific operations of the APF adaptive perception fusion module are as follows: The i-1th level features after bottom-up fusion The kernel size is 3 The strided convolution operation of 3, the result of the strided convolution is the same as Splicing is performed by dimension, and the spliced features are processed with a kernel size of 1. The convolution of 1 adjusts the channel dimension to 2, and performs a Softmax operation to obtain the second spatial attention map. , then use The features output by the strided convolution are multiplied by the second spatial attention map. Multiply it by the second spatial attention map, and then add the two features multiplied by the second spatial attention map to obtain the output feature of the APF adaptive perception fusion module. The calculation formula is as follows: , , , in Represents the i-1th level feature after fusion from bottom to top, represents the strided convolution operation, express The output obtained by strided convolution is Indicates taking out the spatial attention map of the first channel, Indicates taking out the spatial attention map of the second channel, It represents the i-th level feature after fusion from bottom to top, that is, the output feature of the APF adaptive perception fusion module.
7. The small target detection method based on FasterRcnn-S according to claim 6, characterized in that: The specific process of the first stage of the CRPN multi-level region proposal network structure in the RPN part in step S2 is: the output of the Neck part After the convolution operation of the RPN network, the first regression offset of the anchor frame is obtained ; By first regression offset Calculate the anchor frame position and get the corrected anchor frame , thus obtaining the deviation o between the corrected anchor box and the initial anchor box; correcting the features through adaptive convolution operation Align it with the corrected anchor frame to obtain the features corrected by adaptive convolution .
8. The small target detection method based on FasterRcnn-S according to claim 7, characterized in that: The specific process of the second stage of the CRPN multi-level region proposal network structure in the RPN part of step S2 is as follows: the features corrected by adaptive convolution are subjected to convolution operation to obtain the classification score cls and the second regression offset , using the second regression offset Correct the anchor frame again to get the final output fine anchor frame , the final output fine anchor box Mapping back to the original input image to obtain the final generated image object candidate box .
9. The small target detection method based on FasterRcnn-S according to claim 8, characterized in that: The Head module in step S2 specifically includes: The fine-grained decoupling detection head FGDHead includes a classification branch and a regression branch; the classification branch includes two fully connected layers and one classification fully connected layer; the regression branch includes a convolution layer with a convolution kernel size of 3×3, an MDC multi-scale detail capture module and a regression fully connected layer; The final generated image object candidate box After the RoiPooling pooling layer extracts the fixed-size feature map, the pooling feature is obtained ; Pooling features The classification output of the detection head is obtained through the classification branch ; Pooling features The bounding box output of the detection head regression is obtained through the regression branch .
10. The small target detection method based on FasterRcnn-S according to claim 9, characterized in that: The MDC multi-scale detail capture module of the regression branch in the Head module in step S2 is specifically implemented as follows: The pooling features The intermediate features obtained by the 3×3 convolution layer Divide into four groups along the channel dimension and get the first feature , the second feature , the third characteristic And the fourth characteristic ; The first processed features are obtained by processing with depthwise separable convolutions of kernel sizes of 3×3, 13×1, and 1×13. , the second processing feature , the third processing feature , Do nothing; 、 、 as well as Perform splicing to obtain the spliced features ; Features after splicing After a normalization layer and MLP layer, the intermediate features Add up to get the output of the MDC multi-scale detail capture module .
Citation Information
Cited By
A mixed gas infrared spectrum analysis model establishing and detecting method
CN122451705A