An Image Small Target Detection Method and System Based on Gaussian Allocation Strategy

Through the image small object detection method based on the Gaussian allocation strategy, the problems of insufficient positive sample allocation and size difference in small-size object detection are solved, and the performance of small-size object detection is significantly improved and accelerated.

CN114782709BActive Publication Date: 2025-07-22HUNAN SHENFAN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210559371.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-07-22
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

In the existing small-size object detection methods, the label allocation strategy leads to too few positive samples allocated by small-size objects, and the number of positive samples allocated by different size objects is too large, resulting in low detection accuracy and inability to meet the usage requirements.

Method used

The image small object detection method based on the Gaussian allocation strategy is adopted. By loading image features, a multi-scale feature map set is formed, feature fusion and Gaussian allocation strategy are used to allocate positive and negative samples, and the loss is calculated and the network parameters are updated until the network converges. The trained network is used for object detection and non-maximum suppression processing is performed.

Benefits of technology

Improved performance of small target detectors, doubled on average, and even tripled for very tiny targets, with performance up to 3.4% higher than that of state-of-the-art competitors, while achieving 7x acceleration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114782709B_ABST
    Figure CN114782709B_ABST
Patent Text Reader

Abstract

The present invention discloses an image small target detection method and system based on a Gaussian assignment strategy. By loading an image and extracting corresponding features of the image, a multi-scale feature map set is constructed; the extracted low-level feature map and high-level feature map are feature-fused to generate a fused feature map; the category of the target, the coordinate values of the target prediction box, and the centrality of the target box are predicted according to the generated fused feature map; the Gaussian assignment strategy is used to assign positive and negative samples of the image; the network loss is calculated by combining the assigned positive and negative samples and the target category, prediction box coordinates, and centrality output by the network; the network parameters are updated according to the calculated network loss until the network converges; using the trained network, the category of the target, the coordinate values of the target prediction box, and the centrality of the target box in the input image are predicted, and then the detection result is obtained by using a post-processing algorithm. The present invention doubles the performance of the small target detector on average, and for very tiny targets, it can even be tripled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent image processing, and in particular discloses an image small target detection method and system based on a Gaussian distribution strategy. Background Technique

[0002] At present, small-size target detection, as a research difficulty and hotspot in computer vision, is widely applied to driving assistance, remote sensing, and large-scale video surveillance. Driven by the novel network structures and label assignment strategies proposed recently, great progress has been made, but it is still a very challenging problem. Considering the current remote sensing image small target detection scenario, the target images on average less than 13 pixels, and there are even weak small targets with 4 pixels. Existing target detection algorithms are difficult to accurately detect such targets.

[0003] Currently, there are mainly three ideas to solve the problem of small target detection. One is to obtain size-invariant features by normalizing the input image size and multi-scale feature learning (Scale match for tiny person detection: Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision. (2020) 1257{1265). The second is to enhance the feature representation ability of small targets through GAN networks or context information (Perceptual generative adversarial networks for small object detection.: Proceedings of the IEEE conference on computer vision and pattern recognition. (2017) 1222{1230). The third is to improve the performance of small target detection by measuring the similarity of two bounding boxes through specific algorithms (A normalized gaussian wasserstein distance for tiny object detection. arXiv preprint arXiv:2110.13389 (2021)).

[0004] Although these methods have improved the performance of small target detection to a certain extent, the performance improvement is limited, and at the same time, the impact of the label assignment strategy on the performance of small target detection is ignored.

[0005] The current label assignment strategy assigns too few positive samples to small-sized targets, and there is a large difference in the number of positive samples assigned to targets of different sizes, which will result in low detection accuracy for small targets and cannot meet the usage requirements.

[0006] Therefore, the above-mentioned defects existing in the existing small target detection methods are technical problems that need to be solved urgently at present. Summary of the Invention

[0007] The present invention provides an image small target detection method and system based on a Gaussian assignment strategy, aiming to solve the above-mentioned defects existing in the existing small target detection methods.

[0008] One aspect of the present invention relates to an image small target detection method based on a Gaussian assignment strategy, including the following steps:

[0009] Load an image and extract corresponding features of the image to form a multi-scale feature map set;

[0010] Perform feature fusion on the extracted low-level feature map and high-level feature map to generate a fused feature map;

[0011] Perform detection output according to the generated fused feature map, and predict the category of the target, the coordinate values of the target prediction box, and the centrality of the target box;

[0012] Use the Gaussian assignment strategy to assign positive and negative samples of the image;

[0013] Calculate the network loss by combining the assigned positive and negative samples and the target category, prediction box coordinates, and centrality output by the network;

[0014] Update the network parameters according to the calculated network loss, and repeat the above training steps, continuously updating the network parameters until the network converges;

[0015] Use the trained network to predict the category of the target, the coordinate values of the target prediction box, and the centrality of the target box in the input image;

[0016] Perform non-maximum suppression processing on the detection results to remove duplicate predicted target categories and prediction box coordinate values, and obtain the best detection results.

[0017] Further, after the step of loading and extracting corresponding features of the image to form a multi-scale feature map set, it further includes:

[0018] Map the target annotation box to a two-dimensional heat map , where N represents the batch size, C represents the number of categories, H represents the height, W represents the width, and S represents the downsampling rate;

[0019] Encode the training samples using the heatmap generated by the annotation, and use a preset α threshold to control the proportion of positive samples. Samples in the heatmap greater than the α threshold are positive samples, and samples less than or equal to the α threshold in the heatmap are negative samples;

[0020] Apply the encoded training samples to the regression weights.

[0021] Furthermore, map the target annotation box to a two-dimensional heatmap in the step of

[0022] For an annotation box , belonging to the th class, linearly map it to the scale of the feature map;

[0023] Adopt a two-dimensional Gaussian distribution to generate the heatmap ;

[0024] The probability density function of the two-dimensional Gaussian distribution is given by the following formula:

[0025]

[0026] where represents the coordinates of the Gaussian distribution , represents the mean vector of the Gaussian distribution, represents the covariance matrix of the Gaussian distribution;

[0027]

[0028]

[0029] where is used to limit the maximum values of and to balance the number of positive samples assigned to objects of different sizes;

[0030] Update the channel in by performing an element-wise maximum operation on ;

[0031] For a rotated annotation box , adopt a rotated two-dimensional Gaussian distribution to generate the heatmap of the target,

[0032]

[0033] where is the rotation width, represents the rotation height, Indicates the rotation angle.

[0034] Further, in the step of encoding the training samples with the heatmap generated by the annotation and using a preset α threshold to control the proportion of positive samples, where samples in the heatmap greater than the α threshold are positive samples and samples less than or equal to the α threshold are negative samples,

[0035] The centrality of the target of the sample is defined as:

[0036]

[0037] Where, respectively represent the distances from the target center point to the left, right, top, and bottom sides of the target box.

[0038] Further, in the step of applying the encoded training samples to the regression weights,

[0039] Assume At the sub-region of the th annotation box, is the regression sample weight,

[0040]

[0041] Where, is the Gaussian probability at and is the th box area.

[0042] Another aspect of the present invention relates to an image small target detection system based on a Gaussian allocation strategy, including:

[0043] A feature extraction module for loading and extracting corresponding features of the image to form a multi-scale feature map set;

[0044] A feature fusion module for fusing the extracted low-level feature map and high-level feature map to generate a fused feature map;

[0045] A detection module for performing detection output according to the generated fused feature map, predicting the category of the target, the coordinate values of the target prediction box, and the centrality of the target box;

[0046] An allocation module for allocating positive and negative samples of the image using the Gaussian allocation strategy;

[0047] A calculation module for calculating the network loss by combining the allocated positive and negative samples with the target category, prediction box coordinates, and centrality output by the network;

[0048] An update module, configured to update network parameters according to the calculated network loss, and repeat the above training steps to continuously update the network parameters until the network converges;

[0049] A prediction module, configured to use the trained network to predict the category of the target in the input image, the coordinate values of the target prediction box, and the centrality of the target box;

[0050] A post-processing module, configured to perform non-maximum suppression processing on the detection results, remove the target categories and prediction box coordinate values with repeated predictions, and obtain the optimal detection results.

[0051] Furthermore, the small target detection system for images based on the Gaussian assignment strategy further includes:

[0052] A heat map mapping unit, configured to map the target annotation box into a two-dimensional heat map , where N represents the batch size, C represents the number of categories, H represents the height, W represents the width, and S represents the downsampling rate;

[0053] A training sample encoding unit, configured to encode the training samples by using the heat map generated by the annotation, and use a preset α threshold to control the proportion of positive samples. Samples in the heat map greater than the α threshold are positive samples, and samples less than or equal to the α threshold are negative samples;

[0054] A regression weight unit, configured to apply the encoded training samples to the regression weights.

[0055] Furthermore, in the heat map mapping unit, for an annotation box , belonging to the th class, it will be linearly mapped to the scale of the feature map;

[0056] Adopt a two-dimensional Gaussian distribution to generate a heat map ;

[0057] The probability density function of the two-dimensional Gaussian distribution is given by the following formula:

[0058]

[0059] where represents the coordinates of the Gaussian distribution , represents the mean vector of the Gaussian distribution, represents the covariance matrix of the Gaussian distribution;

[0060]

[0061]

[0062] where used to limit and the maximum value of, to balance the number of positive samples assigned to objects of different sizes;

[0063] By performing element-wise maximum processing on to update the channel in;

[0064] For the rotated annotation box , a rotated two-dimensional Gaussian distribution is used to generate the heatmap of the target,

[0065]

[0066] where, is the rotation width, represents the rotation height, represents the rotation angle.

[0067] Furthermore, in the training sample encoding unit, the centrality of the target of the sample is defined as:

[0068]

[0069] where, respectively represent the distances from the center point of the target to the left, right, top, and bottom sides of the target box.

[0070] Furthermore, assume in the sub-region of the th annotation box, is the regression sample weight,

[0071]

[0072] where, is the Gaussian probability at , is the th box area.

[0073] The beneficial effects achieved by the present invention are:

[0074] The present invention provides an image small target detection method and system based on a Gaussian assignment strategy. By loading an image and extracting corresponding features of the image, a multi-scale feature map set is constituted; the extracted low-level feature map and high-level feature map are subjected to feature fusion to generate a fused feature map; detection output is performed according to the generated fused feature map to predict the category of the target, the coordinate values of the target prediction box, and the centrality of the target box; the Gaussian assignment strategy is used to assign positive and negative samples of the image; the network loss is calculated by combining the assigned positive and negative samples with the target category, prediction box coordinates, and centrality output by the network; the network parameters are updated according to the calculated network loss, and the above training steps are repeated, continuously updating the network parameters until the network converges; the trained network is used to predict the category of the target, the coordinate values of the target prediction box, and the centrality of the target box in the input image; non-maximum suppression processing is performed on the detection results to remove duplicate predicted target categories and prediction box coordinate values, and the best detection results are obtained. The image small target detection method and system based on the Gaussian assignment strategy provided by the present invention. The proposed Gaussian assignment strategy can increase the number of positive samples assigned to small targets and balance the number of positive samples assigned to objects of different sizes; the proposed Gaussian assignment strategy can use the normalized Gaussian probability as the sample weight to re-weight the contribution of positive samples during the calculation of the loss and balance the influence of objects of different sizes; the provided algorithm can double the performance of the small target detector on average, and even triple it for very tiny targets; the performance is 3.4% higher than that of the most advanced competitors (from 20.8% to 24.2%), and at the same time, a 7-fold acceleration is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 It is a schematic flowchart of an embodiment of the image small target detection method based on the Gaussian assignment strategy provided by the present invention;

[0076] Figure 2 It is a network structure diagram in the image small target detection method based on the Gaussian assignment strategy provided by the present invention;

[0077] Figure 3 It is a schematic diagram of mapping the annotation box into a two-dimensional Gaussian distribution in the image small target detection method based on the Gaussian assignment strategy provided by the present invention;

[0078] Figure 4 It is a schematic diagram of using a Gaussian heat map to encode training samples in the image small target detection method based on the Gaussian assignment strategy provided by the present invention;

[0079] Figure 5 It is a schematic diagram of the label assignment situation of different label assignment strategies for objects of different sizes in the image small target detection method based on the Gaussian assignment strategy provided by the present invention;

[0080] Figure 6Flowchart of the training phase in the small object detection method for images based on the Gaussian allocation strategy provided by the present invention;

[0081] Figure 7 Flowchart of the inference phase in the small object detection method for images based on the Gaussian allocation strategy provided by the present invention;

[0082] Figure 8 Detection effect diagram in the small object detection method for images based on the Gaussian allocation strategy provided by the present invention;

[0083] Figure 9 Functional block diagram of the first embodiment of the small object detection system for images based on the Gaussian allocation strategy provided by the present invention;

[0084] Figure 10 Functional block diagram of the second embodiment of the small object detection system for images based on the Gaussian allocation strategy provided by the present invention.

[0085] Explanation of the reference numerals in the attached drawings:

[0086] 10. Feature extraction module; 20. Feature fusion module; 30. Detection module; 40. Allocation module; 50. Calculation module; 60. Update module; 70. Prediction module; 80. Post-processing module; 11. Heat map mapping unit; 12. Training sample encoding unit; 13. Regression weight unit. Detailed implementation manners

[0087] To better understand the above technical solutions, the above technical solutions will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0088] As Figures 1 to 8 shown, the first embodiment of the present invention proposes a small object detection method for images based on the Gaussian allocation strategy, including the following steps:

[0089] Step S100: Load and extract the corresponding features of the image to form a multi-scale feature map set.

[0090] Receive a sample image, extract the features of the sample image, and form a multi-scale feature map set; due to the small size of the target, the output of the feature map with 4 times downsampling is increased. Feature extraction can use general backbone networks such as Darknet53, Resnet50, and DLA34.

[0091] Step S200: Perform feature fusion on the extracted low-level feature map and high-level feature map to generate a fused feature map.

[0092] Fuse the extracted low-level feature maps and high-level feature maps to generate a fused feature map with a downsampling rate of 4 times and 128 channels. Add a shortcut connection in the feature fusion module to introduce high-resolution features. The shortcut connection introduces the features of the second, third, and fourth stages of the backbone network, and each connection is implemented by a 3×3 convolutional layer. The numbers of layers in the second, third, and fourth stages are 3, 2, and 1 layer respectively.

[0093] Step S300: Perform detection output based on the generated fused feature map to predict the category of the target, the coordinate values of the target prediction box, and the centrality of the target box.

[0094] Receive the fused feature map generated by the feature fusion module and predict the category of the target, the coordinate values of the target prediction box, and the centrality of the target box. All branches and sub-modules of the detection module are connected to the feature fusion module and are all convolutional networks. The detection module includes a classification branch and a localization branch.

[0095] 1) The classification branch is used to perform convolutional processing on the fused feature map to obtain the category prediction score. The classification branch is connected to the feature fusion module. First, obtain classification features through two 128-channel 3×3 convolutional layers, and then obtain the category confidence through a 3×3 convolutional layer with the number of channels of the target category.

[0096] 2) The localization branch is used to perform convolutional processing on the fused feature map to obtain the coordinates and centrality of the prediction box. The localization branch is connected to the feature fusion module and obtains regression features through two 64-channel 3×3 convolutional layers.

[0097] The localization branch is divided into two modules, one is the target box regression module and the other is the centrality regression module. Among them, the target box regression module uses a 1-layer 4-channel 3×3 convolution to obtain the coordinate values of the target prediction box, and the centrality regression module is used to obtain the centrality of the target through a 1-layer single-channel 3×3 convolutional layer.

[0098] Step S400: Use the Gaussian assignment strategy to assign positive and negative samples of the image.

[0099] Step S500: Calculate the network loss by combining the assigned positive and negative samples and the target category, prediction box coordinates, and centrality output by the network.

[0100] The loss function consists of the following parts.

[0101] (1)

[0102] In formula (1), is the classification loss, and Focalloss is adopted here. 、 are loss coefficients used to adjust the contributions of different losses. Here, take , 。 For the regression loss, IoU loss is adopted; That is is , For the centrality loss, L1 loss is adopted.

[0103] Step S600: Update the network parameters according to the calculated network loss, repeat the above training steps, and continuously update the network parameters until the network converges.

[0104] Calculate the network loss, perform gradient backpropagation, update the network parameters, and continuously iterate the training until the network converges.

[0105] Step S700: Use the trained network to predict the category of the target in the input image, the coordinate values of the target prediction box, and the centrality of the target box.

[0106] Step S800: Perform non-maximum suppression processing on the detection results, remove the target categories and prediction box coordinate values with repeated predictions, and obtain the final detection results.

[0107] Obtain the detection results using the post-processing algorithm. Figure 6 is the detection effect diagram. The upper part of the figure is the detection result of the small target detection algorithm without using the Gaussian distribution strategy, and the lower part is the detection effect of the small target detection using the Gaussian distribution strategy. It can be seen from the figure that the algorithm can significantly improve the detection accuracy of small and weak targets.

[0108] The small object detection method for images based on the Gaussian assignment strategy provided in this embodiment forms a multi-scale feature map set by loading an image and extracting corresponding features of the image; fuses the extracted low-level feature map and high-level feature map to generate a fused feature map; performs detection output based on the generated fused feature map to predict the category of the target, the coordinate values of the target prediction box, and the centrality of the target box; uses the Gaussian assignment strategy to assign positive and negative samples of the image; calculates the network loss by combining the assigned positive and negative samples with the target category, prediction box coordinates, and centrality output by the network; updates the network parameters according to the calculated network loss, repeats the above training steps, and continuously updates the network parameters until the network converges; uses the trained network to predict the category of the target, the coordinate values of the target prediction box, and the centrality of the target box in the input image; performs non-maximum suppression processing on the detection results to remove duplicate predicted target categories and prediction box coordinate values, and obtains the best detection result. The small object detection method for images based on the Gaussian assignment strategy provided in this embodiment. The proposed Gaussian assignment strategy can increase the number of positive samples assigned to small objects and balance the number of positive samples assigned to objects of different sizes; the proposed Gaussian assignment strategy can use the normalized Gaussian probability as the sample weight to re-weight the contribution of positive samples during the calculation of the loss and balance the influence of objects of different sizes; the provided algorithm can on average double the performance of the small object detector, and for very tiny objects, it can even triple the performance; the performance is 3.4% higher than that of the most advanced competitors (from 20.8% to 24.2%), and at the same time, a 7-fold acceleration is achieved.

[0109] After step S100 of the small object detection method for images based on the Gaussian assignment strategy provided in this embodiment, the following steps are further included:

[0110] Step S110: Map the target annotation box to a two-dimensional heat map , where N represents the batch size, C represents the number of categories, H represents the height, W represents the width, and S represents the downsampling rate.

[0111] First, for an annotation box , belonging to the th class, it is linearly mapped to the scale of the feature map;

[0112] Then, a two-dimensional Gaussian distribution is used to generate the heat map ;

[0113] The probability density function of the two-dimensional Gaussian distribution is given by the following formula:

[0114] (2)

[0115] In formula (2), represents the coordinates of the Gaussian distribution , represents the mean vector of the Gaussian distribution, and

[0116] represents the covariance matrix of the Gaussian distribution.

[0116] (3)

[0117] (4)

[0118] In formulas (3) and (4), is used to limit and 's maximum value to balance the number of positive samples assigned to objects of different sizes.

[0119] Finally, by performing element-wise maximum processing on to update the in the channel.

[0120] For the rotated bounding box , a rotated two-dimensional Gaussian distribution is used to generate the heatmap of the target,

[0121] (5)

[0122] In formula (5), is the rotation width, represents the rotation height, represents the rotation angle.

[0123] Figure 3 Figure shows the schematic diagram of mapping the bounding box to a two-dimensional Gaussian distribution, visually demonstrating how to map the bounding box to a two-dimensional Gaussian distribution.

[0124] Step S120: Encode the training samples using the heatmap generated from the annotation, and use a preset α threshold to control the proportion of positive samples. Samples in the heatmap greater than the α threshold are positive samples, and samples less than or equal to the α threshold are negative samples. All samples greater than the α threshold in all heatmaps are positive samples, and samples in other positions are negative samples.

[0125] The centrality of the target of the sample is defined as:

[0126] (6)

[0127] In formula (6), respectively represent the distances from the center point of the target to the left, right, top, and bottom sides of the target box.

[0128] Figure 4It is a schematic diagram of using Gaussian heatmap encoding for training samples. In the figure, the positions where the Gaussian distribution is above the α-threshold plane are the positive sample positions, and other positions are the negative sample positions.

[0129] Step S130: Apply the encoded training samples to the regression weights.

[0130] Suppose At the sub-region of the th annotation box, is the regression sample weight,

[0131] (7)

[0132] In formula (7), is the Gaussian probability at The Gaussian probability at this point, is the area of the th box.

[0133] Applying the Gaussian distribution to the regression weights can make full use of more annotation information contained in large objects and retain the annotation information of small objects; it can also emphasize these samples near the object center and reduce the influence of low-quality samples.

[0134] Since most small targets to be detected are not strict quadrilaterals, there are often a lot of backgrounds near the bounding boxes of small targets. In these objects, foreground pixels are concentrated in the center of the bounding box. The label assignment strategy based on the Gaussian distribution can increase the number of positive samples assigned to small targets and ensure that at least one positive sample is assigned to each target. For normal-sized objects, the strategy used will select high-quality positive samples near the center and reduce the influence of low-quality samples. At the same time, the strategy for small targets effectively reduces the difference in the number of positive samples assigned to objects of different sizes by restricting the maximum radius of the object.

[0135] Figure 5 It is a schematic diagram of different label assignment strategies for label assignment of objects of different sizes. It can be seen from the figure that the assignment strategy based on the Gaussian distribution can increase the number of positive samples assigned to small targets and balance the number of positive samples assigned to objects of different sizes.

[0136] The image small target detection method based on the Gaussian allocation strategy provided in this embodiment maps the target annotation box into a two-dimensional heat map; encodes the training samples with the heat map generated by the annotation, and uses a preset α threshold to control the proportion of positive samples. Samples greater than the α threshold in the heat map are positive samples, and samples less than or equal to the α threshold in the heat map are negative samples; applies the encoded training samples to the regression weights. The image small target detection method based on the Gaussian allocation strategy provided in this embodiment proposes a Gaussian allocation strategy that can increase the number of positive samples allocated to small targets and balance the number of positive samples allocated to objects of different sizes; the proposed Gaussian allocation strategy can use the normalized Gaussian probability as the sample weight to re-weight the contribution of positive samples during the calculation of the loss and balance the influence of objects of different sizes; the provided algorithm can on average double the performance of the small target detector, and for very tiny targets, it can even triple; the performance is 3.4% higher than that of the state-of-the-art competitors (from 20.8% to 24.2%), and at the same time, a 7-fold acceleration is achieved.

[0137] Preferably, see Figure 9 and Figure 10 This invention relates to an image small target detection system based on the Gaussian allocation strategy, including a feature extraction module 10, a feature fusion module 20, a detection module 30, an acquisition module 40, an allocation module 50, a calculation module 60, an update module 70, and a post-processing module 80. Among them, the feature extraction module 10 is used to load and extract the corresponding features of the image to form a multi-scale feature map set; the feature fusion module 20 is used to perform feature fusion on the extracted low-level feature map and high-level feature map to generate a fused feature map; the detection module 30 is used to perform detection output according to the generated fused feature map, and predict the category of the target, the coordinate values of the target prediction box, and the centrality of the target box; the allocation module 40 is used to allocate positive and negative samples of the image using the Gaussian allocation strategy; the calculation module 50 is used to calculate the network loss by combining the allocated positive and negative samples with the target category, prediction box coordinates, and centrality output by the network; the update module 60 is used to update the network parameters according to the calculated network loss, repeat the above training steps, and continuously update the network parameters until the network converges; the prediction module 70 is used to use the trained network to predict the category of the target, the coordinate values of the target prediction box, and the centrality of the target box in the input image; the post-processing module 80 is used to perform non-maximum suppression processing on the detection results, remove the target categories and prediction box coordinate values with repeated predictions, and obtain the final detection results. It should be noted that the post-processing module is not required during the training phase, and only during the inference phase, that is, after the model is trained, the post-processing module is required during the inference phase.

[0138] The feature extraction module 10 receives sample images, extracts the features of the sample images, and constructs a set of multi-scale feature maps; due to the small target size, the output of the feature map with 4 times downsampling is added. For feature extraction, general backbone networks such as Darknet53, Resnet50, and DLA34 can be used.

[0139] The feature fusion module 20 fuses the extracted low-level feature maps and high-level feature maps to generate a fused feature map with a downsampling rate of 4 times and 128 channels. In the feature fusion module, shortcut connections are added to introduce high-resolution features. The shortcut connections introduce the features of the second, third, and fourth stages of the backbone network, and each connection is implemented by a 3×3 convolutional layer. The number of layers in the second, third, and fourth stages are 3, 2, and 1 layers respectively.

[0140] The detection module 30 receives the fused feature map generated by the feature fusion module, and predicts the category of the target, the coordinate values of the target prediction box, and the centrality of the target box. All branches and sub-modules of the detection module are connected to the feature fusion module and are all convolutional networks. The detection module includes a classification branch and a localization branch.

[0141] 1) The classification branch is used to perform convolutional processing on the fused feature map to obtain the category prediction scores. The classification branch is connected to the feature fusion module. First, it obtains classification features through 2 layers of 3×3 convolutional layers with 128 channels, and then obtains the category confidence through a 3×3 convolutional layer with the number of channels of the target categories.

[0142] 2) The localization branch is used to perform convolutional processing on the fused feature map to obtain the coordinates and centrality of the prediction box. The localization branch is connected to the feature fusion module and obtains regression features through two layers of 3×3 convolutional layers with 64 channels.

[0143] The localization branch is divided into two modules, one is the target box regression module, and the other is the centrality regression module. Among them, the target box regression module uses a 1-layer 3×3 convolutional layer with 4 channels to obtain the coordinate values of the target prediction box, and the centrality regression module is used to obtain the centrality of the target through a 1-layer 3×3 convolutional layer with a single channel.

[0144] The acquisition module 40 performs post-processing such as non-maximum suppression on the detection results to remove the target categories and prediction box coordinate values with repeated predictions, and obtains the category of the best prediction target and the coordinate values of the prediction box. In the post-processing module, the centrality and the category confidence are multiplied as the target classification confidence for non-maximum suppression to remove low-quality prediction boxes.

[0145] The update module 60 calculates the network loss, performs gradient backpropagation, updates the network parameters, and continuously iterates the training until the network converges.

[0146] The post-processing module 80 obtains the detection results using the post-processing algorithm. Figure 8This is the detection effect diagram. The upper part of the figure shows the detection results of the small target detection algorithm without using the Gaussian assignment strategy, and the lower part shows the detection effect of the small target detection using the Gaussian assignment strategy. It can be seen from the figure that the algorithm can significantly improve the detection accuracy of small and weak targets.

[0147] The small target detection system for images based on the Gaussian assignment strategy provided in this embodiment forms a multi-scale feature map set by loading an image and extracting corresponding features of the image; fuses the extracted low-level feature map and high-level feature map to generate a fused feature map; performs detection output according to the generated fused feature map to predict the category of the target, the coordinate values of the target prediction box, and the centrality of the target box; uses the Gaussian assignment strategy to assign positive and negative samples of the image; calculates the network loss by combining the assigned positive and negative samples with the target category, prediction box coordinates, and centrality output by the network; updates the network parameters according to the calculated network loss, repeats the above training steps, and continuously updates the network parameters until the network converges; uses the trained network to predict the category of the target, the coordinate values of the target prediction box, and the centrality of the target box in the input image; performs non-maximum suppression processing on the detection results to remove the target categories and prediction box coordinate values with repeated predictions and obtain the best detection results. The Gaussian assignment strategy proposed by the small target detection system for images based on the Gaussian assignment strategy provided in this embodiment can increase the number of positive samples assigned to small targets and balance the number of positive samples assigned to objects of different sizes; the proposed Gaussian assignment strategy can use the normalized Gaussian probability as the sample weight to re-weight the contribution of positive samples during the calculation of the loss and balance the influence of objects of different sizes; the provided algorithm can on average double the performance of the small target detector, and for very small targets, it can even triple the performance; the performance is 3.4% higher than that of the most advanced competitors (from 20.8% to 24.2%), and at the same time, a 7-fold acceleration is achieved.

[0148] Please see Figure 10 , Figure 10 This is the functional block diagram of the second embodiment of the small target detection system for images based on the Gaussian assignment strategy provided by the present invention. On the basis of the first embodiment, the small target detection system for images based on the Gaussian assignment strategy further includes a heatmap mapping unit 41, a training sample encoding unit 42, and a regression weight unit 43.

[0149] The heatmap mapping unit 41 is used to map the target annotation box into a two-dimensional heatmap. , where N represents the batch size, C represents the number of categories, H represents the height, W represents the width, and S represents the downsampling rate.

[0150] First, for an annotation box , belonging to the th class, it is linearly mapped to the scale of the feature map.

[0151] Then, a two-dimensional Gaussian distribution is adopted to generate a heat map ;

[0152] The probability density function of the two-dimensional Gaussian distribution is given by the following formula:

[0153] (8)

[0154] In formula (8), represents the coordinates of the Gaussian distribution , represents the mean vector of the Gaussian distribution, represents the covariance matrix of the Gaussian distribution.

[0155] (9)

[0156] (10)

[0157] In formulas (9) and (10), is used to limit and to balance the number of positive samples assigned to objects of different sizes.

[0158] Finally, the is updated by taking the element-wise maximum of in the channel.

[0159] For the rotated annotation box , a rotated two-dimensional Gaussian distribution is adopted to generate the heat map of the target,

[0160] (11)

[0161] In formula (11), is the rotation width, represents the rotation height, represents the rotation angle.

[0162] The training sample encoding unit 42 is used to encode the training samples with the heat map generated from the annotation, and a preset α threshold is used to control the proportion of positive samples. Samples in the heat map greater than the α threshold are positive samples, and samples less than or equal to the α threshold are negative samples.

[0163] The centrality of the target of the sample is defined as:

[0164] (12)

[0165] In formula (12), They respectively represent the distances from the target center point to the left, right, top, and bottom sides of the target bounding box.

[0166] The regression weight unit 43 is used to apply the encoded training samples to the regression weights.

[0167] Suppose In the sub-region of the th annotation bounding box is the regression sample weight,

[0168] (13)

[0169] In formula (13), is the Gaussian probability at and is the area of the th box.

[0170] Applying the Gaussian distribution to the regression weights can make full use of more annotation information contained in large objects and retain the annotation information of small objects; it can also emphasize these samples near the object center and reduce the impact of low-quality samples.

[0171] Since most of the small targets to be detected are not strict quadrilaterals, there are often a lot of backgrounds near the bounding boxes of small targets. In these objects, the foreground pixels are concentrated in the center of the bounding box. The label assignment strategy based on the Gaussian distribution can increase the number of positive samples assigned to small targets and ensure that at least one positive sample is assigned to each target. For objects of normal size, the applied strategy will select high-quality positive samples near the center and reduce the impact of low-quality samples. At the same time, the strategy for small targets effectively reduces the difference in the number of positive samples assigned to objects of different sizes by restricting the maximum radius of the object.

[0172] The small object detection system for images based on the Gaussian assignment strategy provided in this embodiment maps the target annotation box into a two-dimensional heat map; encodes the training samples using the heat map generated by the annotation, and uses a preset α threshold to control the proportion of positive samples. Samples greater than the α threshold in the heat map are positive samples, and samples less than or equal to the α threshold in the heat map are negative samples; applies the encoded training samples to the regression weights. The small object detection system for images based on the Gaussian assignment strategy provided in this embodiment. The proposed Gaussian assignment strategy can increase the number of positive samples assigned to small objects and balance the number of positive samples assigned to objects of different sizes; the proposed Gaussian assignment strategy can use the normalized Gaussian probability as the sample weight to re-weight the contribution of positive samples during the calculation of the loss and balance the influence of objects of different sizes; the provided algorithm can on average double the performance of the small object detector, and even triple it for very tiny objects; the performance is 3.4% higher than that of the most advanced competitors (from 20.8% to 24.2%), and at the same time achieves a 7-fold acceleration.

[0173] The small object detection system for images based on the Gaussian assignment strategy provided in this embodiment. The entire detection process of small and weak objects based on the Gaussian assignment strategy is divided into two stages: training and inference. The training flow chart is as Figure 7 shown. In the training stage, the following operations are performed:

[0174] 1) Use the preprocessing module to cut the image into sample images of the same size for convenient network processing, and at the same time perform simple data augmentation work such as horizontal or vertical flipping of the image, and load the image and related annotations.

[0175] 2) Use the feature extraction module to receive the sample image, extract the features of the sample image, and form a multi-scale feature map. Since the target size is small, in order to enhance the detection accuracy, a feature map with 4-fold downsampling is added for output.

[0176] 3) Use the feature fusion module to perform feature fusion on the low-level feature map and the high-level feature map extracted by the feature extraction module to generate a fused feature map with a downsampling rate of 4 times and a channel number of 128.

[0177] 4) Use the detection module to receive the fused feature map generated by the feature fusion module, and predict the category of the target, the coordinate values of the target prediction box, and the centrality of the target box.

[0178] 5) Adopt a label assignment module based on Gaussian classification to encode the training samples and obtain the regression weights. The specific implementation steps are as follows. First, map the target annotation box to a two-dimensional heat map . Here, N, C, H, W, and S represent the batch size, the number of categories, height, width, and downsampling rate respectively. For an annotation box , belongs to the The class is first linearly mapped to the scale of the feature map. Then, a two-dimensional Gaussian distribution is adopted to generate the target heat map .

[0179] is used to limit and the maximum value to balance the number of positive samples assigned to objects of different sizes. Finally, is updated by taking the element-wise maximum of to update the in the channel.

[0180] Secondly, the training samples are encoded. The heat maps generated from the annotations are used to encode the training samples. Alpha is used to control the proportion of positive samples. All samples greater than the alpha threshold in all heat maps are positive samples, and the samples in other positions are negative samples. The centrality of the target is defined as:

[0181] (14)

[0182] In formula (14), respectively represent the distances from the center point of the target to the left, right, top, and bottom sides of the target box.

[0183] Finally, the regression weights are obtained. Assume in the sub-region of the th annotation box inside, is the regression sample weight.

[0184] Applying the Gaussian distribution to the regression weights can make full use of the more annotation information contained in large objects and retain the annotation information of small objects. It can also emphasize these samples near the center of the object and reduce the impact of low-quality samples.

[0185] 6) Calculate the losses of the above positive and negative samples, perform gradient backpropagation, update the network parameters, and continuously iterate to train the network until the network converges. The loss function consists of the following parts.

[0186] (15)

[0187] In formula (15), is the classification loss, and Focalloss is adopted here. , are the loss coefficients used to adjust the contributions of different losses. Here, is taken, . is the regression loss, and IoUloss is adopted; is namely ; For the centrality loss, the L1 loss is adopted.

[0188] The flowchart of the inference stage is as Figure 8 shown. In the inference stage, the following operations are performed:

[0189] 1) Import the trained model and initialize the network parameters.

[0190] 2) Use the preprocessing module to preprocess the image and load the image.

[0191] 3) Use the feature fusion module to receive the sample image, extract the features of the sample image, and form a multi-scale feature map.

[0192] 4) Use the feature fusion module to fuse the low-level feature map and the high-level feature map extracted by the feature fusion module, and generate a fused feature map with a downsampling rate of 4 times and 128 channels.

[0193] 5) Use the detection module to receive the fused feature map generated by the feature fusion module, and predict the category of the target, the coordinate values of the target prediction box, and the centrality of the target box.

[0194] 6) Use the post-processing module to mainly perform post-processing such as non-maximum suppression on the detection results. Multiply the centrality and the class confidence as the target classification confidence for non-maximum suppression to remove low-quality prediction boxes.

[0195] In this embodiment, the feature fusion module fuses the low-resolution high-semantic features onto the high-resolution low-semantic feature map.

[0196] The Gaussian label assignment strategy proposed in this embodiment increases the number of positive samples assigned to small targets, and at the same time balances the number of positive samples assigned to objects of different sizes.

[0197] In this embodiment, the normalized Gaussian probability is used as the sample weight, and the contribution of positive samples is re-weighted during the calculation of the loss, balancing the influence of objects of different sizes, and making the positions closer to the center of the bounding box have higher weights.

[0198] The following is illustrated with a specific example:

[0199] This experiment was conducted on the remote sensing image small target detection dataset AI-TOD (tiny object detection dataset), which is a challenging dataset designed for tiny target detection in aerial images. It includes 8 object categories, containing 700621 object instances, distributed in 28036 aerial images of × 800×800 pixels. The average value and standard deviation of the target size in AI-TOD are only 12.8 pixels and 5.9 pixels, which are much smaller than those of general detection datasets.

[0200] During training, 24 epochs of training were performed using the Stochastic Gradient Descent (SGD) optimizer, with a momentum of 0.9, a weight decay of 0.0001, and a batch size of 6. The initial learning rate was set to 0.01 and decayed at the 16th and 22nd epochs. During testing, a preset score of 0.05 was used to filter out background bounding boxes, and the Non-Maximum Suppression (NMS) was applied with an Intersection Over Union (IOU, the ratio of the intersection to the union of the predicted object bounding box and the ground truth bounding box) threshold of 0.65. All images were resized to 800×800 for training and testing. No other data augmentation methods were used except for simple horizontal flipping in all experiments. Additionally, to speed up training, the floating-point 16 (FP16) algorithm was adopted. The inference speed was tested on a server with 1 GTX2080Ti GPU (Graphics Processing Unit).

[0201] Effect of the positive sample threshold α of Instance 1 on performance.

[0202] This parameter is used to select positive samples from candidates. Several experiments were conducted to study the influence of hyperparameters. In this comparative experiment, the hyperparameter maximum size was set to infinity.

[0203] Table 1 shows the influence of different thresholds α on performance. When α = 1.0, the assignment strategy is the same as that of CenterNet, with only the center point being a positive sample and the others being negative samples. When α decreases, both the mean and variance of the number of positive samples assigned to each object gradually increase, and the corresponding performance shows a trend of rising first and then falling. Because the more positive samples there are, the more supervision signals there are, but an imbalance in the number of positive samples will cause small objects to be overwhelmed by large objects. When α = 0.1, the best performance of 20.6% AP was obtained.

[0204] Table 1 Influence of hyperparameter α on performance

[0205]

[0206] Instance 2 Maximum size Effect on performance.

[0207] Hyperparameter is used to balance the number of positive samples assigned to large and small objects. Table 2 shows the influence of different hyperparameters on performance, and several experiments were conducted using different In this experiment, the hyperparameter α was set to 0.1. An increase in will lead to an increase in the positive samples assigned to large targets, but have little impact on small objects. When = 7, only one positive sample is specified for all objects. When increases, the mean and variance of the positive samples increase slowly, and its performance also improves accordingly. When = 16 (slightly larger than the average size of the targets), the maximum number limit of positive sampling is 9, and the standard deviation of positive sampling is reduced to 1.7. The best performance of 21.7% AP is obtained.

[0208] Table 2 Influence of hyperparameter β on performance

[0209]

[0210] Example 3 Influence of regression reweighting on performance.

[0211] Assigning different weights to each positive sample according to the Gaussian distribution can highlight the positive samples near the target center, reduce the influence of blurred and low-quality samples. It can also make full use of the more annotation information contained in large objects and retain the annotation information of small objects. Another method opposite to this is to normalize by the annotation box. Table 3 shows the influence of regression reweighting on performance. After regression weighting, the performance of TTFNet and ATSS-S is improved by 1.1% and 0.4% respectively.

[0212] Table 3 Influence of regression reweighting on performance

[0213]

[0214] In summary, the beneficial effects achieved by the small object detection method and system for images based on the Gaussian assignment strategy provided in this embodiment are as follows:

[0215] The Gaussian assignment strategy proposed in this embodiment can increase the number of positive samples assigned to small targets and balance the number of positive samples assigned to objects of different sizes.

[0216] The Gaussian assignment strategy proposed in this embodiment can use the normalized Gaussian probability as the sample weight to reweight the contribution of positive samples during the calculation of the loss and balance the influence of objects of different sizes.

[0217] The algorithm of this embodiment can on average double the performance of the small object detector, and even triple it for very tiny targets. The performance of the present invention is 3.4% higher than that of the most advanced competitors (from 20.8% to 24.2%), and at the same time achieves a 7-fold acceleration.

[0218] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. An image small target detection method based on a Gaussian distribution strategy, characterized in that, Including the following steps: Loading an image and extracting corresponding features of the image to form a set of multi-scale feature maps; Performing feature fusion on the extracted low-level feature map and high-level feature map to generate a fused feature map; Performing detection output according to the generated fused feature map, predicting the category of the target, the coordinate values of the target prediction box, and the centrality of the target box; Using a Gaussian assignment strategy to assign positive and negative samples of the image; Combining the assigned positive and negative samples with the target category, prediction box coordinates, and centrality output by the network to calculate the network loss; Updating the network parameters according to the calculated network loss, repeating the above steps, and continuously updating the network parameters until the network converges; Using the trained network to predict the category of the target, the coordinate values of the target prediction box, and the centrality of the target box in the input image; Performing non-maximum suppression processing on the detection results, removing the target categories and prediction box coordinate values with repeated predictions, and obtaining the final detection results; After the step of loading and extracting corresponding features of the images to form a set of multi-scale feature maps, the method further includes: mapping the target annotation box into a two-dimensional heat map , where N represents the batch size, C represents the number of categories, H represents the height, W represents the width, and S represents the downsampling rate; Encoding the training samples with the heat map generated by the annotation, using a preset α threshold to control the proportion of positive samples, where the samples in the heat map greater than the α threshold are positive samples, and the samples in the heat map less than or equal to the α threshold are negative samples; Applying the encoded training samples to the regression weights; Map the target annotation box to a two-dimensional heat map including the following steps: For an annotation box , belonging to the category, it will be linearly mapped to the scale of the feature map; Adopt a two-dimensional Gaussian distribution Generate a heat map ; The probability density function of the two-dimensional Gaussian distribution is given by the following formula: where x represents the coordinate of the Gaussian distribution 、 represents the mean vector of the Gaussian distribution, represents the covariance matrix of the Gaussian distribution; Among them, used to limit and the maximum value to balance the number of positive samples allocated to objects of different sizes; Update the by performing element-wise maximum processing on in the channel; For the rotated bounding box , a rotated two-dimensional Gaussian distribution is used to generate the heatmap of the target. Among them, is the rotation width, represents the rotation height, represents the rotation angle.

2. The method for detecting small targets in images based on the Gaussian distribution strategy according to claim 1, wherein, In the step of encoding the training samples with the heat map generated by the annotation, using a preset α threshold to control the proportion of positive samples, where the samples in the heat map greater than the α threshold are positive samples, and the samples in the heat map less than or equal to the α threshold are negative samples, The centrality of the target of the sample is defined as: Among them, respectively represent the distances from the target center point to the left, right, top, and bottom sides of the target box.

3. The method for detecting small targets in images based on the Gaussian distribution strategy according to claim 2, characterized in that, In the step of applying the encoded training samples to the regression weights, Hypothesis In the sub-region of the n-th bounding box is the regression sample weight Among them, is the Gaussian probability at , and is the area of the th box.

4. An image small target detection system based on a Gaussian distribution strategy, characterized in that, Including: A feature extraction module (10) for loading and extracting corresponding features of the image to form a set of multi-scale feature maps; A feature fusion module (20) for performing feature fusion on the extracted low-level feature map and high-level feature map to generate a fused feature map; A detection module (30) for performing detection output according to the generated fused feature map, predicting the category of the target, the coordinate values of the target prediction box, and the centrality of the target box; An assignment module (40) for using a Gaussian assignment strategy to assign positive and negative samples of the image; A calculation module (50) for combining the assigned positive and negative samples with the target category, prediction box coordinates, and centrality output by the network to calculate the network loss; An update module (60) for updating the network parameters according to the calculated network loss, repeating the above steps, and continuously updating the network parameters until the network converges; A prediction module (70) for using the trained network to predict the category of the target, the coordinate values of the target prediction box, and the centrality of the target box in the input image; A post-processing module (80) for performing non-maximum suppression processing on the detection results, removing the target categories and prediction box coordinate values with repeated predictions, and obtaining the final detection results; The small object detection system for images based on the Gaussian distribution strategy further includes: a heatmap mapping unit (11) for mapping the target annotation box into a two-dimensional heatmap , where N represents the batch size, C represents the number of categories, H represents the height, W represents the width, and S represents the downsampling rate; A training sample encoding unit (12) for encoding the training samples with the heat map generated by the annotation, using a preset α threshold to control the proportion of positive samples, where the samples in the heat map greater than the α threshold are positive samples, and the samples in the heat map less than or equal to the α threshold are negative samples; A regression weight unit (13) for applying the encoded training samples to the regression weights; In the heat map mapping unit (11): For a bounding box , belonging to the category will be linearly mapped to the scale of the feature map; Adopt a two-dimensional Gaussian distribution Generate a heat map ; The probability density function of the two-dimensional Gaussian distribution is given by the following formula: where x represents the coordinate of the Gaussian distribution and represents the mean vector of the Gaussian distribution, and represents the covariance matrix of the Gaussian distribution; Among them, used to limit and the maximum value, in order to balance the number of positive samples allocated to objects of different sizes; Update the by performing element-wise maximum processing on for the channel; For the rotated bounding box , use a rotated two-dimensional Gaussian distribution to generate the heatmap of the target. Among them, is the rotation width, represents the rotation height, represents the rotation angle.

5. The image small target detection system based on the Gaussian distribution strategy according to claim 4, characterized in that, In the training sample encoding unit (12), the centrality of the target of the sample is defined as: Among them, respectively represent the distances from the target center point to the left, right, top, and bottom sides of the target box.

6. The image small target detection system based on the Gaussian distribution strategy according to claim 5, characterized in that In the regression weight unit (13), assume that In the sub-region of the th bounding box,[[]] is the regression sample weight. Among them, is the Gaussian probability at , and is the th box area.

Citation Information

Patent Citations

  • Small target detection method based on distribution distance

    CN113378905A