A Small Target Detection Method Based on Attention Mechanism
Through the improved ResNet network and CBAM module combined with FPN, the problem of low small object detection accuracy is solved, and high-precision and robust small object detection is achieved, which is suitable for practical application scenarios.
Patent Information
- Application Number
- CN202111504006.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-09
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-12-09
AI Technical Summary
The prior art is difficult to effectively detect small targets, especially small targets that occupy fewer pixels in the image and are easily obstructed or overlapped, resulting in low detection accuracy and high leakage detection rate, which cannot meet the actual application needs.
The convolutional neural network based on attention mechanism is adopted to improve the accuracy and robustness of small object detection by constructing an improved ResNet network, CBAM module and feature pyramid network FPN, and combining multi-scale feature fusion and regression networks.
It realizes high-precision detection of small targets, can operate stably in a multi-scale environment, and is suitable for actual production environments, and the detection speed and accuracy meet engineering needs.
Smart Images

Figure CN114202672B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of biometric authentication, and relates to a small target detection method based on an attention mechanism. Background Art
[0002] Object detection is also one of the four basic tasks in computer vision and has very broad application prospects. Object detection technology has great application value in both military and civilian fields. For example, it is applied in important scenarios such as airports, railway stations, ports, and UAV ground detection, as well as in video surveillance, face recognition, intelligent transportation, etc., and has achieved good results. At the same time, it also provides a technical basis for tasks such as image analysis, understanding, and behavior recognition. However, this technology is not perfect and there are some difficult problems to solve, such as the problem of difficult detection of small targets. This problem is common in daily life, such as relatively small vehicles and pedestrians in surveillance videos, pedestrians and vehicles that need to be recognized at a long distance in autonomous driving, and numerous small targets in satellite images. Small targets usually have a small pixel ratio in the image because the target to be detected in the scene is far from the camera or has a small actual physical size. Therefore, in the process of object detection, due to the different feature representation capabilities of targets of different sizes, it is difficult to learn multi-scale features, and finally, the detection accuracy of small-sized targets is low or even a large number of missed detections occur. At present, the detection effect of these small targets cannot be applied to daily life and industrial production at all, and it still needs to be greatly improved to be applied. Based on such a development background, the detection of small-sized targets has always been a very challenging and important branch in the object detection task.
[0003] Small target detection technology is to determine whether there are small targets in a given image and mark the positions of the small targets. Generally, rectangular boxes are used for marking. Small target detection has extensive and important applications in fields such as autonomous driving, medical detection, industrial production, satellite remote sensing, and criminal investigation. In the field of autonomous driving, cars often collect high-resolution scene photos through devices such as cameras. However, due to reasons such as distance, pedestrian targets or traffic signs in the photos are unlikely to be very large. But the accurate detection of these small targets profoundly affects the realization of safe autonomous driving; in the medical field, the successful detection of tiny masses in medical images is an important prerequisite for early and accurate tumor diagnosis; defect detection in industrial production can detect and locate small defects on the surface of materials to discover problems as soon as possible, which also reflects the advantages of small target detection; in satellite remote sensing images, it is necessary to effectively annotate targets such as cars, ships, and houses. However, due to distance reasons, these targets often appear as small targets, and there is also an urgent need for small target detection methods to detect such targets; in criminal investigation images, abnormal small packages, small pedestrians, small pendants in cars, small signs on clothes, and some small furnishings indoors are all key clues for solving cases. In addition, there are many other application scenarios, so small target detection has great value.
[0004] Since small target objects occupy few pixels in the image and there is little available information. The difficulties of small target detection lie in the following three aspects: First, small targets occupy few pixels. After multiple convolution and pooling operations in the deep neural network, the detector extracts fewer features. Even a small target object may become a single pixel point and cannot be detected. Second, because small targets are small, during the detection process, they will be blocked or overlapped by other nearby targets, making it difficult to segment them from other targets and achieve the positioning and classification of small targets. Third, the sizes and aspect ratios of the anchor boxes in the existing anchor box-based object detection methods are set based on medium and large targets, causing small targets to be ignored throughout the learning process, and the receptive fields in general object detection are not very friendly to small targets. The receptive field of small target features mapped back to the original image may be larger than the size of the small target in the original image, resulting in poor detection effects.
[0005] Traditional object detection methods mainly consist of region selection, feature extraction, and classifier design. First, candidate regions are selected on the image, and there can be multiple candidate boxes of different sizes. Then, feature extraction is performed on each candidate region, and the extracted features are put into the classifier for class judgment and regression processing to obtain the final detection result. This method often uses manually selected features such as Haar features, HOG features, and integral image features. However, different features need to be selected for different detection tasks, making it difficult to meet requirements in terms of generality, robustness, and portability.
[0006] With the development of deep learning technology, deep learning methods have been applied to object detection. In 2014, Girshick, Donahue and others first introduced deep learning into object detection and proposed the R-CNN network. Subsequently, techniques such as Fast R-CNN and Faster R-CNN, which are called two-stage methods, emerged. These techniques have greatly improved the accuracy of object detection. However, due to the use of the two-stage method, their speed is not very good. Therefore, single-stage techniques such as YOLO v1, YOLO v2, YOLO v3, YOLO v4, SSD, and DSSD have emerged. Although the detection accuracy of these techniques may be slightly inferior to that of the two-stage method, their detection speed is superior to that of the two-stage method. However, these methods are limited in that they are all designed for medium and large targets. Although they can detect small targets, the detection effect is not very satisfactory. Some scholars have proposed the FPN network to detect targets at different scales, thereby achieving the detection of small targets, and the detection performance of small targets has been greatly improved. However, the FPN network simply superimposes the feature maps obtained from the backbone network and the feature maps obtained from top-down upsampling to obtain new feature maps, and the spatial information and channel information in the feature maps are not fully utilized. Summary of the Invention
[0007] The purpose of the present invention is to provide a small target detection method based on an attention mechanism with high detection accuracy and good robustness.
[0008] The principle of the present invention is as follows: A dataset is constructed by using datasets such as COCO and PASCAL VOC and self-annotated images, and then the dataset is divided into a training set, a test set, and a validation set; then a preprocessing network is constructed to preprocess the input images, and then a feature extraction network, a feature fusion network, and a small target regression network are constructed, and the networks are initialized. Then, the networks are trained using the data in the training set, the test set, and the validation set to obtain optimal network parameters; then the trained networks are used to process the input images to regress the position bounding boxes of small targets.
[0009] The technical solution for achieving the purpose of the present invention is: A small target detection method based on an attention mechanism, which specifically includes the following steps:
[0010] Step 1: Use a method that combines object detection datasets and self-annotated image data to construct a small target detection dataset, preprocess the images in the dataset, and then divide them into a training set, a test set, and a validation set according to a set ratio;
[0011] Step 2: Construct the network structure of the convolutional neural network, including a feature extraction network, a feature fusion network, and a small target prediction network, and initialize the parameters; use the improved Resnet network as the feature extraction network, and decompose the Bottle Net network architecture of the Resnet network into multiple uniform branch structures; the feature fusion network adopts a module based on channel and spatial attention, namely the CBAM module, and embeds the CBAM module into the Feature Pyramid Network (FPN) for multi-scale prediction to fuse the information between multiple layers.
[0012] Step 3: Input the training samples in the training set into the initialized convolutional neural network, calculate the losses of each part according to the network propagation process, and adjust each parameter according to the losses to obtain the optimal network parameters; then test in the test set and verify in the validation set to finally obtain the trained neural network model.
[0013] Step 4: Use the trained deep convolutional neural network model to detect small targets in the image, obtain the small target detection boxes, classification, and confidence information, and mark them on the image.
[0014] Compared with the prior art, the present invention has the following significant advantages: (1) The small target detection method constructed by using deep learning has high detection accuracy, is not sensitive to changes in the actual detection environment, has good robustness, and can be applied to the actual production environment; (2) Since a multi-scale detection method is used in the network, the entire network can not only detect small targets, but also detect medium and large targets, and both the detection speed and detection accuracy can well meet the detection requirements in engineering. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a specific implementation flowchart of the present invention.
[0016] Figure 2 It is a schematic diagram of the ResNet residual module and the improved ResNet module.
[0017] Figure 3 It is a schematic diagram of bilinear interpolation.
[0018] Figure 4 It is a schematic diagram of the channel and spatial attention module.
[0019] Figure 5 It is a schematic diagram of the channel attention module.
[0020] Figure 6 It is a schematic diagram of the spatial attention module.
[0021] Figure 7 It is a schematic diagram of the FPN module with an attention mechanism added.
[0022] Figure 8 It is a training flow chart. Specific implementation manners
[0023] A small target detection method based on an attention mechanism according to the present invention specifically includes the following steps:
[0024] Step 1: Construct a small target detection dataset by combining a target detection dataset and self-annotated image data, preprocess the images in the dataset, and then divide them into a training set, a test set, and a validation set according to a set ratio;
[0025] Step 2: Construct a network structure of a convolutional neural network, including a feature extraction network, a feature fusion network, and a small target prediction network, and initialize the parameters; use an improved Resnet network as the feature extraction network, and decompose the Bottle Net network architecture of the Resnet network into multiple uniform branch structures; the feature fusion network adopts a module based on channel and spatial attention, namely the CBAM module, embed the CBAM module into the Feature Pyramid Network (FPN) for multi-scale prediction, and fuse the information between multiple layers;
[0026] Step 3: Input the training samples in the training set into the initialized convolutional neural network, calculate the losses of each part according to the network propagation process, and adjust each parameter according to the losses to obtain the optimal network parameters; then test in the test set and verify in the validation set to finally obtain a trained neural network model;
[0027] Step 4: Use the trained deep convolutional neural network model to detect small targets in the image, obtain the small target detection boxes, classification, and confidence information, and mark them on the image.
[0028] Further, the step 1 specifically includes the following steps:
[0029] (1.1) Obtain target detection images and construct a small target detection dataset. Although there is no dedicated dataset for general small target detection nowadays, there are a large number of small target objects in the COCO dataset, and these image data can be collected to construct a small target detection dataset.
[0030] (1.2) Preprocess the small target dataset. Since there are significant differences between the image data collected in natural scenes and the image data in the dataset and the expected samples, and the width and height do not meet the input requirements, the data collected in step one needs to be processed, mainly including scaling, padding, and normalization, etc.; in small target detection training, the input image required by the network is 512*512, and most of the images in our dataset do not meet the network input, so the size needs to be modified. This method is to simply scale the image size proportionally and then fill it with 0 to obtain a 512*512 input image.
[0031] The normalization process in the preprocessing method is to convert the image data format into a unified image data format and adopt the normalization formula Normalize each pixel point in the image sample.
[0032] (1.3) When dividing the training set, test set, and validation set, it is necessary to divide them in different ways according to the size of the dataset. If the data volume is not very large (less than ten thousand), the training set, validation set, and test set are divided into 3:1:1; if the data is very large, the ratio of the training set, validation set, and test set can be adjusted to 98:1:1; but when the available data is very little, some methods such as K-fold cross-validation can be used for training and validation, etc.
[0033] Furthermore, in step 2, construct a feature extraction network, a feature fusion network, and a small target regression network; specifically, it includes the following sub-steps:
[0034] (2.1) Construct a feature extraction network, which can extract the deep and shallow semantic features of the input image.
[0035] (2.2) Construct a feature fusion network, upsample the deep semantic information obtained by the feature extraction network and then fuse it with the shallow detail information to obtain the final feature map.
[0036] (2.3) Construct a small target prediction network. The small target prediction network is divided into two parts. One is a regression task module, which is used to locate the target box, and the other is a classification module, which is used to classify the target box. Using the feature map obtained by the feature fusion network as the input, the small target detection network obtains the final result through these features.
[0037] Furthermore, the sub-step (2.1) specifically includes:
[0038] Constructing a Feature Extraction Network: The improved Resnet network is used in the feature extraction network. The entire feature extraction network is composed of multiple residual modules. The forward propagation formula of a general residual module is as follows:
[0039] y = F(x, w) + x (1)
[0040] Where x and y are the input and output respectively, F(x, w) is the forward propagation formula of a general neural network, and w is the parameter related to propagation.
[0041] Decompose the BottleNet network architecture of the Resnet network into multiple uniform branch structures. Refer to depthwise separable convolution and use grouped convolution to control the number of groups through variable cardinality, that is, the number of channels of the feature map generated by each branch is n, where n > 1.
[0042] Then its forward propagation formula is:
[0043]
[0044] Where x and y are the input and output respectively, F(x, w i ) is the forward propagation formula of the neural network of each branch, and w i is the parameter related to the propagation of each branch, that is, the parameter to be trained in the network.
[0045] The method involves convolution and pooling operations. The purpose of the convolution operation is to extract the features of the image. Different feature extraction maps will be obtained according to different convolution kernels and different calculation methods. The pooling layer is sandwiched between consecutive convolution layers and is used to compress the amount of data and parameters, reducing overfitting. In short, if the input is an image, the most important role of the pooling layer is to compress the image. It has feature invariance and feature dimensionality reduction, thereby removing redundant information and extracting the most important features. In addition, the pooling operation can prevent overfitting to a certain extent and is more convenient for optimization.
[0046] The feature extraction network also includes a convolution module and a pooling module: The purpose of the convolution module is to extract the features of the image, and different feature extraction maps are obtained according to different convolution kernels and different calculation methods; the pooling module is sandwiched between consecutive convolution modules and is used to compress the amount of data and parameters;
[0047] Construct the feature extraction network according to the format of Table 1 with the above convolution module, pooling module, and improved residual module, where conv1, conv2_x, conv3_x, conv4_x, conv5_x respectively represent five modules composed of multiple convolution layers, max pooling represents max pooling, and stride is the pooling stride;
[0048] Table 1
[0049]
[0050] As shown in Table 1, the feature extraction network has a total of 49 convolutional neural network layers and one max pooling layer.
[0051] Furthermore, the sub-step (2.2) includes:
[0052] Construct a feature fusion network: The features extracted by the shallow network in the deep convolutional network have a higher resolution and stronger representation ability than those extracted by the deep network, but the semantic information they contain is very little. Although the features of the deep network have a low resolution, their feature maps contain rich semantic information. Using only the feature maps of the shallow network or the deep network alone cannot obtain satisfactory results. Therefore, a feature fusion method is needed to fuse the features of the shallow network and the deep network, so as to combine the advantages of the two types of networks to obtain a satisfactory small target detection effect.
[0053] ① In the process of feature fusion, the upsampling method needs to be used to implement it. The upsampling method used in the invention is the bilinear interpolation method. Its schematic diagram is as shown in the appendix Figure 3 Shown. Bilinear interpolation is to perform two linear transformations. First, perform a linear transformation on the X-axis to find the R points of each row:
[0054]
[0055] Then, find the P point in this area through another linear transformation:
[0056]
[0057] where (x, y) represents the position to be inserted, P 11 , P 12 , P 21 , P 22 are the 4 corner points of the position to be inserted in the bilinear interpolation method, and their coordinates are (x 1 , y 1 ), (x 1 , y 2 ), (x 2 , y 1 ), (x 2 , y 2 ), f(·) represents the pixel value at ·, T 1 is the midpoint of P 11 and P 21 , T 2 is the midpoint of P 11 and P 22 .
[0058] ②When performing feature map fusion, in order to make full use of information in different channels and spaces, a module based on channel and spatial attention (CBAM) is adopted in the invention. The structure of the CBAM module is as shown in Figure 4 , which contains two independent sub-modules, the channel attention module (CAM) (whose structure is as shown in Figure 5 ) and the spatial attention module (SAM) (whose structure is as shown in Figure 6 ), respectively performing information aggregation in channels and spaces. This not only saves parameters and computing power but also ensures that it can be integrated into the existing network architecture.
[0059] The formula for the channel attention module is:
[0060]
[0061] where σ represents the sigmoid function, W 1 , W 0 are the weights of the MLP network, and W 1 , W 0 share the ReLU activation function after W 0 .
[0062] And the formula for the spatial attention module is:
[0063]
[0064] where σ represents the sigmoid function, f 7×7 is the convolution operation, whose convolution kernel is 7*7, represents the feature map obtained through average pooling, represents the feature map obtained through max pooling;
[0065] ③The specific process of CBAM is divided into two stages: first is the channel attention module, and then is the spatial attention module.
[0066] The input feature map F(H×W×C) is respectively passed through global max pooling and global average pooling to obtain two 1×1×C feature maps. Then, they are respectively fed into a two-layer neural network, and the two layers of this neural network are shared. The number of neurons in the first layer is C / rate (rate is the reduction rate), and ReLU is used as the activation function. The number of neurons in the second layer is C. Then, the features output by the two-layer neural network are subjected to an element-wise addition operation, and then through a sigmoid activation operation to generate the final channel attention feature map. Finally, the attention feature map and the input feature map F are subjected to an element-wise multiplication operation to generate the input feature required by the Spatial attention module.
[0067] Take the feature map output by the channel attention module as the input feature map of this module. First, perform global maximum pooling and global average pooling based on channels to obtain two feature maps of H×W×1, and then concatenate these two feature maps based on channels. Then, perform a 7×7 convolution operation to reduce the dimension to 1 channel. After that, generate a spatial attention feature map through sigmoid. Finally, multiply the spatial attention feature map by the input feature of this module to obtain the finally generated feature.
[0068] After passing through the attention module, during the feature fusion process, only concatenation is required to achieve feature fusion. Moreover, this feature fusion module not only reduces the model complexity but also improves the detection performance of the model.
[0069] ④ Embed the attention module CBAM into the Feature Pyramid Network (FPN). The FPN network consists of a bottom-up and a top-down part. Add the attention module before each address where feature fusion is performed. Feature fusion in FPN consists of two parts. One part is from the feedforward backbone, and downsampling with a stride of 2 is used for each level going up. Select the last layer feature map of each level as the corresponding layer number of the bottom-up path. First, pass through the attention module, and then obtain the feature map after passing through a 1x1 convolution. The top-down process upsamples the small feature map at the top layer to the size of the feature map of the previous stage. Figure 1 Perform a concatenation operation on the feature map obtained after the 1x1 convolution and the feature map obtained by upsampling from top to bottom to obtain the final feature map for prediction. Then, perform prediction and regression at three scales to obtain the results.
[0070] Further, sub-step (2.3) includes:
[0071] Construct a small object prediction network: Since the entire model will output prediction results at three scales, not only a small object prediction network will be constructed, but also prediction networks for medium and large objects will be constructed. However, these three networks have the same network structure.
[0072] Taking the small object prediction network as an example, use convolutional layers and pooling layers to construct the small object prediction network. The constructed prediction network consists of two parts. One is a binary classification task network for judging whether the candidate box generated by the anchor box is an object, and the other is a regression task network for performing bounding box regression on the candidate box. The two sub-networks of the prediction network are both composed of convolutional layers with a convolution kernel of 3×3, and finally both have two output channels, but the meanings represented are different, representing the regression box of the detected small object, as well as the classification information and confidence of the object.
[0073] Further, in step 3, the following operations are performed: input the training set data into the network for training to finally obtain a trained neural network model, which specifically includes:
[0074] Send the images in the training set into the network designed in step 2. The specific training process of the images is as follows: Pass the 512×512-sized images through a convolutional layer with a 7×7 convolutional kernel as shown in Table 1, and then sequentially pass through the convolutional layers shown in the table. Through the entire network model, multiple prediction boxes are predicted. Then, calculate the loss based on these prediction boxes and the truly labeled boxes, thereby guiding the change of various parameters and finally obtaining the optimal model parameters.
[0075] Classification and regression are integrated into one network, so the loss function must be multi-task:
[0076]
[0077] where p i is the probability that the anchor is predicted as the target, is the probability of the GT box, and t i is a vector representing the four parameterized coordinates of the prediction box, is the parameterized coordinate corresponding to the positive sample box, N cls is the size of the mini-batch, and λ
[0078] is the weight of the regression loss;
[0079] The loss function can be divided into two parts. The left side is the classification loss value, and the right side is the regression loss value.
[0080] First, consider the classification loss, where is:
[0081]
[0082] The classification loss is cross-entropy, and its formula is:
[0083]
[0084] When is 0:
[0085]
[0086] When is 1:
[0087]
[0088] In the case of ordinary cross-entropy, for positive samples, the greater the output probability, the smaller the loss; for negative samples, the smaller the output probability, the smaller the loss. At this time, the loss function is relatively slow during the iteration of a large number of simple samples and may not be optimized to the optimal. The focal loss Focal Loss is introduced to solve this problem, and the formula for the focal loss Focal Loss is:
[0089]
[0090] On this basis, a balance factor α is introduced to balance the problem of unbalanced positive and negative samples, and its formula is:
[0091]
[0092] Among them, α takes 0.25 and γ takes 2.
[0093] The loss of the second part is the regression loss: when is 0, the regression loss is 0. When is 1, the regression loss needs to be considered. The regression loss formula is:
[0094]
[0095] Among them, R is:
[0096]
[0097] The RPN network of Faster RCNN is used to obtain candidate boxes. The specific training process is as follows: First, initialize the model parameters and train the RPN network independently. Then use the trained RPN network to train the feature extraction network and the feature fusion network. Then freeze the trained feature extraction network and the feature fusion network, and retrain the RPN network. Finally, the parameters of the trained RPN network need to be frozen, and then the feature extraction and feature fusion networks are retrained.
[0098] During the training process of the above convolutional network, one iteration process (as shown in the appendix Figure 8 ) includes: fitting the object detection through the backpropagation and gradient descent algorithms, reducing the error of the detection target position, bias, and category to reduce the error of the entire convolutional neural network, and then updating the weights in the model through forward propagation. After each 10,000 iterations or when the error between the output of the neural network and the real target is less than the set value, the training of this round is terminated.
[0099] Furthermore, the regression predicts the position, category, and confidence of small targets, including:
[0100] After the trained neural network obtained according to the above steps is input with the image to be measured, the position of the small target can be obtained through regression, and at the same time, the positions of other medium and large targets that can be regressed can also be obtained.
[0101] The present invention will be further illustrated below in conjunction with the accompanying drawings of the specification. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, those skilled in the art's various equivalent modifications of the present invention all fall within the scope defined by the appended claims of this application.
[0102] Embodiment
[0103] As Figure 1 shown, the implementation of the present invention mainly includes four steps:
[0104] Step 1: First, preprocess the images in the input image dataset and divide them into a training set, a test set, and a validation set according to a certain ratio.
[0105] Step 2: Construct the network structure of the convolutional neural network, including a feature extraction network, a feature fusion network, and a small target regression network.
[0106] Step 3: Input the training set data into the network for training, and finally obtain a trained neural network model.
[0107] Step 4: Use the trained deep convolutional neural network model to detect small targets in the image and obtain the detection boxes of the small targets at accurate positions.
[0108] In Step 1, it can be further divided into the following sub-steps:
[0109] (1.1) Obtain image data to construct a small target dataset.
[0110] Although there is not yet a dedicated dataset for small target detection, a small target detection dataset can be constructed by collecting publicly available object detection image datasets (such as the COCO dataset, the Pascal VOC dataset, etc.) and the image information marked by oneself.
[0111] (1.2) Preprocess the small target dataset.
[0112] Since there are significant differences between the image data collected in natural scenes and the image data in the dataset compared to the expected samples, and the width and height do not meet the input requirements, the data collected in Step 1 needs to be processed, mainly including scaling, padding, and normalization, etc. In small target detection training, the required input image size for the network is 512*512. Most of the images in our dataset do not meet the network input size, so we need to modify the size. This method simply scales the image size proportionally and then fills it with 0 to obtain an input image of 512*512. The specific operation is to scale the input image with width iw and height ih, and the formula is as follows:
[0113] scale=min(w / iw,h / ih) (1)
[0114] nw=iw×scale (2)
[0115] nh=ih×scale (3)
[0116] where w and h are the expected width and height, which are 512 in the invention, scale is the scaling ratio, nw and nh are the width and height after scaling respectively. Then, the scaled image is placed in the center and the boundaries are filled with 0.
[0117] The normalization process in the preprocessing method is to convert the image data format into a unified image data format and use the normalization formula to normalize each pixel point in the image sample, where x ij represents the pixel value at the point in the (i,j) position, and x min ,x max represent the minimum and maximum values of all pixels in the image sample.
[0118] (1.3) When dividing the training set, test set, and validation set, it needs to be divided in different ways according to the size of the dataset. If the data volume is not very large (less than ten thousand), the training set, validation set, and test set are divided into 3:1:1; if the data is very large, the ratio of the training set, validation set, and test set can be adjusted to 98:1:1; but when the available data is very little, some methods such as K-fold cross-validation can be used for training and validation, etc.
[0119] In Step 2, it can be further divided into the following three sub-steps: constructing a feature extraction network, a feature fusion network, and a small target regression network; specifically including the following steps:
[0120] (2.1) Construct a feature extraction network.
[0121] The improved Resnet network used in the feature extraction network, as Figure 2 shown, the entire feature extraction network is composed of multiple residual modules, and the forward propagation formula of each residual module is as follows:
[0122] y = F(x, w) + x (4)
[0123] where x and y are the input and output respectively, F(x, w) is the forward propagation formula of a general neural network, and w is the parameter related to propagation.
[0124] The improved Resnet network module refers to depthwise separable convolution and uses grouped convolution to control the number of groups through variable cardinality. That is, the number of channels of the feature map generated by each branch is n (n > 1).
[0125] Then its forward propagation formula is:
[0126]
[0127] where x and y are the input and output respectively, F(x, w i ) is the forward propagation formula of the neural network of each branch, and w i is the parameter related to the propagation of each branch, that is, the parameter to be trained in the network.
[0128] The method involves convolution and pooling operations. The purpose of the convolution operation is to extract the features of the image. Different feature extraction maps will be obtained according to different convolution kernels and different calculation methods. The pooling layer is sandwiched between consecutive convolutional layers and is used to compress the amount of data and parameters, reducing overfitting. In short, if the input is an image, the most important role of the pooling layer is to compress the image. It has feature invariance and feature dimensionality reduction, thereby removing redundant information and extracting the most important features. In addition, the pooling operation can prevent overfitting to a certain extent and is more convenient for optimization.
[0129] The above convolution module, pooling module, and improved residual module can be used to construct the feature extraction network according to the following table format. The convolution kernels of each layer used specifically are shown in Table 1.
[0130] Table 1 Feature Extraction Network Structure
[0131]
[0132] As shown in the above table, the feature extraction network has a total of 49 convolutional neural network layers and also has one max pooling layer. For the parameter initialization of this network, the number of network layers of this network can be appropriately increased or decreased in specific implementations.
[0133] (2.2) Construct the feature fusion block
[0134] In a deep convolutional network, the features extracted by the shallow network have a higher resolution and stronger representation ability than those extracted by the deep network. However, the semantic information they contain is very little, while the features of the deep network, although having a low resolution, have rich semantic information in their feature maps. Using only the feature maps of the shallow network or the deep network alone cannot obtain satisfactory results. Therefore, a feature fusion method is needed to fuse the features of the shallow network and the deep network, so as to combine the advantages of the two types of networks to obtain a satisfactory small target detection effect.
[0135] In the process of feature fusion, the method of upsampling needs to be used to achieve it. The upsampling method used in the invention is the method of bilinear interpolation. Its schematic diagram is as shown in the appendix Figure 3 shown. Bilinear interpolation is to perform two linear transformations. First, perform a linear transformation on the X-axis to find the R points of each row:
[0136]
[0137] Then, find the P point in this area through another linear transformation:
[0138]
[0139] When performing feature map fusion, in order to make full use of the information of different channels and spaces, a module based on channel and spatial attention (CBAM) is adopted in the invention. The structure of the CBAM module is as shown in Figure 4 shown. It contains 2 independent sub-modules, the channel attention module (CAM) (its structure is as shown in Figure 5 shown) and the spatial attention module (SAM) (its structure is as shown in Figure 6 shown), respectively performing attention on channels and spaces. This not only saves parameters and computing power, but also ensures that it can be integrated into the existing network architecture.
[0140] The formula of the channel attention module is:
[0141]
[0142] where σ(·) is the feature fusion function, and the sigmoid function is used. W 1 , W 0 are the weights of the MLP network, and W 1 , W 0 share W 0 and then the ReLU function is used as the activation function. F represents the feature map, AvgPool(·) is the average pooling function, and MaxPool(·) is the maximum pooling function;
[0143] And the formula of the spatial attention module is:
[0144]
[0145] where σ represents the sigmoid function, f 7×7 is a convolution operation with a convolution kernel of 7*7, represents the feature map obtained through average pooling, represents the feature map obtained through max pooling;
[0146] The specific process of CBAM is divided into two stages: first is the channel attention module, and then is the spatial attention module.
[0147] The input feature map F(H×W×C) is respectively passed through global max pooling and global average pooling to obtain two feature maps of 1×1×C. Then, they are respectively fed into a two-layer neural network, and the two layers of this neural network are shared. The number of neurons in the first layer is C / rate (rate is the reduction rate), and ReLU is used as the activation function. The number of neurons in the second layer is C. Then, the features output by the two-layer neural network are subjected to an element-wise addition operation, and then passed through a sigmoid activation operation to generate the final channel attention feature map. Finally, the attention feature map and the input feature map F are subjected to an element-wise multiplication operation to generate the input feature required by the Spatial attention module.
[0148] The feature map output by the channel attention module is used as the input feature map of this module. First, a global maximum pooling and global average pooling based on channels are performed to obtain two feature maps of H×W×1. Then, these two feature maps are concatenated based on channels. Then, a 7×7 convolution operation is performed to reduce the dimension to 1 channel. Then, a sigmoid is passed through to generate the spatial attention feature map. Finally, the spatial attention feature map and the input feature of this module are multiplied to obtain the finally generated feature.
[0149] After passing through the attention module, during the process of feature fusion, only concatenation is required to achieve feature fusion. Moreover, this feature fusion module not only reduces the model complexity but also improves the detection performance of the model.
[0150] Such as Figure 7As shown in the figure, the attention module CBAM is embedded into the Feature Pyramid Network (FPN). The FPN network contains the original feature maps obtained from the backbone network and the newly generated feature maps obtained during the top-down process. An attention module is added before each feature fusion. Each layer of the original feature maps first passes through an attention module and then is adjusted by a 1×1 convolution to obtain an improved original feature map with fused attention. The feature map to be fused with it is a deeper layer of the feature map corresponding to the original feature map in the newly generated feature maps. This feature map is first enlarged to the same size as the improved original feature map using bilinear interpolation. Finally, a 1x1 convolution is used to fuse the two feature maps of the same size to obtain the final improved feature pyramid.
[0151] (2.3) Construct a small object prediction network. Since the entire model will output prediction results at three scales, not only a small object prediction network will be constructed, but also prediction networks for medium and large objects will be constructed. However, these three networks have the same network structure.
[0152] Taking the small object prediction network as an example, a small object prediction network is constructed using convolutional layers and pooling layers. The constructed prediction network consists of two parts. One is a binary classification task network for judging whether the candidate box generated by the anchor is an object, and the other is a regression task network for performing bounding box regression on the candidate box. The two sub-networks of the prediction network are both composed of convolutional layers with a convolution kernel of 3×3 and finally have two output channels, but they represent different meanings, representing the regression box of the detected small object and the classification information and confidence of the object respectively.
[0153] In step three, the following mainly occurs: the input training set data enters the network for training, and finally a trained neural network model is obtained;
[0154] The images in the training set are sent into the network designed in step B. The specific training process of the images is as follows: The 512×512-sized images pass through a convolutional layer with a 7×7 convolution kernel as shown in Table 1, and then pass through the convolutional layers shown in the table in sequence. Through the entire network model, multiple prediction boxes are predicted, and then the loss is calculated using these prediction boxes and the boxes marked with the ground truth, thereby guiding the change of various parameters and finally obtaining the optimal model parameters.
[0155] Classification and regression are integrated into one network, so the loss function must be multi-task:
[0156]
[0157] where p i is the probability that the anchor predicts an object, The probability of the GT box, t i is a vector representing the four parameterized coordinates of the predicted box, and is the parameterized coordinates corresponding to the positive sample box. N cls is the size of the mini - batch. λ is the weight of the regression loss.
[0158] The loss function can be divided into two parts. The left side is the loss value of classification, and the right side is the loss value of regression.
[0159] First, consider the classification loss, where is:
[0160]
[0161] And the classification loss is cross - entropy, and its formula is:
[0162]
[0163] When is 0:
[0164]
[0165] When is 1:
[0166]
[0167] For ordinary cross - entropy, for positive samples, the larger the output probability, the smaller the loss. For negative samples, the smaller the output probability, the smaller the loss. At this time, the loss function is relatively slow during the iteration of a large number of simple samples and may not be optimized to the optimal.
[0168] Therefore, Focal Loss is introduced to solve this problem. The formula of Focal Loss is:
[0169]
[0170] And on this basis, a balance factor α is introduced to balance the problem of unbalanced positive and negative samples. Its formula is:
[0171]
[0172] where α is taken as 0.25 and γ is taken as 2.
[0173] The loss of the second part is the regression loss: When is 0, the regression loss is 0. When is 1, the regression loss needs to be considered. The regression loss formula is:
[0174]
[0175] where R is:
[0176]
[0177] The RPN network of Faster RCNN is used to obtain candidate boxes. The specific training process is as follows: First, initialize the model parameters and train the RPN network independently. Then, use the trained RPN network to train the feature extraction network and the feature fusion network. Then, freeze the trained feature extraction network and the feature fusion network, and retrain the RPN network. Finally, freeze the parameters of the trained RPN network, and then retrain the feature extraction and feature fusion networks.
[0178] During the training process of the above convolutional network, one iteration process (as shown in the appendix Figure 8 ) includes: fitting the object detection through the backpropagation and gradient descent algorithms, reducing the error of the entire convolutional neural network by reducing the errors of the detection target position, bias, and category, and then updating the weights in the model through forward propagation. Each time after 10,000 iterations or the error between the output of the neural network and the real target is less than the set value, terminate the training of this round.
[0179] Step 4: After inputting the test image into the trained neural network obtained according to the above steps, the position of the small target can be obtained through regression, and at the same time, the positions of other medium and large targets can also be obtained through regression.
Claims
1. A small target detection method based on the attention mechanism, characterized in that: This method specifically includes the following steps: Step 1: Use a method that combines a target detection dataset and self-annotated image data to construct a small target detection dataset. Preprocess the images in the dataset, and then divide them into a training set, a test set, and a validation set according to a set ratio; Step 2: Construct the network structure of a convolutional neural network, including a feature extraction network, a feature fusion network, and a small target prediction network, and initialize the parameters; Use the improved Resnet network as the feature extraction network, and decompose the Bottle Net network architecture of the Resnet network into multiple uniform branch structures; The feature fusion network adopts a module based on channel and spatial attention, namely the CBAM module, and embeds the CBAM module into the Feature Pyramid Network (FPN) for multi-scale prediction and fuses the information between multiple layers; Specifically includes the following steps: (2.1) Construct a feature extraction network, which extracts the deep and shallow semantic features of the input image; (2.2) Construct a feature fusion network, upsample the deep semantic information obtained by the feature extraction network, and then fuse it with the shallow detailed information to obtain the final feature map; Specifically as follows: ① In the process of feature fusion, use the bilinear interpolation upsampling method. Bilinear interpolation is to perform two linear transformations. First, perform a linear transformation on the X-axis to find the R points of each row; ② When performing feature map fusion, adopt a module based on channel and spatial attention, called the CBAM module. The CBAM module contains 2 independent sub-modules, the channel attention module (CAM) and the spatial attention module (SAM); ③ The processing flow of the CBAM module is divided into two stages: First is the channel attention module, and then the spatial attention module; ④ After passing through the CBAM module, splice the features to achieve feature fusion: Embed the CBAM module into the Feature Pyramid Network (FPN); (2.3) Construct a small target prediction network. The small target prediction network is divided into two parts. One is the regression task module, which is used to locate the target box, and the other is the classification module, which is used to classify the target box; Use the feature map obtained by the feature fusion network as the input, and the small target detection network obtains the final detection result through these features; Step 3: Input the training samples in the training set into the initialized convolutional neural network, calculate the losses of each part according to the network propagation process, and adjust each parameter according to the losses to obtain the best network parameters; Then test in the test set and verify in the validation set to finally obtain a trained neural network model; Step 4: Use the trained neural network model to detect small targets in the image, obtain the small target detection box, classification, and confidence information and annotate them in the image.
2. The small target detection method based on the attention mechanism according to claim 1, characterized in that, Step 1 specifically includes the following steps: (1.1) Obtain the target detection image and construct the small target detection dataset: Collect the image data of small target objects in the COCO dataset to construct the small target detection dataset; (1.2) Preprocess the small target detection dataset: Process the acquired image data, including scaling, padding, and normalization; Normalization refers to converting the image data format into a unified image data format and normalizing each pixel point in the image sample using the normalization formula; (1.3) Divide the training set, test set, and validation set: Divide them in different ways according to the size of the dataset. When the data volume is no more than ten thousand, divide the training set, validation set, and test set into 3:1:1; If the data volume is more than ten thousand, adjust the ratio of the training set, validation set, and test set to 98:1:
1.
3. The small target detection method based on the attention mechanism according to claim 1, characterized in that, The construction of the feature extraction network in step (2.1) is specifically as follows: The feature extraction network uses an improved Resnet network. The entire feature extraction network is composed of multiple residual modules. The forward propagation formula of the traditional residual module is as follows: y = F(x, w) + x (1) where x and y are the input and output respectively, F(x, w) is the forward propagation formula of a general neural network, and w is the parameter related to the propagation; Decompose the BottleNet network architecture of the Resnet network into multiple uniform branch structures, refer to depthwise separable convolution, and use grouped convolution to control the number of groups through the variable cardinality, that is, the number of channels of the feature map generated by each branch is n, n > 1; Then the forward propagation formula of the residual module is: where x and y are the input and output respectively, and F i (x, w i ) is the forward propagation formula of the neural network for each branch, and w i are the parameters related to the propagation of each branch, that is, the parameters to be trained in the network; The feature extraction network also includes a convolution module and a pooling module: The purpose of the convolution module is to extract the features of the image. Different feature extraction maps can be obtained according to different convolution kernels and different calculation methods; The pooling module is sandwiched between consecutive convolution modules and is used to compress the amount of data and parameters; Construct the feature extraction network in the format of Table 1 with the above convolution module, pooling module, and improved residual module, where conv1, conv2_x, conv3_x, conv4_x, conv5_x respectively represent five modules composed of multiple convolutional layers, maxpooling represents max pooling, and stride is the pooling stride; Table 1 As shown in Table 1, the feature extraction network has a total of 49 convolutional neural network layers and one max pooling layer.
4. The small target detection method based on the attention mechanism according to claim 1, characterized in that, In ① of step (2.2), the upsampling method of bilinear interpolation is used in the process of feature fusion. Bilinear interpolation is to perform two linear transformations. First, perform a linear transformation on the X axis to find the R points of each row, specifically as follows: Then find the P point in the region through another linear transformation: where (x, y) represents the position to be inserted, P 11 , P 12 , P 21 , P 22 are respectively the four corner points at the position to be inserted in the bilinear interpolation method, and their coordinates are (x 1 , y 1 ), (x 1 , y 2 ), (x 2 , y 1 ), (x 2 , y 2 ), f(·) represents the pixel value at ·, T 1 is the midpoint of P 11 and P 21 , and T 2 is the midpoint of P 11 and P 22 ; When performing feature map fusion as described in ②, a module based on channel and spatial attention is adopted, called the CBAM module. The CBAM module contains two independent sub-modules, the channel attention module, namely CAM, and the spatial attention module, namely SAM, as follows: The formula for the channel attention module is: where σ(·) is the feature fusion function, and the sigmoid function is used, W 1 , W 0 are the weights of the MLP network, and W 1 , W 0 share W 0 and then the ReLU function is used as the activation function, F represents the feature map, AvgPool(·) is the average pooling function, and MaxPool(·) is the max pooling function; And the formula for the spatial attention module is: where σ represents the sigmoid function, and f 7×7 is a convolution operation with a convolution kernel of 7*7, represents the feature map obtained through average pooling, represents the feature map obtained through max pooling; The processing flow of the CBAM module described in ③ is divided into two stages: first is the channel attention module, and then the spatial attention module; specifically as follows: The input feature map F, H×W×C is respectively passed through global max pooling and global average pooling to obtain two feature maps of 1×1×C, which are respectively fed into a two-layer neural network and share this two-layer neural network; the number of neurons in the first layer is C / rate, where rate is the reduction rate, and ReLU is used as the activation function; the number of neurons in the second layer is C; then, the features output by the two-layer neural network are subjected to an element-wise multiplication-based summation operation, and then passed through a sigmoid activation operation to generate the final channel attention feature map; finally, the channel attention feature map and the input feature map F are subjected to an element-wise multiplication operation to generate the input feature map required for the spatial attention module; The feature map output by the channel attention module is used as the input feature map of the spatial attention module; first, a channel-based global maximum pooling and global average pooling are performed to obtain two feature maps of H×W×1, and then these two feature maps are concatenated based on the channel; then, a 7×7 convolution operation is performed to reduce the dimension to 1 channel; Then, through a sigmoid activation operation, a spatial attention feature map is generated; Finally, the spatial attention feature map and the input feature map of the spatial attention module are multiplied to obtain the finally generated feature.
5. The small target detection method based on the attention mechanism according to claim 2, wherein, The construction of the small target prediction network in step (2.3) is specifically as follows: The small target prediction network is constructed using convolutional layers and pooling layers. The constructed prediction network consists of two parts, one is a binary classification task network for determining whether the candidate box generated by the anchor box anchor is a target, and the other is a regression task network for performing bounding box regression on the candidate box; both sub-networks of the prediction network are composed of convolutional layers, with a convolutional kernel of 3×3, and finally both have two output channels, one output channel is used to output the position of the regression box of the small target, and the other output channel is used to output the classification information and confidence information of the corresponding regression box.
6. The small target detection method based on the attention mechanism according to claim 1, wherein, The specific process of step 3 is as follows: The images in the training set are fed into the convolutional neural network constructed in step 2. The specific training process of the images is as follows: The images with a size of 512×512 pass through a convolutional layer with a 7×7 convolutional kernel, and then pass through convolutional layers in sequence. Multiple prediction boxes are predicted through the entire network model. Then, the loss is calculated based on these prediction boxes and the truly labeled boxes, thereby guiding the change of various parameters, and finally obtaining the optimal model parameters; Classification and regression are integrated into one network, so the loss function is multi-task: where p i is the probability that the anchor is predicted as the target, is the probability of the GT box, t i is a vector representing the four parameterized coordinates of the predicted box, is the parameterized coordinate corresponding to the positive sample box, N cls is the size of the mini - batch, and λ is the weight of the regression loss; The loss function is divided into two parts. The left side is the loss value of classification, and the right side is the loss value of regression; First, consider the classification loss, where is The classification loss is the cross-entropy loss, and the formula is: When is 0: When is 1: Given that for positive samples, the larger the output probability, the smaller the loss; for negative samples, the smaller the output probability, the smaller the loss; the Focal Loss is introduced to solve this problem, and its mathematical expression is as follows: And on this basis, a balance factor α is introduced to balance the problem of unbalanced positive and negative samples. The formula is: where α is taken as 0.25 and γ is taken as 2; The loss of the second part is the regression loss: When is 0, the regression loss is 0. When is 1, the regression loss needs to be considered. The regression loss formula is: where R is: Use the RPN network in the Faster R-CNN model to obtain candidate boxes. The specific training process is as follows: First, initialize the model parameters and train the RPN network independently; then use the trained RPN network to train the feature extraction network and the feature fusion network; Then freeze the trained feature extraction network and feature fusion network, and retrain the RPN network; finally, freeze the parameters of the trained RPN network, and then retrain the feature extraction and feature fusion networks; In the training process of the above convolutional neural network, the process of one iteration includes: fitting object detection through backpropagation and gradient descent algorithms, and then updating the weights in the model through forward propagation. Each time after 10,000 iterations or when the error between the output of the neural network and the true target is less than the set value, the training of this round is terminated.
7. The small target detection method based on the attention mechanism according to claim 1, characterized in that regressing and predicting the position, category and confidence of the small target candidate box, including: After inputting the test image into the trained neural network obtained, through regression, the position of the small target is obtained, and at the same time, the positions of other medium and large targets that can be regressed and obtained.
Citation Information
Patent Citations
Improved vehicle detection method based on anchor-frame-free detection network
CN112966747A
Rare earth mining high-resolution image recognition and positioning method
CN113033315A
Cited By
A misaligned RGBT real-time fusion detection method
CN122551119A