Small Target Detection Method Based on Central Surrounding Mechanism
By using a central surround mechanism method in small object detection, the context features of the target are extracted and fused, and the problems of missing feature information and large scale span in small object detection are solved, and small object detection with high accuracy and real-time are achieved.
Patent Information
- Application Number
- CN202310234913.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-13
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-03-13
AI Technical Summary
Small object detection in images is difficult to detect effectively due to the lack of feature information and large scale spans.
Using a small object detection method based on the central surround mechanism, by building a backbone network, a multi-scale feature fusion module, a context feature extraction module based on the central surround mechanism and an object detection head module, the context features of the target are directly extracted from the current scale, and an adaptive target features and context feature fusion method are designed.
While maintaining the original inference speed, the accuracy of small object detection is effectively improved and real-time computing is realized on the embedded platform.
Smart Images

Figure CN116229228B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a small target detection method based on a center surround mechanism, belonging to the technical fields of computer vision and image processing. Background Art
[0002] With the rapid development of computer hardware performance and software technology, target detection technology has received increasing attention in the field of computer vision applications and demonstrated broad application prospects in many practical application scenarios, such as intelligent monitoring, autonomous driving, robot vision, etc. Small target detection is one of the important research directions of target detection, aiming to detect targets with fewer pixels (below 32 pixels × 32 pixels) in an image. Due to the small number of pixels, small targets have problems of missing appearance feature information and large scale span, making it difficult for deep learning models to detect them from the background.
[0003] Regarding the problem of missing feature information of small targets, the existing solutions can be divided into two categories:
[0004] The first category is to use a generative adversarial network method to perform super-resolution on the image, resample the original image, generate a high-resolution image, and perform target detection on the high-resolution image. There are two detection methods for high-resolution images, namely direct detection and block detection. Direct detection takes the entire image as the input of the target detection network, with a large amount of floating-point operations in the model, high requirements for computer computing power, and extremely high computing costs. Block detection divides the original image into image blocks of the same size, detects each image block separately, and the computing cost is relatively low, but the target fusion of different block detections is more difficult. Currently, both of these methods cannot achieve real-time calculation on embedded platforms with limited computing power.
[0005] The second category of methods is to utilize the context feature information of small targets to assist the model in inferring the position and category of small targets. This type of method believes that the deep features of the image extracted by the model contain the context information of the shallow features, and combines the upsampled deep features with the shallow features to achieve the fusion of target features and context features. However, the scale span of small targets is large. During the training process of the model, the deep features are responsible for detecting large targets, and the extraction module of the deep features will be optimized in the direction of extracting large targets, weakening the included shallow context feature information.
[0006] Regarding the problem of large scale span of small targets, usually, detection scales with smaller strides are added, and smaller preset anchor boxes are designed. The model starts detecting at a level with a smaller stride. This type of method can effectively improve the small target detection accuracy of the model. However, for detection scales with smaller strides, the width and height of the features are usually twice that of the next layer, and the floating-point operation amount is at least 4 times that of the original model, making it difficult to achieve real-time operation. Summary of the Invention
[0007] The present invention aims to overcome the problem of missing feature information in small target detection, and proposes a small target detection method based on a center surround mechanism, which uses the center surround mechanism to extract context feature information to make up for the missing small target features.
[0008] The technical solution adopted by the present invention is as follows:
[0009] The small target detection method based on the center surround mechanism includes the following steps:
[0010] In the first step, a backbone network is constructed to extract image features at different scales;
[0011] In the second step, a multi-scale feature fusion module is constructed to achieve feature fusion at different scales;
[0012] In the third step, a context feature extraction module based on the center surround mechanism is constructed to extract the context features of the target from the current scale;
[0013] In the fourth step, a target detection head module is constructed to transform the multi-scale features into target categories and target positions.
[0014] The effects and benefits of the present invention: Aiming at the problem of missing features of small targets, the present invention proposes a small target detection method based on the center surround mechanism, directly extracts the context features of the target from the current scale, and designs an adaptive method for fusing target features and context features to enhance the context features of the weak part of the target features. Aiming at the problem of large scale span of small targets, on the basis of the original detection step range, the present invention improves the detection step granularity of the model and adopts a two-way parallel method to improve the inference speed of the fine-grained detection head. The present invention uses a context feature fusion based on the center surround mechanism and a two-way parallel model structure to effectively improve the accuracy of small target detection while maintaining the original inference speed. Description of the Drawings
[0015] Figure 1 It is a schematic diagram of the small target detection model structure based on the center surround mechanism of the present invention.
[0016] Figure 2 It is a comparison chart of the number of parameters and the accuracy mAP@0.5 between the present invention and an open-source target detection algorithm on the VisDrone dataset.
[0017] Figure 3 It is a comparison chart of the number of parameters and the accuracy mAP@0.5:0.95 between the present invention and an open-source target detection algorithm on the VisDrone dataset.
[0018] Figure 4It is a comparison chart of the floating-point operation volume and mAP@0.5 between the present invention and the open-source object detection algorithm on the VisDrone dataset.
[0019] Figure 5 It is a comparison chart of the floating-point operation volume and the accuracy rate mAP@0.5:0.95 between the present invention and the open-source object detection algorithm on the VisDrone dataset. Specific implementation manners
[0020] The following further illustrates the specific implementation manners of the present invention in combination with the accompanying drawings and technical solutions.
[0021] As Figure 1 shown, the small object detection method based on the center surround mechanism includes the following steps:
[0022] The first step is to construct a backbone network to extract image features at different scales.
[0023] For the small object detection covering multiple scales in the present invention, a backbone network is constructed to extract image features of 5 different scales. The structure of the backbone network is as Figure 1 shown in the backbone network part. The backbone network consists of 1 Focus module and 6 Downsampling (DS) modules. The Focus module is a 6×6 convolutional block (with a stride of 2 in this embodiment), aiming to reduce the resolution of the feature map and improve the model inference speed. Among the 6 Downsampling modules, three are 2-fold downsampling modules and the other three are 3-fold downsampling modules; the Focus module is serially connected with the three 2-fold downsampling modules in sequence. The input of the 3-fold downsampling module is the output of the corresponding 2-fold downsampling module in sequence (in the figure, DS1, DS2, and DS3 are 2-fold downsampling modules, and DS4, DS5, and DS6 are 3-fold downsampling modules). In this embodiment, the downsampling rates of the output features of each module relative to the input image are 4, 8, 16, 12, 24, and 48 in sequence. The calculation methods of the output features of each module are as follows:
[0024]
[0025]
[0026] Among them, F s (x) represents the output feature with a stride of s, DS k represents the downsampling module with the subscript k, k is the downsampling module number, and Focus(x) is the output feature of the Focus module. According to the output stride, the DS kThe inference results of the modules will be output to the next layer for multi-scale feature fusion. Because among the output features, the output features of the DS1 module contain too much noise and are only used as the input for other modules in this step to further extract effective features without being output to the next layer; the inference results of the DS2, DS3, DS4, DS5, and DS6 modules will be output to the next layer for multi-scale feature fusion.
[0027] In the second step, construct a multi-scale feature fusion module to achieve feature fusion at different scales.
[0028] The features at the bottom layer have higher resolution, contain more location and detail information, but have passed through fewer convolutional layers, lower semanticity, and more noise; the features at the high layer contain stronger semantic information but have low resolution and lack detailed features. Using the feature pyramid structure, the multi-scale features are fused to transfer the deep semantic information to the shallow layer and the shallow detailed information to the deep layer.
[0029] The image features with different downsampling rates are divided into two groups according to the multiples of the downsampling module. The first group is the features generated by the 2x downsampling module, and the second group is the features generated by the 3x downsampling module. Each group uses the features with different downsampling rates obtained in the first step to build a feature pyramid structure. The smaller the downsampling rate, the deeper the layer it is in. In the structure, first perform bottom-up feature fusion, that is, the deep features are upsampled by 2 times and concatenated with the shallow features, and then perform top-down feature fusion, that is, the shallow features are downsampled by 2 times and concatenated with the deep features. After the concatenation is completed, the output features are obtained. The parallel feature fusion structure is as Figure 1 shown in the dual-path parallel multi-scale feature fusion module.
[0030] After the feature pyramid is established, the features are output for context feature enhancement.
[0031] In the third step, construct a context feature extraction module based on the center surround mechanism to extract the context features of the target from the current scale.
[0032] The context feature extraction module based on the center surround mechanism is called the Hole module, which directly extracts the context information of the target from the current scale. The module structure can be seen in the Figure 1 Hole module part of. The Hole module consists of 4 1×1 convolutions, 3 dilated convolutions, 1 feature concatenation module, and 1 feature fusion module.
[0033] The input feature x first enters two different 1×1 convolutions to extract features f1(x) and f2(x) respectively. f1(x) is used as the input of the dilated convolution to extract context feature information; f2(x) is passed to the end of the module for fusion. The calculation method of the feature is as follows:
[0034] f1(x) = Conv1×1 (x) (3)
[0035] f2(x) = Conv 1×1 (x) (4)
[0036] where x is the output feature in the second step, f1(x) and f2(x) are convolution outputs, and Conv 1×1 is a convolution with a 1×1 convolution kernel and a stride of 1.
[0037] Three dilated convolutions are connected in series in sequence to further extract the above context feature information, and the extracted features are concatenated. The feature F context (x) extracted by the dilated convolution is calculated as follows:
[0038] F context (x) = Concat(DConv i (f1(x))), i = 1, 2, 3 (5)
[0039] where f1(x) is the feature extracted by formula (3), DConv i is the dilated convolution, i is the flag of different dilated convolutions, and Concat is the feature concatenation module that concatenates features at the channel scale.
[0040] The concatenated features need to be input into two different 1×1 convolutions to generate context features and bias features, which are then passed to the end of the module for fusion. The context features and bias features are calculated as follows:
[0041] w, w bia s = Conv 1×1 (F context (x)) (6)
[0042] where w and w bias are the context feature and the bias feature respectively, Conv 1×1 is a convolution with a 1×1 convolution kernel and a stride of 1, and F context (x) is the feature extracted by formula (5).
[0043] The feature fusion module fuses the target feature, context feature, and bias feature. During the fusion process, the bias feature is used to enhance the target original feature when it is weak. The exponential function is used to adaptively evaluate the feature strength of the target. When the feature response is high, the coefficient e approaches 0; when the feature response is weak, the coefficient e will amplify the bias feature. The calculation method of the feature fusion module is as follows:
[0044]
[0045] where F(x) is the final output feature of the Hole module.
[0046] In the fourth step, a target detection head module is constructed to transform multi-scale features into target categories and target positions.
[0047] The target detection head uses a 1×1 convolution to reduce the number of channels of the features to 1 + 4 + the number of categories. Among them, the first channel represents whether there is an object at this position, and the second to fifth channels are the position information of small targets, including the horizontal axis, vertical axis, target width, and target height of the center point of the target. The remaining number of channels is the number of target categories included in the training data. Each position represents a category and is used to indicate the category to which the detected small target belongs.
[0048] The present invention is verified on the publicly available small target dataset Visdrone, and the small target detection method based on the center surround mechanism is compared with the open-source target detection algorithm YOLOv5. The comparison contents include: the number of parameters, the number of floating-point operations, the mean average precision (mAP), and the speed.
[0049] The number of parameters can evaluate the size of the model, and the number of floating-point operations refers to the number of calculations during the inference process of the model. The mean average precision is the mean of the average precision (AP) of all categories in the test dataset. The average precision is related to the precision-recall curve generated by the model during the test.
[0050] Precision represents the proportion of correctly predicted positive samples among the positive samples predicted by the model, and the calculation method is as follows:
[0051]
[0052] Among them, TP represents the number of samples predicted as positive and with the true label also being positive, and All detections represents all the labels predicted by the model.
[0053] Recall represents the ability of the model to predict positive samples, and the calculation method is as follows:
[0054]
[0055] Among them, All GT represents all the positive samples existing in the inference image.
[0056] Precision and recall are contradictory to each other, and a Precision-Recall curve can be generated. By calculating the mean of the precision corresponding to each recall value, the average accuracy (AP) can be obtained. The mean average precision (mAP) represents the mean of the average accuracies of all categories.
[0057] The present invention verifies models of 4 different models (n, s, m, l) with the same structure. Among them, the model of model l is the standard model, and the models of other models are obtained by multiplying the width and height of the l model by a coefficient. The width and height of the model will affect the number of parameters and the amount of floating-point operations of the model. The larger the width and depth, the larger the number of parameters and the amount of floating-point operations, and the higher the accuracy. However, the inference time is also longer. You can select a model of a suitable model according to your needs. The corresponding relationship between the model model and the coefficient is shown in Table 1.
[0058] Table 1: Width and height of models of different sizes
[0059] Model Width Coefficient Height Coefficient n 0.25 0.33 s 0.5 0.33 m 0.75 0.67 l 1.0 1.0
[0060] Figure 2 It is a comparison chart of the number of parameters and the accuracy mAP@0.5 between the present invention and the open-source object detection algorithm yolov5 on the VisDrone dataset. The small object detection method based on the center surround mechanism has higher accuracy at each scale. The improvements of mAP@0.5 at each scale are 3.7, 4.4, 3.1, and 3.6 respectively. And at each scale, this model has only half of the number of parameters, and its accuracy is close to that of the yolov5 model at the next scale.
[0061] Figure 3 It is a comparison chart of the number of parameters and the accuracy mAP@0.5:0.95 between the present invention and the open-source object detection algorithm yolov5 on the VisDrone dataset. The present invention has uniform improvements at each scale, which are 2.1, 2.2, 1.6, and 1.8 respectively.
[0062] Figure 4 and Figure 5 They are respectively the comparison charts of the amount of floating-point operations and the accuracy mAP@0.5 and mAP@0.5:0.95 between the present invention and the open-source object detection algorithm yolov5 on the VisDrone dataset. As can be seen from the figure, the method of the present invention has significant improvements at different scales.
[0063] The above results show that the present invention has good performance in the small object detection task, and the improvement is more prominent in models with a smaller scale. At each scale, it can achieve a smaller number of parameters and the amount of operations, making the small object detection accuracy close to that of the open-source model at the next stage, which is convenient for real-time small object detection and deployment on embedded platforms with limited computing power.
Claims
1. A small target detection method based on a center surround mechanism, characterized in that, It includes the following steps: In the first step, construct a backbone network to extract image features at different scales; Construct a backbone network to extract image features at 5 different scales. The backbone network consists of 1 focusing module and 6 downsampling modules DS. The focusing module is a 6×6 convolutional block. Among the 6 downsampling modules, three are 2-fold downsampling modules and the other three are 3-fold downsampling modules. The focusing module is connected in series with the three 2-fold downsampling modules in sequence. The input of the 3-fold downsampling module is the output of the corresponding 2-fold downsampling module in sequence. The calculation method of the output features of each module is as follows: Among them, F s (x) represents the output feature with a step size of s, and DS k represents the downsampling module with subscript k, where k is the downsampling module number, and Focus(x) is the output feature of the focus module; according to the output stride, the inference results of the DS k modules corresponding to k = 2, 3, 4, 5, 6 will be output to the next layer for multi-scale feature fusion; In the second step, construct a multi-scale feature fusion module to achieve feature fusion at different scales; Divide the image features with different downsampling rates into two groups according to the multiples of the downsampling modules. The first group is the features generated by the 2-fold downsampling modules, and the second group is the features generated by the 3-fold downsampling modules. Each group uses the features with different downsampling rates obtained in the first step to build a feature pyramid structure. The smaller the downsampling rate, the deeper the layer. In the structure, first perform bottom-up feature fusion, that is, the deep features are upsampled by 2 times and concatenated with the shallow features, and then perform top-down feature fusion, that is, the shallow features are downsampled by 2 times and concatenated with the deep features. After the concatenation is completed, output the features; After the feature pyramid is established, output the features and perform context feature enhancement; In the third step, construct a context feature extraction module based on the center surround mechanism to extract the context features of the target from the current scale; The context feature extraction module based on the center surround mechanism is called the Hole module, which directly extracts the context information of the target from the current scale. The Hole module consists of 4 1×1 convolutions, 3 dilated convolutions, 1 feature concatenation module and 1 feature fusion module; The input feature x first enters two different 1×1 convolutions to extract features f1(x) and f2(x) respectively. f1(x) is used as the input of the dilated convolution to extract context feature information. f2(x) is passed to the end of the module for fusion. The calculation method of the features is as follows: f1(x) = Conv 1×1 (x) (3) f2(x) = Conv 1×1 (x) (4) where x is the output feature in the second step, f1(x) and f2(x) are convolution outputs, and Conv 1×1 is a convolution with a 1×1 convolution kernel and a stride of 1; Three dilated convolutions are connected in series in turn to further extract the context feature information, and the extracted features are concatenated. The feature F context (x) is calculated as follows: F context (x) = Concat(DConv i (f1(x))), i = 1, 2, 3 (5) Among them, f1(x) is the feature extracted by formula (3), DConv i is dilated convolution, i is the flag of different dilated convolutions, Concat is the feature concatenation module, and feature concatenation is performed at the channel scale; The concatenated features need to be input into two different 1×1 convolutional blocks to generate context features and bias features, which are passed to the end of the module for fusion. The calculation methods of the context features and bias features are as follows: w, W bias = Conv 1×1 (F context (x)) (6) Among them, w and w bias are context feature and bias feature respectively, Conv 1×1 is a convolution with a convolution kernel of 1×1 and a stride of 1, and F context (x) is the feature extracted from formula (5); The feature fusion module fuses the target features, context features and bias features. During the fusion process, the bias features are used to enhance the target original features when they are weak, and an exponential function is used to adaptively evaluate the feature strength of the target. The calculation method of the feature fusion module is as follows: Among them, F(x) is the final output feature of the Hole module; In the fourth step, construct a target detection head module to transform the multi-scale features into target categories and target positions; The target detection head uses a 1×1 convolution to reduce the number of channels of the features to 1 + 4 + the number of categories. Among them, the first channel represents whether there is an object at this position, and the second to fifth channels are the position information of small targets, including the horizontal axis, vertical axis, target width and target height of the target center point. The remaining number of channels is the number of target categories included in the training data. Each position represents a category and is used to represent the category to which the detected target belongs.
Citation Information
Patent Citations
A bridge vehicle wheel detection method based on a multilayer feature fusion neural network model
CN109886312A
Small target detection method based on Center Net improved multi-scale feature fusion
CN115631400A