Target detection method for small targets based on improved SSD network

By introducing the FTT module, Attention module and feature fusion module into the SSD network, the SSD network structure is improved, which solves the problems of insufficient resolution and semantic information in small target detection and improves the small target recognition accuracy and detection performance.

CN115841611BActive Publication Date: 2025-09-23SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211576273.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-09
Publication Date
2025-09-23
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

Existing technologies have low recognition accuracy in small target detection, and most SSD-based target recognition networks do not consider adding an Attention module. There is still room for structural improvement, especially in the problem of too small resolution and insufficient semantic information in feature maps.

Method used

The FTT module, Attention module and feature fusion module are introduced to improve the structure of the SSD network. ResNet-101 is used as the backbone network, and the last feature layer conv11_2 is discarded. Super-resolution and attention mechanism are combined to improve feature extraction and fusion capabilities.

Benefits of technology

The recognition accuracy and detection performance of small target objects are improved. The improved SSD network performs better in small target detection and has practical promotion value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115841611B_ABST
    Figure CN115841611B_ABST
Patent Text Reader

Abstract

The present invention discloses a target detection method for small targets based on an improved SSD network. The method replaces the backbone of the original SSD network with ResNet-101, which is conducive to the feature extraction of small target objects; introduces an Attention module, an FTT module, and a feature fusion module; and modifies and trains the SSD network structure. The present invention introduces the FTT module, the Attention module, and the feature fusion module for super-resolution restoration, effectively utilizing multi-scale feature fusion, which helps to improve the recognition accuracy of small target objects. In addition, the last feature layer conv11_2 is discarded to better identify small target objects. The present invention innovatively transforms the SSD network, making the improved SSD network more accurate in recognizing small target objects, and has practical promotion and application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to a target detection method for small targets based on an improved SSD network. Background Art

[0002] Small object detection has long been a challenging and research hotspot in computer vision. Small objects are those that occupy very few pixels in an image. For example, in the object definition of COCO, a common dataset for object detection, small objects are those smaller than 32×32 pixels (medium objects are between 32*32 and 96*96 pixels, and large objects are larger than 96*96 pixels).

[0003] The difficulty in recognizing small objects lies in the paucity of semantic information in the original image, making it difficult for the network to learn their features. Current object detection algorithms mostly use feature fusion to improve small object detection accuracy. This is primarily due to the fact that low-level feature maps contain relatively scarce semantic information but accurately represent the target location, while high-level feature maps contain rich but inaccurate semantic information. Combining these two methods can significantly improve the recognition accuracy of small objects.

[0004] Despite numerous measures, the recognition accuracy of small objects remains low. Furthermore, most current SSD-based object recognition networks do not incorporate an attention module, or their architectures could be improved. For example, they lack integration with other image processing techniques, such as super-resolution.

[0005] Super-resolution restoration is a hot field in image processing. It can effectively restore high-resolution images from low-resolution images. This feature can be well combined with the field of small target object detection to effectively solve the key problem in the field of small target object detection, that is, the resolution of small target objects in the feature map is too small and the semantic information contained is too little.

[0006] Therefore, it is of practical significance to propose a target detection method that combines the super-resolution field and the Attention mechanism. Summary of the Invention

[0007] This invention aims to overcome the shortcomings and deficiencies of existing technologies and proposes a small-target object detection method based on an improved SSD network. This method incorporates an FTT module, an Attention module, and a feature fusion module, effectively utilizing multi-scale feature fusion to improve the recognition accuracy of small objects. Furthermore, the last feature layer, conv11_2, is discarded to further enhance small-target object recognition. This method innovatively improves the SSD network, resulting in even higher recognition accuracy for small objects.

[0008] To achieve the above objectives, the technical solution provided by the present invention is: a target detection method for small targets based on an improved SSD network. The improved SSD network is an improvement of the original SSD network in four parts. The first part is to construct an FTT module for super-resolution restoration and integrate it into the SSD network to improve the resolution of the feature layer; the second part is to construct an Attention module to enhance the feature extraction capability of the network; the third part is to use a feature fusion module to enhance the multi-scale information fusion capability of the network, and perform feature fusion on the output processed by the Attention module and the FTT module; the fourth part is to improve the original SSD network structure, specifically, to change the backbone network used in the SSD network to ResNet-101 to better extract small target features. In addition, the last feature layer extracted by the SSD network is discarded because the resolution of the last feature layer is very small, does not contain small target object information, and also provides incorrect position information;

[0009] The specific implementation steps of this target detection method are as follows:

[0010] 1) Collect images for object annotation, build a dataset, divide the dataset into training set, test set, and validation set, and finally convert the dataset into XML format;

[0011] 2) Construct an SSD network and its required Attention module, FTT module, and feature fusion module. Modify the SSD network structure, changing its backbone network to ResNet-101. Then integrate the constructed Attention module, FTT module, and feature fusion module into the modified SSD network. At the same time, discard the last feature layer conv11_2 of the SSD network to obtain the final SSD network, i.e., the improved SSD network.

[0012] 3) Use the training set and validation set to train and validate the improved SSD network, iterating until the total loss of the validation set reaches the minimum state, and obtain the optimal network after training and validation;

[0013] 4) Send the images in the test set to the optimal network obtained in step 3) for testing, and set the corresponding score threshold to obtain detection targets above the score threshold, while targets below the score threshold will not be displayed.

[0014] Furthermore, in step 1), industrial cameras are used to collect data, and the collected data is divided into a training set, a test set, and a validation set; then, the data is annotated using online or offline tools, mainly by setting a pre-selection box to surround the object to be detected with the pre-selection box. After the annotation is completed, the data set is output as an XML format data set.

[0015] Furthermore, in step 2), the backbone network of the SSD network is changed to ResNet-101. ResNet-101 consists of six layers, namely layer1, layer2, layer3, layer4, layer5, and layer6; layer1 is a 7×7 convolutional layer with 64 convolution kernels and stride=2. The output of layer1 passes through a 3×3 maxpooling layer and then passes through the remaining five layers to obtain the final output feature map; layer2, layer3, layer4, layer5, and layer6 are all composed of multiple convolutional layers and residual structures. After obtaining the output of ResNet-101, as with the original SSD network, additional feature extraction layers Conv6, Conv7, Conv8, and Conv9 are added to extract features of different scales.

[0016] Furthermore, in step 2), it is necessary to construct an FTT module for super-resolution restoration of small feature layers and an Attention module based on the attention mechanism, and integrate the two modules into the SSD network. The specific structures of the FTT module and the Attention module are as follows:

[0017] The function of the FTT module is to perform super-resolution restoration on the input features to obtain a higher resolution feature map. It has two inputs. The first input is a reference feature map P that helps to restore the details. ref , P ref The resolution of the target feature map is the same as that of the target feature map, and its role is to help restore the detailed texture part of the feature map; the other input is the low-resolution feature map P that needs to increase the resolution low , P low After the input, it first passes through a Content Extrator module for rough information extraction. The Content Extrator module is composed of multiple Conv layers, BN layers and Relu layers stacked together, which can extract the object contour information required for super-resolution restoration. After being processed by the Content Extrator module, a sub-pixel convolution operation is required to increase the resolution to the same level as P ref same;

[0018] After the above steps, we get ref After the output of the same resolution, the output is divided into two branches, one of which is connected to P refAfter splicing, the spliced ​​feature map is input into the Texture Extrator module to extract the feature of the detail texture part of the feature map. The Texture Extrator module is composed of multiple Conv layers, BN layers and ReLU layers stacked together. Its function is to extract the detail texture information required for super-resolution restoration;

[0019] After obtaining the output of the Texture Extrator module, perform a matrix addition operation on it and the output of another branch after the previous sub-pixel convolution to form a residual structure. The overall formula of the FTT module is as follows:

[0020] Out FTT =E t (P ref ||E c (P low )↑ 2× )+E c (P low )↑ 2× (1)

[0021] Where, Out FTT Indicates the output of the FTT module, P low Represents the low-resolution feature map of the input, P ref Represents the high-resolution reference feature map of the input, E c Indicates the Content Extrator module, E t Indicates the Texture Extrator module, ↑ 2× It represents the upsampling operation of sub-pixel convolution;

[0022] The function of the Attention module is to use the attention mechanism to make the network's feature extraction ability for objects more targeted and enhance the network's learning ability. Its input is the feature map that needs to be Attention operated. The Attention structure of the conv4_3 and fc7 feature layers is different from that of other feature layers. This is because the resolution of the conv4_3 and fc7 feature layers is high, so the Attention coefficient matrix required by the attention mechanism is formed by these two feature layers. Specifically, the Attention module structure of the conv4_3 and fc7 feature layers is as follows:

[0023] After the Attention module of the conv4_3 and fc7 feature layers receives the input, it first passes through a residual block and then divides into two branches. The first branch consists of two residual blocks, and the other branch first passes through an Hourglass structure to generate a heatmap of the feature map, and then passes through multiple BN layers, Conv layers, and Relu layers. Then, a sigmoid operation is performed to obtain the required Attention coefficient matrix. The obtained Attention coefficient matrix is ​​then matrix multiplied with the output of the first branch. The output is then added to the first branch to form a residual structure. Finally, a residual block, an L2 norm layer, and a Relu layer are used to obtain the final output. The specific formula is as follows:

[0024] P1=res(res(res(x))) (2)

[0025] ATTmap=sig(2×δ(Hourglass(res(x)))) (3)

[0026] Out ATT =relu(L2(res((P1×ATTmap)+P1))) (4)

[0027] Where P1 represents the output of the first branch, res represents the residual block, x represents the input of the Attention module, ATTmap represents the Attention coefficient matrix, sig represents the sigmoid operation, Hourglass represents the hourglass structure used, δ represents the combination of BN layer, Relu layer and 1×1 convolution layer, Out FTT Represents the final output of the ATT module;

[0028] Compared with the Attention modules of the conv4_3 and fc7 feature layers, the Attention modules of the remaining feature layers lack the step of generating the Attention coefficient matrix. The specific structure is as follows: first, the input feature map is received, processed continuously by three residual blocks, and the output is multiplied by the Attention coefficient matrix. Specifically, whether it is multiplied by the Attention coefficient matrix formed by the conv4_3 feature layer or the Attention coefficient matrix formed by the fc7 feature layer depends on the output dimension and resolution of the final feature fusion module. The operation after multiplication is the same as the Attention module of the conv4_3 and fc7 feature layers.

[0029] Furthermore, in step 2), it is necessary to build a feature fusion module to fuse the outputs of multiple Attention modules built in the above steps to enhance the expressive power of the model. The specific structure of the feature fusion module is as follows:

[0030] The input of the feature fusion module is three feature maps that need to be fused. Since the FTT module constructed in the previous step includes the upsampling process, the three feature maps received by the feature fusion module have the same resolution. The feature map with the original resolution, that is, the highest resolution before upsampling, is regarded as the main feature map, and the other two feature maps are secondary feature maps. Two 1×1 convolution kernels are designed respectively, and the number of channels of the secondary feature map is converted to half of the main feature map, while the resolution remains unchanged. In this way, the weight of the main feature map is higher than that of the secondary feature map. Finally, the three processed feature maps are spliced ​​to obtain the final output.

[0031] Furthermore, in step 2), it is necessary to add the modules constructed in the above steps into the SSD network and discard the last feature layer conv11_2 to obtain the final network structure. The specific steps are as follows:

[0032] All feature layers except conv4_3 and conv10_2 need to add FTT modules. The input reference feature map P of each FTT module is ref For the previous feature layer, all feature layers need to add the Attention module. The output of the Attention module of the conv4_3, fc7 and conv8_2 feature layers is input into the first feature fusion module for multi-scale feature fusion. The output of the Attention module of the fc7, conv8_2 and conv9_2 feature layers is input into the second feature fusion module to obtain the second feature fusion output. In addition, conv11_2 is discarded to obtain the final network structure. This is because conv11_2 does not contain information about small target objects and also provides incorrect position information.

[0033] Furthermore, in step 3), the training set and validation set of the prepared data set are fed into the improved SSD network in batches. After the images in the training set are extracted through the network features, the probability value and position value of the detection image classification are obtained through the corresponding convolutional layer. The binary cross entropy is combined with the initial multibox_loss in the SSD network to calculate the loss value. The optimizer adjusts the network parameters according to the number of iterations in the training process, updates the learning rate, and verifies the training effect of the network with the validation set after each specific number of network trainings. The iteration is performed until the total loss of the validation set reaches the minimum state, and finally the optimal network after training and verification is obtained.

[0034] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0035] 1. We designed and constructed an FTT module based on the super-resolution restoration principle, which helps to expand low-resolution feature maps into higher-resolution feature maps, which is beneficial for small target object detection.

[0036] 2. Designed and constructed an Attention module based on the attention mechanism, which helps to improve the network's feature extraction capability.

[0037] 3. A feature fusion module was designed and constructed, which can fuse multi-scale feature information and improve the detection performance of the network.

[0038] 4. The overall structure of the SSD network was modified, including changing the backbone network and removing the last feature layer conv11_2, which improved the network detection speed and accuracy.

[0039] 5. The present invention innovatively transforms the SSD network, making the improved SSD network more accurate in recognizing small target objects, and has practical promotion and application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is a structural diagram of the FTT module.

[0041] Figure 2 This is a structural diagram of the Attention module.

[0042] Figure 3 Schematic diagram of the feature fusion module.

[0043] Figure 4 Schematic diagram of the improved SSD network structure. DETAILED DESCRIPTION

[0044] The present invention will be described in further detail below with reference to the embodiments and drawings, but the embodiments of the present invention are not limited thereto.

[0045] This embodiment discloses a target detection method for small targets based on an improved SSD network. The improved SSD network makes four improvements to the original SSD network. The first part is to construct an FTT module for super-resolution restoration and integrate it into the SSD network to improve the resolution of the feature layer; the second part is to construct an Attention module to increase the feature extraction capability of the network; the third part is to use a feature fusion module to increase the multi-scale information fusion capability of the network, and perform feature fusion on the output processed by the Attention module and the FTT module; the fourth part is to improve the original SSD network structure, specifically, the backbone network used by the SSD network is changed to ResNet-101 to better extract small target features. In addition, the last feature layer extracted by the SSD network is discarded. This is because the resolution of the last feature layer is too small, basically does not contain small target object information, and also provides incorrect position information. The specific situation of this method is as follows:

[0046] 1) Collect your own dataset or use a public dataset from the Internet. Use a standard industrial camera to capture images. The camera resolution must be greater than 300×300. Divide the dataset into training, test, and validation sets. Label the dataset by using online or offline tools to pre-select the area containing the target. Export the dataset in XML format.

[0047] 2) Construct the original SSD network. The backbone network used in the original SSD network is VGG. Experiments have shown that the feature extraction ability of the VGG network is not as good as ResNet. Therefore, its backbone network is changed to ResNet-101. ResNet-101 mainly consists of six layers: layer1, layer2, layer3, layer4, layer5, and layer6. Layer1 is a 7×7 convolutional layer with 64 convolution kernels and stride = 2. The output of layer1 passes through a 3×3 maxpooling layer and then passes through the remaining five layers to obtain the final output feature map. Layer2, layer3, layer4, layer5, and layer6 are all composed of multiple convolutional layers and residual structures.

[0048] The SSD network also adds several additional feature extraction convolutional layers Conv6, Conv7, Conv8, and Conv9, each of which consists of two convolutional layers. The first convolutional layer has a convolution kernel size of 1×1, and the second convolutional layer has a convolution kernel size of 3×3.

[0049] The number of pre-selected boxes on each scale feature map is (4, 6, 6, 6, 4), and the size of the pre-selected box is calculated according to the formula To set, where s k is the size of the pre-selected box of the k-th feature layer, s min =0.2,s max = 0.9, m is the number of feature layers used for prediction.

[0050] like Figure 1 As shown in Figure 2, the FTT module has two inputs. The first input is a reference feature map that helps recover the details. Its resolution is the same as that of the target feature map, and its function is to help recover the detailed texture part of the feature map.

[0051] The other input is a low-resolution feature map that needs to be increased in resolution. After input, it first passes through a ContentExtrator module for coarse information extraction. The ContentExtrator module is composed of multiple Conv layers, BN layers, and Relu layers stacked together. It can extract information such as the approximate outline of the object required for super-resolution restoration. After being processed by the ContentExtrator module, a sub-pixel convolution operation is required to increase the resolution to the same as the reference feature map;

[0052] After the above steps, the output with the same resolution as the reference feature map is obtained. The output is divided into two branches, one of which is spliced ​​with the reference feature map. The spliced ​​feature map is input into the TextureExtrator module to extract the features of the detail texture part of the feature map. The TextureExtrator module is also composed of multiple Conv layers, BN layers and ReLU layers. Its function is to extract the detail texture and other information required for super-resolution restoration.

[0053] After obtaining the output of the Texture Extrator module, perform a matrix addition operation on it and the output of another branch after the previous sub-pixel convolution to form a residual structure. The overall formula of the FTT module is as follows:

[0054] Out FTT =E t (P ref ||E c (P low )↑ 2× )+E c (P low )↑ 2× (1)

[0055] Where, Out FTT Indicates the output of the FTT module, P low Represents the low-resolution feature map of the input, P ref Represents the high-resolution reference feature map of the input, Ec Indicates the Content Extrator module, E t Indicates the Texture Extrator module, ↑ 2× It represents the upsampling operation of sub-pixel convolution;

[0056] There are two structures of the Attention module. The first is the Attention module that the conv4_3 and fc7 feature layers need to pass through, and the second is the Attention module that the other feature layers pass through. The first Attention module needs to construct the Attention coefficient matrix that generates its own feature layer, that is, the Attentionmap. The other feature layers do not need to generate their own Attentionmaps. Instead, after upsampling, they directly use the Attentionmaps generated by the previous conv4_3 and fc7 feature layers to perform the Attention operation. This is because the resolution of the feature layers other than the conv4_3 and fc7 feature layers is too small, resulting in too little information about small target objects on the generated Attentionmaps, which are of no value.

[0057] The Attention module structure of conv4_3 and fc7 feature layers can be found in Figure 2 As shown, its input is the feature map that needs to be processed by Attention. After receiving the input, it first passes through a residual block and then divides into two branches. The first branch consists of two residual blocks, and the other branch first passes through an Hourglass structure to generate a heatmap of the feature map, and then passes through multiple BN layers, Conv layers and Relu layers, and then a sigmoid operation is performed to obtain the required Attention coefficient matrix. The obtained Attention coefficient matrix is ​​then matrix multiplied with the output of the first branch, and the obtained output is added to the first branch to form a residual structure. Finally, it passes through a residual block, an L2 norm layer and a Relu layer to obtain the final output. The specific formula is as follows:

[0058] P1=res(res(res(x))) (2)

[0059] ATTmap=sig(2×δ(Hourglass(res(x)))) (3)

[0060] Out ATT =relu(L2(res((P1×ATTmap)+P1))) (4)

[0061] Where P1 represents the output of the first branch, res represents the residual block, x represents the input of the Attention module, ATTmap represents the Attention coefficient matrix, sig represents the sigmoid operation, Hourglass represents the hourglass structure used, and represents the combination of BN layer, Relu layer and 1×1 convolution layer. ATT Represents the final output of the ATT module;

[0062] Compared with the Attention modules of the conv4_3 and fc7 feature layers, the Attention modules of the other feature layers lack the step of generating the Attention coefficient matrix. The specific structure is as follows: first, the input feature map is received, processed by three residual blocks in succession, and the output is multiplied by the Attention coefficient matrix. Specifically, whether it is multiplied by the Attention coefficient matrix formed by the conv4_3 feature layer or the Attention coefficient matrix formed by the fc7 feature layer depends on which feature fusion module is finally input. The operation after multiplication is the same as the Attention module of the conv4_3 and fc7 feature layers.

[0063] like Figure 3 As shown in the figure, the input of the feature fusion module is three feature maps that need to be fused. Since the FTT module constructed in the previous step includes the upsampling process, the three feature maps received by the feature fusion module have the same resolution. The feature map with the original resolution, that is, the largest resolution before upsampling, is regarded as the main feature map, and the other two feature maps are secondary feature maps. Two 1×1 convolution kernels are designed respectively, and the number of channels of the secondary feature map is converted to half of the main feature map, while the resolution remains unchanged. In this way, the weight of the main feature map can be higher than that of the secondary feature map. Finally, the three processed feature maps are spliced ​​to obtain the final output.

[0064] like Figure 4 As shown in the figure, after building the feature fusion module, all modules built in the previous steps need to be added to the SSD network, and the FTT module is added after all feature layers except conv4_3 and conv10_2. All feature layers except conv10_2 are subjected to Attention processing, and the outputs of the three feature layers conv4_3, fc7 and conv8_2 are input into the first feature fusion module, and the three feature layers fc7, conv8_2 and conv9_2 are input into the second feature fusion module to generate two outputs that need to be detected. In addition, the last feature layer con11_2 is discarded to obtain a complete network structure.

[0065] 3) The training set and validation set of the prepared data set are sent to the improved SSD network in batches. After the images in the training set are extracted through the network features, the probability value and position value of the detection image classification are obtained through the corresponding convolutional layer. The regression loss of all positive label box prediction results, the cross entropy loss of all positive label prediction results, and the cross entropy loss of the prediction results of a certain type of negative label are calculated respectively, and added according to a certain ratio to obtain the final loss value. The optimizer adjusts the network parameters according to the number of iterations in the training process, updates the learning rate, and verifies the training effect of the network with the validation set after each specific number of network training. The iteration is performed until the total loss of the validation set reaches the minimum state, and finally the optimal network after training and verification is obtained.

[0066] 4) Send the images in the test set to the optimal network obtained in step 3) for testing, and set the corresponding score threshold to obtain detection targets above the score threshold, while targets below the score threshold will not be displayed.

[0067] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. The target detection method for small targets based on the improved SSD network is characterized by: The improved SSD network makes four improvements to the original SSD network. The first part is to build an FTT module for super-resolution restoration and integrate it into the SSD network to increase the resolution of the feature layer. The second part is to build an Attention module to enhance the network's feature extraction capability. The third part is to use a feature fusion module to enhance the network's multi-scale information fusion capability, fusing the outputs processed by the Attention and FTT modules. The fourth part is to improve the original SSD network structure. Specifically, the backbone network used in the SSD network is changed to ResNet-101 to better extract small target features. In addition, the last feature layer extracted by the SSD network is discarded because the last feature layer has a very low resolution, does not contain information about small targets, and also provides incorrect location information. The specific implementation steps of this target detection method are as follows: 1) Collect images for object annotation, build a dataset, divide the dataset into training set, test set, and validation set, and finally convert the dataset into XML format; 2) Construct an SSD network and its required Attention module, FTT module, and feature fusion module. Modify the SSD network structure, changing its backbone network to ResNet-101. Then integrate the constructed Attention module, FTT module, and feature fusion module into the modified SSD network. At the same time, discard the last feature layer conv11_2 of the SSD network to obtain the final SSD network, i.e., the improved SSD network. 3) Use the training set and validation set to train and validate the improved SSD network, iterating until the total loss of the validation set reaches the minimum state, and obtain the optimal network after training and validation; 4) Send the images in the test set to the optimal network obtained in step 3) for testing, and set the corresponding score threshold to obtain detection targets above the score threshold, while targets below the score threshold will not be displayed.

2. The target detection method for small targets based on the improved SSD network according to claim 1, characterized in that: In step 1), industrial cameras are used to collect data, and the collected data is divided into a training set, a test set, and a validation set. Then, the data is annotated using online or offline tools. The main task is to set a pre-selection box to surround the object to be detected. After the annotation is completed, the data set is output as an XML format data set.

3. The target detection method for small targets based on the improved SSD network according to claim 1, characterized in that: In step 2), the backbone network of the SSD network is changed to ResNet-101. ResNet-101 consists of six layers, namely layer1, layer2, layer3, layer4, layer5, and layer6. Layer1 is a 7×7 convolutional layer with 64 convolution kernels and stride=2. The output of layer1 passes through a 3×3 max pooling layer and then passes through the remaining five layers to obtain the final output feature map. Layer2, layer3, layer4, layer5, and layer6 are all composed of multiple convolutional layers and residual structures. After obtaining the output of ResNet-101, as with the original SSD network, additional feature extraction layers Conv6, Conv7, Conv8, and Conv9 are added to extract features of different scales.

4. The target detection method for small targets based on the improved SSD network according to claim 3, characterized in that: In step 2), it is necessary to build an FTT module for super-resolution restoration of small feature layers and an Attention module based on the attention mechanism, and integrate the two modules into the SSD network. The specific structures of the FTT module and the Attention module are as follows: The function of the FTT module is to perform super-resolution restoration on the input features to obtain a higher resolution feature map. It has two inputs. The first input is a reference feature map P that helps to restore the details. ref , P ref The resolution of the target feature map is the same as that of the target feature map, and its role is to help restore the detailed texture part of the feature map; the other input is the low-resolution feature map P that needs to increase the resolution low , P low After the input, it first passes through a Content Extrator module for rough information extraction. The Content Extrator module is composed of multiple Conv layers, BN layers and Relu layers stacked together, which can extract the object contour information required for super-resolution restoration. After being processed by the Content Extrator module, a sub-pixel convolution operation is required to increase the resolution to the same level as P ref same; After the above steps, we get ref After the output of the same resolution, the output is divided into two branches, one of which is connected to P ref After splicing, the spliced ​​feature map is input into the Texture Extrator module to extract the feature of the detail texture part of the feature map. The Texture Extrator module is composed of multiple Conv layers, BN layers and ReLU layers stacked together. Its function is to extract the detail texture information required for super-resolution restoration; After obtaining the output of the Texture Extrator module, perform a matrix addition operation on it and the output of another branch after the previous sub-pixel convolution to form a residual structure. The overall formula of the FTT module is as follows: Out FTT =E t (P ref ||E c (P low )↑ 2× )+E c (P low )↑ 2× (1) Where, Out FTT Indicates the output of the FTT module, P low Represents the low-resolution feature map of the input, P ref Represents the high-resolution reference feature map of the input, E c Indicates the Content Extrator module, E t Indicates the Texture Extrator module, ↑ 2× It represents the upsampling operation of sub-pixel convolution; The function of the Attention module is to use the attention mechanism to make the network's feature extraction ability for objects more targeted and enhance the network's learning ability. Its input is the feature map that needs to be Attention operated. The Attention structure of the conv4_3 and fc7 feature layers is different from that of other feature layers. This is because the resolution of the conv4_3 and fc7 feature layers is high, so the Attention coefficient matrix required by the attention mechanism is formed by these two feature layers. Specifically, the Attention module structure of the conv4_3 and fc7 feature layers is as follows: After the Attention module of the conv4_3 and fc7 feature layers receives the input, it first passes through a residual block and then divides into two branches. The first branch consists of two residual blocks, and the other branch first passes through an Hourglass structure to generate a heatmap of the feature map, and then passes through multiple BN layers, Conv layers, and Relu layers. Then, a sigmoid operation is performed to obtain the required Attention coefficient matrix. The obtained Attention coefficient matrix is ​​then matrix multiplied with the output of the first branch. The output is then added to the first branch to form a residual structure. Finally, a residual block, an L2 norm layer, and a Relu layer are used to obtain the final output. The specific formula is as follows: P1=res(res(res(x))) (2) ATTmap=sig(2×δ(Hourglass(res(x)))) (3) Out ATT =relu(L2(res((P1×ATTmap)+P1))) (4) Where P1 represents the output of the first branch, res represents the residual block, x represents the input of the Attention module, ATTmap represents the Attention coefficient matrix, sig represents the sigmoid operation, Hourglass represents the hourglass structure used, δ represents the combination of BN layer, Relu layer and 1×1 convolution layer, Out FTT Represents the final output of the ATT module; Compared with the Attention modules of the conv4_3 and fc7 feature layers, the Attention modules of the remaining feature layers lack the step of generating the Attention coefficient matrix. The specific structure is as follows: first, the input feature map is received, processed continuously by three residual blocks, and the output is multiplied by the Attention coefficient matrix. Specifically, whether it is multiplied by the Attention coefficient matrix formed by the conv4_3 feature layer or the Attention coefficient matrix formed by the fc7 feature layer depends on the output dimension and resolution of the final feature fusion module. The operation after multiplication is the same as the Attention module of the conv4_3 and fc7 feature layers.

5. The target detection method for small targets based on the improved SSD network according to claim 4, characterized in that: In step 2), it is necessary to build the feature fusion module used to fuse the outputs of multiple Attention modules built in the above steps to enhance the expressive power of the model. The specific structure of the feature fusion module is as follows: The input of the feature fusion module is three feature maps that need to be fused. Since the FTT module constructed in the previous step includes the upsampling process, the three feature maps received by the feature fusion module have the same resolution. The feature map with the original resolution, that is, the highest resolution before upsampling, is regarded as the main feature map, and the other two feature maps are secondary feature maps. Two 1×1 convolution kernels are designed respectively, and the number of channels of the secondary feature map is converted to half of the main feature map, while the resolution remains unchanged. In this way, the weight of the main feature map is higher than that of the secondary feature map. Finally, the three processed feature maps are spliced ​​to obtain the final output.

6. The target detection method for small targets based on the improved SSD network according to claim 5, characterized in that: In step 2), the modules constructed in the above steps need to be added to the SSD network, and the last feature layer conv11_2 is discarded to obtain the final network structure. The specific steps are as follows: All feature layers except conv4_3 and conv10_2 need to add FTT modules. The input reference feature map P of each FTT module is ref For the previous feature layer, all feature layers need to add the Attention module. The output of the Attention module of the conv4_3, fc7 and conv8_2 feature layers is input into the first feature fusion module for multi-scale feature fusion. The output of the Attention module of the fc7, conv8_2 and conv9_2 feature layers is input into the second feature fusion module to obtain the second feature fusion output. In addition, conv11_2 is discarded to obtain the final network structure. This is because conv11_2 does not contain information about small target objects and also provides incorrect position information.

Citation Information

Patent Citations

  • Method for detecting open-pit mine field in remote sensing image based on deep learning

    CN112270280A

  • Method for detecting and identifying floating objects on water based on improved SSD (Solid State Disk) algorithm

    CN114782772A