Object Detection Method, Device, Equipment, Storage Medium and Program Product
Through downsampling, hierarchy and attention module fusion processing in the feature extraction network, the problem of incomplete feature extraction of convolutional networks is solved, and a higher target detection accuracy is achieved.
Patent Information
- Application Number
- CN202210671790.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-15
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-06-15
AI Technical Summary
In the existing object detection methods, the convolutional network is not comprehensive enough to extract image features, resulting in low detection accuracy.
A feature extraction network is adopted, including a cascading downsampling network, a hierarchical network, an attention network and a feature fusion network. The hierarchical feature map is extracted and fused through multiple attention modules to improve the retention and weight of features.
It improves the accuracy of target detection, enhances the weight of important features, ensures that the feature map is more comprehensive, and improves the accuracy of detection.
Smart Images

Figure CN115063658B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to an object detection method, apparatus, device, storage medium, and program product. Background Art
[0002] In recent years, there has been an increasing amount of research on detecting objects using computer image processing technology. Object detection is a hot and difficult point in computer vision research. The main problems to be solved in object detection are whether there is an object to be detected in an image or video frame and the position of the detected object in the image or video frame.
[0003] Currently, object detection is mainly achieved through convolutional networks. Convolutional processing of an image can extract most of the features of the image and then perform object detection work. However, this method has the problem of incomplete feature extraction in the image, resulting in low object detection accuracy. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide an object detection method, apparatus, device, storage medium, and program product that can improve the accuracy.
[0005] In a first aspect, this application provides an object detection method. The method includes:
[0006] Input the target image into a feature extraction network. The feature extraction network includes a cascaded downsampling network, a hierarchical network, an attention network, and a feature fusion network. The attention network includes multiple attention modules; perform downsampling feature extraction on the target image through the downsampling network, perform hierarchical processing on the feature map output by the downsampling network through the hierarchical network, perform feature extraction on the multiple hierarchical feature maps output by the hierarchical network through the multiple attention modules, and perform fusion processing on the multiple feature maps output by the multiple attention modules and the feature map output by the downsampling network through the feature fusion network to obtain a target feature map, which is used for object detection.
[0007] In one embodiment, the multiple attention modules are in one-to-one correspondence with the multiple hierarchical features Figure 1 Perform feature extraction on the multiple hierarchical feature maps output by the hierarchical network through the multiple attention modules, including: input each hierarchical feature map into the corresponding attention module; perform feature extraction on the input hierarchical feature map respectively through each attention module.
[0008] In one embodiment, the attention module includes a cascaded first extraction sub-network, second extraction sub-network, and third extraction sub-network. Feature extraction is performed on the input hierarchical feature maps through each attention module, including: for each attention module, feature extraction is performed on the input hierarchical feature maps through the first extraction sub-network; for the feature maps output by the first extraction sub-network, feature extraction is performed through the first feature extraction branch and the second feature extraction branch in the second extraction sub-network to obtain two candidate feature maps, and the two candidate feature maps are fused through the second extraction sub-network; and feature extraction is performed on the feature maps output by the second extraction sub-network through the third extraction sub-network.
[0009] In one embodiment, the first extraction sub-network includes two asymmetric first convolutional layers and a first fusion layer. Feature extraction is performed on the input hierarchical feature maps through the first extraction sub-network, including: feature extraction is performed on the hierarchical feature maps through the two first convolutional layers respectively, and the feature maps output by the two first convolutional layers are fused through the first fusion layer.
[0010] In one embodiment, the first feature extraction branch includes a second convolutional layer, and the second feature extraction branch includes a second convolutional layer, a cascaded global pooling layer, a first fully connected layer, an activation function layer, and a second fully connected layer.
[0011] In one embodiment, the third extraction sub-network includes two asymmetric third convolutional layers and a second fusion layer. Feature extraction is performed on the feature maps output by the second extraction sub-network through the third extraction sub-network, including: feature extraction is performed on the feature maps output by the second extraction sub-network through the two third convolutional layers respectively, and the feature maps output by the two third convolutional layers are fused through the second fusion layer.
[0012] In one embodiment, the attention module further includes a parameter normalization sub-network. After feature extraction is performed on the feature maps output by the second extraction sub-network through the third extraction sub-network, the method further includes: parameter normalization processing is performed on the feature maps output by the third extraction sub-network through the parameter normalization sub-network.
[0013] In one embodiment, hierarchical processing is performed on the feature maps output by the downsampling network through a hierarchical network, including: hierarchical processing is performed on the feature maps output by the downsampling network by the hierarchical network based on the number of channels; wherein, the sizes of the multiple hierarchical feature maps obtained after hierarchical processing are the same as those of the feature maps output by the downsampling network, and the sum of the number of channels of the multiple hierarchical feature maps is equal to the number of channels of the feature maps output by the downsampling network.
[0014] In one embodiment, the feature fusion network performs fusion processing on multiple feature maps output by multiple attention modules and the feature map output by the downsampling network, including: performing fusion processing on the multiple feature maps output by the multiple attention modules through the feature fusion network, and performing fusion processing again on the feature map obtained after the fusion processing and the feature map output by the downsampling network through the feature fusion network.
[0015] In a second aspect, the present application also provides an object detection device. The device includes:
[0016] An input module, configured to input a target image into a feature extraction network, where the feature extraction network includes a cascaded downsampling network, a hierarchical network, an attention network, and a feature fusion network, and the attention network includes multiple attention modules;
[0017] A feature extraction module, configured to perform downsampling feature extraction on the target image through the downsampling network, perform hierarchical processing on the feature map output by the downsampling network through the hierarchical network, perform feature extraction on the multiple hierarchical feature maps output by the hierarchical network through the multiple attention modules, and perform fusion processing on the multiple feature maps output by the multiple attention modules and the feature map output by the downsampling network through the feature fusion network to obtain a target feature map, where the target feature map is used for object detection.
[0018] In one embodiment, the multiple attention modules correspond one-to-one with the multiple hierarchical features Figure 1 Specifically, the feature extraction module is configured to input each hierarchical feature map into the corresponding attention module; and perform feature extraction on the input hierarchical feature map through each attention module.
[0019] In one embodiment, the attention module includes a cascaded first extraction sub-network, a second extraction sub-network, and a third extraction sub-network. Specifically, for each attention module, the feature extraction module is configured to perform feature extraction on the input hierarchical feature map through the first extraction sub-network, perform feature extraction on the feature map output by the first extraction sub-network through the first feature extraction branch and the second feature extraction branch in the second extraction sub-network to obtain two candidate feature maps, perform fusion processing on the two candidate feature maps through the second extraction sub-network, and perform feature extraction on the feature map output by the second extraction sub-network through the third extraction sub-network.
[0020] In one embodiment, the first extraction sub-network includes two asymmetric first convolutional layers and a first fusion layer. Specifically, the feature extraction module is configured to perform feature extraction on the hierarchical feature map through the two first convolutional layers respectively, and perform fusion processing on the feature maps output by the two first convolutional layers through the first fusion layer.
[0021] In one embodiment, the first feature extraction branch includes a second convolutional layer, and the second feature extraction branch includes a second convolutional layer, a cascaded global pooling layer, a first fully-connected layer, an activation function layer, and a second fully-connected layer.
[0022] In one embodiment, the third extraction sub-network includes two asymmetric third convolutional layers and a second fusion layer. The feature extraction module is specifically configured to perform feature extraction on the feature maps output by the second extraction sub-network through the two third convolutional layers, and perform fusion processing on the feature maps output by the two third convolutional layers through the second fusion layer.
[0023] In one embodiment, the attention module further includes a parameter normalization sub-network. The feature extraction module is specifically configured to perform parameter normalization processing on the feature maps output by the third extraction sub-network through the parameter normalization sub-network.
[0024] In one embodiment, the feature extraction module is specifically configured to perform hierarchical processing on the feature maps output by the downsampling network based on the number of channels through a hierarchical network; wherein, the sizes of the multiple hierarchical feature maps obtained after hierarchical processing are the same as those of the feature maps output by the downsampling network, and the sum of the number of channels of the multiple hierarchical feature maps is equal to the number of channels of the feature maps output by the downsampling network.
[0025] In one embodiment, the feature extraction module is specifically configured to perform fusion processing on the multiple feature maps output by the multiple attention modules and the feature maps output by the downsampling network through a feature fusion network, including: performing fusion processing on the multiple feature maps output by the multiple attention modules through the feature fusion network, and performing fusion processing on the feature maps obtained after the fusion processing and the feature maps output by the downsampling network again through the feature fusion network.
[0026] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the object detection method described in any one of the first aspects above is implemented.
[0027] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the object detection method described in any one of the first aspects above is implemented.
[0028] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the object detection method described in any one of the first aspects above is implemented.
[0029] The above-mentioned object detection method, device, equipment, storage medium and program product. First, input the target image into the feature extraction network, where the feature extraction network includes a cascaded downsampling network, a hierarchical network, an attention network and a feature fusion network, and the attention network includes multiple attention modules. Then, perform downsampling feature extraction processing on the target image through the downsampling network, then perform hierarchical processing on the feature map output by the downsampling network through the hierarchical network, extract features from the multiple hierarchical feature maps output by the hierarchical network through multiple attention modules, and finally, perform fusion processing on the multiple feature maps output by the multiple attention modules and the feature map output by the downsampling network through the feature fusion network to obtain the target feature map for object detection. In this way, through hierarchical feature extraction after downsampling the target image, the downsampled target image is passed through the attention module for feature extraction to increase the weight of the features in the downsampled target image, and then the multiple feature maps output by the attention module and the feature map after downsampling processing are fused. At this time, the target feature map has both the features of the target image before hierarchical processing and the features after hierarchical processing. Therefore, the features in the target feature map are retained more comprehensively, and the weight of important features is increased, thereby improving the accuracy of object detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a schematic flowchart of the object detection method in one embodiment;
[0031] Figure 2 is a schematic flowchart of the object detection method in another embodiment;
[0032] Figure 3 is a schematic flowchart of the object detection method in another embodiment;
[0033] Figure 4 is a schematic flowchart of the object detection method in another embodiment;
[0034] Figure 5 is a schematic flowchart of the object detection method in another embodiment;
[0035] Figure 6 is a flowchart of the object detection method in another embodiment;
[0036] Figure 7 is a flowchart of the attention module in the object detection method in another embodiment;
[0037] Figure 8 is a structural block diagram of the object detection device in another embodiment;
[0038] Figure 9 is an internal structure diagram of a computer device in another embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] In order to make the objectives, technical solutions and advantages of this application more clear and understandable, the following further details this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not used to limit this application.
[0040] In one embodiment, as Figure 1 shown, a target detection method is provided. Taking the case where this method is applied to a terminal as an example for illustration, it can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. This method includes the following steps:
[0041] Step 101, the terminal inputs the target image into the feature extraction network.
[0042] Among them, the feature extraction network includes a cascaded downsampling network, a hierarchical network, an attention network and a feature fusion network. The attention network includes multiple attention modules. The target image is an image to be subjected to target detection.
[0043] Step 102, the terminal performs downsampling feature extraction on the target image through the downsampling network, performs hierarchical processing on the feature map output by the downsampling network through the hierarchical network, performs feature extraction on the multiple hierarchical feature maps output by the hierarchical network through multiple attention modules, and performs fusion processing on the multiple feature maps output by the multiple attention modules and the feature map output by the downsampling network through the feature fusion network to obtain a target feature map.
[0044] Among them, the target feature map is used for target detection, and the terminal extracts the target feature map by performing feature extraction on the target image through the feature extraction network.
[0045] First, the terminal performs downsampling processing on the target image through the downsampling network. Optionally, the downsampling network includes 5 downsampling sub-networks, and each downsampling sub-network includes a convolutional layer, an activation function layer and a max pooling layer. After passing through the downsampling network, the size of the target image is reduced and the number of channels increases. Then, the feature map output by the downsampling network is input into the hierarchical network for hierarchical processing. Optionally, the feature map can be divided into 4 layers to obtain 4 corresponding hierarchical feature maps. The 4 hierarchical feature maps are respectively input into the attention module for feature extraction. The attention module can assign different weights to different input features, thereby enhancing the representation ability of the features. Finally, the feature fusion network fuses the multiple feature maps output by the multiple attention modules and the feature map output by the downsampling network to obtain a target feature map. At this time, the target feature map has both the features of the target image before layering and the weighted features after being processed by the attention module after layering.
[0046] In the above object detection method, first, the target image is input into the feature extraction network, where the feature extraction network includes a cascaded downsampling network, a hierarchical network, an attention network, and a feature fusion network. The attention network includes multiple attention modules. Then, the target image is subjected to downsampling feature extraction processing by the downsampling network, the feature map output by the downsampling network is subjected to hierarchical processing by the hierarchical network, the multiple hierarchical feature maps output by the hierarchical network are subjected to feature extraction by multiple attention modules, and the multiple feature maps output by the multiple attention modules and the feature map output by the downsampling network are subjected to fusion processing by the feature fusion network to obtain a target feature map for object detection. In this way, after the target image is downsampled, the features are hierarchically processed, the hierarchically processed target image is subjected to feature extraction by the attention module to increase the weight of the features in the hierarchically processed target image, and then the multiple feature maps output by the attention module and the feature map after downsampling processing are fused. At this time, a target feature map with both the features of the target image before hierarchical processing and the features after hierarchical processing is obtained. Therefore, the features in the target feature map are retained more comprehensively, and the weight of the important features is increased, thereby improving the accuracy of object detection.
[0047] In one embodiment, in order to extract the features of the hierarchical feature maps, multiple attention modules correspond one-to-one with multiple hierarchical features Figure 1 as Figure 2 shown. The steps of the multiple attention modules extracting the features of the multiple hierarchical feature maps output by the hierarchical network include:
[0048] Step 201, the terminal inputs each hierarchical feature map into the corresponding attention module.
[0049] Optionally, the target image can be divided into 4 layers after being processed by the hierarchical network, that is, 4 hierarchical feature maps are obtained. Therefore, 4 attention modules can be used to process the 4 hierarchical feature maps, and the terminal inputs the 4 hierarchical feature maps into 4 corresponding attention modules respectively.
[0050] Step 202, the terminal extracts the features of the input hierarchical feature maps through each attention module respectively.
[0051] Among them, the attention module can include a cascaded first extraction sub-network, a second extraction sub-network, and a third extraction sub-network, as Figure 3 shown. The steps of each attention module extracting the features of the input hierarchical feature map specifically include:
[0052] Step 301, the terminal extracts the features of the input hierarchical feature map through the first extraction sub-network.
[0053] Among them, the first extraction sub-network includes two asymmetric first convolutional layers and a first fusion layer.
[0054] First, the terminal performs feature extraction on the hierarchical feature maps through two first convolutional layers respectively. The first convolutional layer includes convolutional processing and activation function processing. The convolutional processing of the two first convolutional layers is asymmetric convolution. Since asymmetric convolution has fewer parameters, it can speed up the training process. For example, the asymmetric convolutions are (1*3) and (3*1). Optionally, the activation function used is ReLU (The Rectified Linear Unit), and ReLU can speed up the training process and prevent gradient vanishing. Gradient vanishing will affect the efficiency of object detection.
[0055] Then, the terminal performs fusion processing on the feature maps output by the two first convolutional layers through the first fusion layer, that is, adds the feature maps output by the two first convolutional layers. Among them, the size of the feature map remains unchanged, and the number of channels is summed up. For example, the feature map output by each first convolutional layer is (C*H*W), where C represents the number of channels of the target image, H represents the height of the target image, and W represents the width of the target image. Adding the feature maps output by the two first convolutional layers, the resulting feature map is (2C*H*W).
[0056] Step 302, the terminal performs feature extraction on the feature map output by the first extraction sub-network through the first feature extraction branch and the second feature extraction branch in the second extraction sub-network, and obtains two candidate feature maps.
[0057] Optionally, the first feature extraction branch includes a second convolutional layer, and the second feature extraction branch includes a second convolutional layer and a cascaded global pooling layer, a first fully connected layer, an activation function layer, and a second fully connected layer. Among them, the second convolutional layer includes convolutional processing and activation function processing. Optionally, the convolution of the second convolutional layer is (1*1), and the activation function can use ReLU. The activation function layer in the second feature extraction branch can use the Sigmod activation function.
[0058] Step 303, the terminal performs fusion processing on the two candidate feature maps through the second extraction sub-network.
[0059] The first feature extraction branch in the second extraction sub-network performs feature extraction processing on the feature map output by the first extraction sub-network to obtain a candidate feature Figure 1 , and the second feature extraction branch in the second extraction sub-network performs feature extraction processing on the feature map output by the first extraction sub-network to obtain a candidate feature Figure 2 , and fuses the two candidate feature maps, that is, multiplies the candidate feature Figure 1 and the candidate feature Figure 2 . At this time, the weight of the target feature in the resulting feature map increases.
[0060] Step 304: The terminal extracts features from the feature map output by the second extraction sub-network through the third extraction sub-network.
[0061] Among them, the third extraction sub-network includes two asymmetric third convolutional layers and a second fusion layer. First, the terminal extracts features from the feature map output by the second extraction sub-network through the two third convolutional layers respectively. Among them, the third convolutional layer includes convolutional processing and activation function processing. The convolutional processing in the two third convolutional layers is asymmetric convolution, and the parameters of the third convolutional layer can be the same as those of the first convolutional layer. Then, the terminal performs fusion processing on the feature maps output by the two third convolutional layers, that is, adds the feature maps output by the two third convolutional layers.
[0062] In the above embodiment, the hierarchical feature map is input into the attention module for feature extraction. The attention module can assign different weights to different input features, thereby enhancing the feature representation ability and increasing the weight of the target feature in the feature map.
[0063] In one embodiment, the attention module further includes a parameter normalization sub-network. After the terminal extracts features from the feature map output by the second extraction sub-network through the third extraction sub-network, the method further includes: performing parameter normalization processing on the feature map output by the third extraction sub-network through the parameter normalization sub-network.
[0064] Optionally, the parameter normalization processing is to normalize the output of the previous layer, which can avoid the problem of gradient disappearance to a certain extent. The parameter normalization sub-network performs parameter normalization processing on the feature map output by the third extraction sub-network, and then continues the next operation as the output of the attention module.
[0065] In the above embodiment, through parameter standard processing, the stability of the data feature distribution can be ensured, and the problem of gradient disappearance can be reduced.
[0066] In one embodiment of the present application, in order to extract the hierarchical features of the target image, first, it is necessary to perform hierarchical processing on the feature map output by the downsampling network through the hierarchical network. The hierarchical network performs hierarchical processing on the feature map output by the downsampling network based on the number of channels.
[0067] Among them, the sizes of the multiple hierarchical feature maps obtained after hierarchical processing are the same as those of the feature maps output by the downsampling network. At the same time, the sum of the number of channels of the multiple hierarchical feature maps is equal to the number of channels of the feature maps output by the downsampling network. For example, the target image output by the downsampling network is (32C * H / 32 * W / 32). Dividing this target image into 4 layers, that is, 4 hierarchical feature maps are obtained. Each hierarchical feature map is (8C * H / 32 * W / 32), that is, the size, namely the height and width, of each hierarchical feature map is the same as that of the feature maps output by the downsampling network. The sum of the number of channels of the 4 hierarchical feature maps is 32C, which is equal to the number of channels of the feature maps output by the downsampling network.
[0068] In the above embodiment, the feature maps output by the downsampling network are hierarchically processed to obtain multiple hierarchical feature maps, so that feature extraction can be continued for the hierarchical feature maps.
[0069] In one embodiment, as Figure 4 shown, after obtaining the features of each hierarchical feature map, the steps of fusing the multiple feature maps output by multiple attention modules and the feature maps output by the downsampling network through the feature fusion network include:
[0070] Step 401, the terminal fuses the multiple feature maps output by multiple attention modules through the feature fusion network.
[0071] Among them, the feature fusion network includes a third fusion layer, a fourth convolutional layer, and a fourth fusion layer. The third fusion layer fuses the multiple feature maps output by multiple attention modules, that is, sums the multiple feature maps. At this time, the size of the obtained feature map remains unchanged, and the number of channels is the sum of the number of channels of the multiple feature maps. For example, the feature maps output by 4 attention modules are (8C * H / 32 * W / 32). After being summed by the third fusion layer, the output feature map is (32C * H / 32 * W / 32).
[0072] Step 402, the terminal fuses the feature map obtained after fusion processing and the feature map output by the downsampling network again through the feature fusion network.
[0073] The feature map after being summed by the third fusion layer is input into the fourth convolutional layer for processing. The fourth convolutional layer includes a convolution operation and an activation function operation. Then, the feature map output by the fourth convolutional layer and the feature map output by the downsampling network are fused through the fourth fusion layer, that is, the feature maps are summed. At this time, the size of the obtained feature map remains unchanged, and the number of channels is the sum of the number of channels of the two feature maps. At this time, the target feature map is obtained. Inputting the target feature map into the classification and detection network can perform target detection.
[0074] In the above embodiments, through the feature fusion network for fusion processing, a target feature map is obtained. At this time, the target feature map has both the target image features before layering and the weighted features processed by the attention module after layering.
[0075] In an embodiment of the present application, please refer to Figure 5 which shows a flowchart of a target detection method provided by an embodiment of the present application. The target detection method includes the following steps:
[0076] Step 501, the terminal inputs the target image into the feature extraction network.
[0077] Step 502, the terminal performs downsampling feature extraction on the target image through the downsampling network.
[0078] Step 503, the terminal performs layering processing on the feature map output by the downsampling network based on the number of channels through the layering network to obtain each layered feature map.
[0079] Step 504, the terminal inputs each layered feature map into the corresponding attention module.
[0080] Step 505, the terminal performs feature extraction on the input layered feature maps through each attention module respectively.
[0081] Step 506, the terminal performs fusion processing on the multiple feature maps output by the multiple attention modules through the feature fusion network.
[0082] Step 507, the terminal performs fusion processing on the feature map obtained after fusion processing and the feature map output by the downsampling network again through the feature fusion network to obtain the target feature map.
[0083] To facilitate the reader's understanding of the technical solution provided by the embodiment of the present application, an example of the target detection algorithm of the present application is given. Please refer to Figure 6 Figure 6 which is a flowchart of the target detection algorithm. The specific steps are as follows:
[0084] (1) Input the target image (C*H*W) into the downsampling network and perform 5 downsampling operations. Each downsampling operation includes: convolution conv(3*3), activation function relu, and max pooling maxpool(3*3).
[0085] (2) Perform layering processing on the feature map (32C*H / 32*W / 32) processed by the downsampling network to obtain 4 layered feature maps (8C*H / 32*W / 32).
[0086] (3) Input the 4 layered feature maps into the attention module for processing respectively, and then sum them to obtain the feature map (32C*H / 32*W / 32).
[0087] (4) Perform a convolution operation on the feature map output in step (3), including: convolution conv(3*3) and activation function relu. Then sum it with the feature map output in step (1) to obtain the target feature map (64C*H / 32*W / 32).
[0088] (5) Input the target feature map into the classification and detection module for object detection to obtain the object detection result.
[0089] Among them, the attention module can increase the weight of the target features in the target image. The processing flow of the attention module is as Figure 7 shown, and the specific steps include:
[0090] (1) Process the feature map (C*H*W) input into the attention module in two paths through two first convolutional layers. The first convolutional layer includes: asymmetric convolutions conv(1*3) and conv(3*1), and activation function relu. Then sum the feature maps output by the two first convolutional layers to output a feature map (2C*H*W).
[0091] (2) Process the feature map (2C*H*W) through the first feature extraction branch, including: convolution conv(1*1) and activation function relu.
[0092] (3) Process the feature map (2C*H*W) output in step (1) through the second feature extraction branch, including: convolution conv(1*1), activation function relu, global pooling global pool, fully connected layer FC, and activation function sigmod.
[0093] (4) Multiply the feature maps output in step (2) and step (3) to obtain a feature map (C*H*W).
[0094] (5) Process the feature map output in step (4) in two paths through two third convolutional layers. The third convolutional layer includes: asymmetric convolutions conv(1*3) and conv(3*1), and activation function relu. Then sum the feature maps output by the two third convolutional layers.
[0095] (6) Perform parameter normalization processing on the feature map output in step (5), including: parameter normalization LayerNorm. Then use it as the output of the attention module to continue with the next step of the object detection algorithm.
[0096] Table 1 shows the experimental comparison results of the object detection accuracy and detection time consumption of the object detection algorithm of this application and other models. The mAP (mean Average Precision) value of the object detection algorithm of this application is higher than that of other models, indicating that the object detection algorithm of this application has higher detection accuracy, while the detection time consumption of this application is lower, indicating that the object detection algorithm of this application has a faster detection speed compared to other models.
[0097] Comparison Chart of Object Detection Accuracy and Time Consumption of Each Model in Table 1
[0098] Model mAP Time Consumed (ms) SSD 45.4 61 Yolo3 57.9 51 R-FCN 51.9 85 The Method of This Application 65.2 49
[0099] Among them, mAP is an index to measure the recognition accuracy in object detection.
[0100] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.
[0101] Based on the same inventive concept, the embodiments of this application also provide an object detection device for implementing the above-mentioned object detection method. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the object detection device provided below can refer to the limitations on the object detection method in the above text, and will not be repeated here.
[0102] In one embodiment, as Figure 8 shown, an object detection device 800 is provided, including: an input module 801 and a feature extraction module 802.
[0103] Among them, the input module 801 is used to input the target image into the feature extraction network. The feature extraction network includes a cascaded downsampling network, a hierarchical network, an attention network, and a feature fusion network. The attention network includes multiple attention modules;
[0104] The feature extraction module 802 is configured to perform downsampling feature extraction on the target image through the downsampling network, perform hierarchical processing on the feature map output by the downsampling network through the hierarchical network, perform feature extraction on the multiple hierarchical feature maps output by the hierarchical network through the multiple attention modules, and perform fusion processing on the multiple feature maps output by the multiple attention modules and the feature map output by the downsampling network through the feature fusion network to obtain a target feature map, where the target feature map is used for object detection.
[0105] In one embodiment, the multiple attention modules are in one-to-one correspondence with the multiple hierarchical features Figure 1 The feature extraction module 802 is specifically configured to input each hierarchical feature map into the corresponding attention module; and perform feature extraction on the input hierarchical feature map through each attention module.
[0106] In one embodiment, the attention module includes a cascaded first extraction sub-network, a second extraction sub-network, and a third extraction sub-network. The feature extraction module 802 is specifically configured to, for each attention module, perform feature extraction on the input hierarchical feature map through the first extraction sub-network, perform feature extraction on the feature map output by the first extraction sub-network through the first feature extraction branch and the second feature extraction branch in the second extraction sub-network respectively to obtain two candidate feature maps, perform fusion processing on the two candidate feature maps through the second extraction sub-network, and perform feature extraction on the feature map output by the second extraction sub-network through the third extraction sub-network.
[0107] In one embodiment, the first extraction sub-network includes two asymmetric first convolutional layers and a first fusion layer. The feature extraction module 802 is specifically configured to perform feature extraction on the hierarchical feature map through the two first convolutional layers respectively, and perform fusion processing on the feature maps output by the two first convolutional layers through the first fusion layer.
[0108] In one embodiment, the first feature extraction branch includes a second convolutional layer, and the second feature extraction branch includes a second convolutional layer and a cascaded global pooling layer, a first fully-connected layer, an activation function layer, and a second fully-connected layer.
[0109] In one embodiment, the third extraction sub-network includes two asymmetric third convolutional layers and a second fusion layer. The feature extraction module 802 is specifically configured to perform feature extraction on the feature map output by the second extraction sub-network through the two third convolutional layers respectively, and perform fusion processing on the feature maps output by the two third convolutional layers through the second fusion layer.
[0110] In one embodiment, the attention module further includes a parameter normalization sub-network. The feature extraction module 802 is specifically configured to perform parameter normalization processing on the feature map output by the third extraction sub-network through the parameter normalization sub-network.
[0111] In one embodiment, the feature extraction module 802 is specifically configured to perform hierarchical processing on the feature map output by the downsampling network based on the number of channels through a hierarchical network; wherein, the sizes of the multiple hierarchical feature maps obtained after hierarchical processing are the same as the feature map output by the downsampling network, and the sum of the number of channels of the multiple hierarchical feature maps is equal to the number of channels of the feature map output by the downsampling network.
[0112] In one embodiment, the feature extraction module 802 is specifically configured to perform fusion processing on the multiple feature maps output by the multiple attention modules and the feature map output by the downsampling network through a feature fusion network, including: performing fusion processing on the multiple feature maps output by the multiple attention modules through the feature fusion network, and performing fusion processing again on the feature map obtained after the fusion processing and the feature map output by the downsampling network through the feature fusion network.
[0113] Each module in the above object detection device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0114] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 9 shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an object detection method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0115] Those skilled in the art can understand,Figure 9 The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0116] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0117] Input the target image into the feature extraction network. The feature extraction network includes a cascaded downsampling network, a hierarchical network, an attention network, and a feature fusion network. The attention network includes multiple attention modules; perform downsampling feature extraction on the target image through the downsampling network, perform hierarchical processing on the feature map output by the downsampling network through the hierarchical network, perform feature extraction on the multiple hierarchical feature maps output by the hierarchical network through the multiple attention modules, and perform fusion processing on the multiple feature maps output by the multiple attention modules and the feature map output by the downsampling network through the feature fusion network to obtain the target feature map, and the target feature map is used for target detection.
[0118] In one embodiment, when the processor executes the computer program, the following steps are further implemented: The multiple attention modules are in one-to-one correspondence with the multiple hierarchical features Figure 1 Perform feature extraction on the multiple hierarchical feature maps output by the hierarchical network through the multiple attention modules, including: input each hierarchical feature map into the corresponding attention module; perform feature extraction on the input hierarchical feature map through each attention module.
[0119] In one embodiment, when the processor executes the computer program, the following steps are further implemented: The attention module includes a cascaded first extraction sub-network, a second extraction sub-network, and a third extraction sub-network. Perform feature extraction on the input hierarchical feature map through each attention module, including: for each attention module, perform feature extraction on the input hierarchical feature map through the first extraction sub-network, perform feature extraction on the feature map output by the first extraction sub-network through the first feature extraction branch and the second feature extraction branch in the second extraction sub-network respectively to obtain two candidate feature maps, and perform fusion processing on the two candidate feature maps through the second extraction sub-network, and perform feature extraction on the feature map output by the second extraction sub-network through the third extraction sub-network.
[0120] In one embodiment, when the processor executes the computer program, the following steps are further implemented: The first extraction sub-network includes two asymmetric first convolutional layers and a first fusion layer. Feature extraction is performed on the input hierarchical feature map through the first extraction sub-network, including: Feature extraction is performed on the hierarchical feature map through the two first convolutional layers, and the feature maps output by the two first convolutional layers are fused through the first fusion layer.
[0121] In one embodiment, when the processor executes the computer program, the following steps are further implemented: The first feature extraction branch includes a second convolutional layer, and the second feature extraction branch includes a second convolutional layer and a cascaded global pooling layer, a first fully connected layer, an activation function layer, and a second fully connected layer.
[0122] In one embodiment, when the processor executes the computer program, the following steps are further implemented: The third extraction sub-network includes two asymmetric third convolutional layers and a second fusion layer. Feature extraction is performed on the feature map output by the second extraction sub-network through the third extraction sub-network, including: Feature extraction is performed on the feature map output by the second extraction sub-network through the two third convolutional layers, and the feature maps output by the two third convolutional layers are fused through the second fusion layer.
[0123] In one embodiment, when the processor executes the computer program, the following steps are further implemented: The attention module further includes a parameter normalization sub-network. After feature extraction is performed on the feature map output by the second extraction sub-network through the third extraction sub-network, the method further includes: Parameter normalization processing is performed on the feature map output by the third extraction sub-network through the parameter normalization sub-network.
[0124] In one embodiment, when the processor executes the computer program, the following steps are further implemented: Hierarchical processing is performed on the feature map output by the downsampling network through the hierarchical network, including: Hierarchical processing is performed on the feature map output by the downsampling network by the hierarchical network based on the number of channels; wherein, the sizes of the multiple hierarchical feature maps obtained after hierarchical processing are the same as the feature map output by the downsampling network, and the sum of the number of channels of the multiple hierarchical feature maps is equal to the number of channels of the feature map output by the downsampling network.
[0125] In one embodiment, when the processor executes the computer program, the following steps are further implemented: Fusion processing is performed on the multiple feature maps output by the multiple attention modules and the feature map output by the downsampling network through the feature fusion network, including: Fusion processing is performed on the multiple feature maps output by the multiple attention modules through the feature fusion network, and the feature map obtained after the fusion processing and the feature map output by the downsampling network are fused again through the feature fusion network.
[0126] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the object detection method provided in each of the above method embodiments is implemented.
[0127] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the object detection method provided in each of the above method embodiments is implemented.
[0128] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties.
[0129] Those of ordinary skill in the art can understand that all or part of the processes in the above-described method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the method embodiments as described above. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0130] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0131] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A target detection method, characterized in that, The method includes: Inputting a target image into a feature extraction network, where the feature extraction network includes a cascaded downsampling network, a hierarchical network, an attention network, and a feature fusion network, and the attention network includes a plurality of attention modules; the downsampling network includes a downsampling sub-network, and the downsampling sub-network includes a convolutional layer, an activation function layer, and a max pooling layer; Performing downsampling feature extraction on the target image through the downsampling network, performing hierarchical processing on the feature map output by the downsampling network through the hierarchical network, performing feature extraction on the multiple hierarchical feature maps output by the hierarchical network through the multiple attention modules, and performing fusion processing on the multiple feature maps output by the multiple attention modules and the feature map output by the downsampling network through the feature fusion network to obtain a target feature map, where the target feature map is used for target detection; The multiple attention modules correspond one-to-one to the multiple hierarchical feature maps, and the performing feature extraction on the multiple hierarchical feature maps output by the hierarchical network through the multiple attention modules includes: Inputting each of the hierarchical feature maps into the corresponding attention module; Performing feature extraction on the input hierarchical feature maps respectively through each of the attention modules.
2. The method according to claim 1, characterized in that, The attention module includes a cascaded first extraction sub-network, a second extraction sub-network, and a third extraction sub-network, and the performing feature extraction on the input hierarchical feature maps respectively through each of the attention modules includes: For each of the attention modules, performing feature extraction on the input hierarchical feature map through the first extraction sub-network, performing feature extraction on the feature map output by the first extraction sub-network respectively through a first feature extraction branch and a second feature extraction branch in the second extraction sub-network to obtain two candidate feature maps, performing fusion processing on the two candidate feature maps through the second extraction sub-network, and performing feature extraction on the feature map output by the second extraction sub-network through the third extraction sub-network.
3. The method according to claim 2, characterized in that The first extraction sub-network includes two asymmetric first convolutional layers and a first fusion layer, and the performing feature extraction on the input hierarchical feature map through the first extraction sub-network includes: Performing feature extraction on the hierarchical feature map respectively through the two first convolutional layers, and performing fusion processing on the feature maps output by the two first convolutional layers through the first fusion layer.
4. The method according to claim 2, characterized in that, The first feature extraction branch includes a second convolutional layer, and the second feature extraction branch includes the second convolutional layer and a cascaded global pooling layer, a first fully-connected layer, an activation function layer, and a second fully-connected layer.
5. The method according to claim 2, wherein The third extraction sub-network includes two asymmetric third convolutional layers and a second fusion layer, and the performing feature extraction on the feature map output by the second extraction sub-network through the third extraction sub-network includes: Performing feature extraction on the feature map output by the second extraction sub-network respectively through the two third convolutional layers, and performing fusion processing on the feature maps output by the two third convolutional layers through the second fusion layer.
6. The method according to claim 2, wherein The attention module further includes a parameter normalization sub-network. After the third extraction sub-network extracts features from the feature map output by the second extraction sub-network, the method further includes: Performing parameter normalization processing on the feature map output by the third extraction sub-network through the parameter normalization sub-network.
7. According to the method as claimed in any one of claims 1 to 6, characterized in that, The hierarchical processing of the feature map output by the downsampling network through the hierarchical network includes: Performing hierarchical processing on the feature map output by the downsampling network through the hierarchical network based on the number of channels; Wherein, the sizes of the multiple hierarchical feature maps obtained after hierarchical processing are the same as the feature map output by the downsampling network, and the sum of the number of channels of the multiple hierarchical feature maps is equal to the number of channels of the feature map output by the downsampling network.
8. The method according to any one of claims 1 to 6, characterized in that, The fusion processing of the multiple feature maps output by the multiple attention modules and the feature map output by the downsampling network through the feature fusion network includes: Fusing the multiple feature maps output by the multiple attention modules through the feature fusion network, and fusing the feature map obtained after the fusion processing and the feature map output by the downsampling network again through the feature fusion network.
9. A target detection device, characterized in that, The device includes: An input module, configured to input a target image into a feature extraction network, where the feature extraction network includes a cascaded downsampling network, a hierarchical network, an attention network, and a feature fusion network, and the attention network includes multiple attention modules; the downsampling network includes a downsampling sub-network, and the downsampling sub-network includes a convolutional layer, an activation function layer, and a max pooling layer; A feature extraction module, configured to perform downsampling feature extraction on the target image through the downsampling network, perform hierarchical processing on the feature map output by the downsampling network through the hierarchical network, perform feature extraction on the multiple hierarchical feature maps output by the hierarchical network through the multiple attention modules, and perform fusion processing on the multiple feature maps output by the multiple attention modules and the feature map output by the downsampling network through the feature fusion network to obtain a target feature map, where the target feature map is used for target detection; The multiple attention modules correspond to the multiple hierarchical feature maps one by one. The feature extraction module includes: inputting each of the hierarchical feature maps into the corresponding attention module; respectively performing feature extraction on the input hierarchical feature maps through each of the attention modules.
10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Target detection method and moving target tracking method using same
CN114092820A