A multi-task multi-branch attention network structure

By introducing a multi-branch attention mechanism into the multi-task neural network, the problems of negative migration and inefficiency of existing multi-task networks when handling multi-tasks are solved, and higher accuracy and detection speed are achieved.

CN114897149BActive Publication Date: 2025-06-10山西清众科技股份有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210705174.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-21
Publication Date
2025-06-10
Estimated Expiration
2042-06-21

AI Technical Summary

Technical Problem

When existing multitasking neural networks process multiple tasks in a picture, there are negative migration problems and problems such as large calculations and slow detection speed.

Method used

A multi-task multi-branch attention network structure is adopted, including a multi-branch feature extraction network, an attention-based feature selection module and an attention-based multi-branch prediction module. Through the channel attention and branch attention modules, the weighting of feature maps and prediction results are integrated to improve the accuracy of the network.

Benefits of technology

Information sharing and integration between different tasks and different branches is realized, the accuracy and detection speed of the network are improved, and negative migration problems are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114897149B_ABST
    Figure CN114897149B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-task multi-branch attention network structure, belonging to the technical field of deep learning. It solves the problem that the existing multi-task network structure enables the model to process multiple tasks simultaneously, but at the same time, due to the different features required for each task, it also brings the problem of network negative transfer. The technical solution adopted to solve the above technical problems is as follows: This structure has three modules: a multi-branch feature extraction network, an attention-based feature selection module, and an attention-based multi-branch prediction module. The multi-branch feature extraction and output network can extract feature maps of pictures at different stages of the network. The feature selection module is used to weight the feature maps to provide task-related feature maps for each task, while the attention-based multi-branch prediction module can integrate the prediction results of the same task from different branches, thereby improving the accuracy of the network. The present invention is applied to image classification processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention provides a multi-task multi-branch attention network structure, belonging to the technical field of deep learning in computer technology. Background Art

[0002] A convolutional neural network is a type of neural network that includes convolution, pooling, activation function calculations, and has a certain depth structure, and is one of the representative algorithms in the field of deep learning. At present, it has been confirmed by a large number of research examples that it has strong performance in the fields of object classification, localization, and detection, and has made breakthrough progress in the field of object classification with multi-level feature learning and rich feature expression capabilities.

[0003] In recent years, one classic network model after another has emerged in the field of image classification. GoogLeNet proposed the Inception module, which widened the network, thereby enhancing the network's feature extraction ability in another direction; ResNet proposed the residual structure, effectively alleviating problems such as gradient disappearance and network degradation, so that the neural network can reach a depth of hundreds of layers; DensenNet proposed the concept of dense connection, connecting the input of each layer before the network to all subsequent layers, thereby achieving efficient feature reuse and also slowing down problems such as gradient disappearance. In the field of object classification, many scholars have applied the above models to image classification, but most of the models they constructed are single-task models, which can only perform one task at a time. In real life, a picture usually needs to simultaneously perform multiple task judgments. In this case, it is usually necessary to train multiple single-task models to perform classification calculations for each task, resulting in a large amount of computation and slow detection speed.

[0004] Currently, for the situation where there are multiple tasks in a picture, there are already various multi-task neural network frameworks. Multi-task learning (MTL) is a comprehensive learning method that is achieved by simultaneously training several tasks and sharing some parameters among the tasks. In a multi-task network, multiple tasks share a structure, and information from different tasks can be utilized. When the losses of all tasks tend to flatten, this structure is equivalent to integrating the information of all tasks. Generally speaking, a multi-task network has stronger generalization ability than a single-task network. However, most multi-task networks require a strong connection between different tasks, otherwise it will cause negative transfer of the multi-task network. In practical applications, its accuracy is not high and there are many limitations. Summary of the Invention

[0005] In order to overcome the deficiencies in the prior art, the technical problem to be solved by the present invention is to provide an improvement to the multi-task multi-branch attention network structure.

[0006] To solve the above technical problems, the technical solution adopted by the present invention is: a multi-task multi-branch attention network structure, including the following modules:

[0007] Multi-branch feature extraction network: used to extract features from the preprocessed image, and divide the network into multiple branches, and each branch outputs the feature map extracted by the feature extraction network at the current stage;

[0008] Attention-based feature selection module: uses channel attention to perform a weighting operation on the channels of the feature maps output by each branch feature extraction network, and generates task-related feature maps for each task;

[0009] Attention-based multi-branch prediction module: inputs each weighted feature map by channel attention into a fully connected layer for multi-task prediction, and then uses a branch attention module and an intra-task attention module to integrate the prediction results of the same task of different branches. Finally, the classification results of each task are output.

[0010] The feature extraction network uses the modified ResNet50 as the backbone network. The modified ResNet50 includes a total of 1 input part and 4 Blocks. The input part consists of 3 cascaded 3×3 convolutional layers. The Blocks are composed of 3, 4, 6, and 3 Layers respectively. Each Layer includes a 1×1 convolutional layer, a batch normalization layer, a rectified linear unit layer, a 3×3 convolutional layer, a batch normalization layer, a rectified linear unit layer, a 1×1 convolutional layer, a batch normalization layer, and a rectified linear unit layer.

[0011] The attention-based feature selection module is represented by the following formula:

[0012] ;

[0013] In the above formula: f i represents the original feature map output by the i-th branch of the feature extraction network; AvgPool and MaxPool represent pooling operations, and MLP represents a multi-layer perceptron with one hidden layer. represents the sigmoid function. represents the specific task feature map generated by the attention module on the i-th branch.

[0014] The branch attention module assigns weights to each branch by generating an attention mask, and assigns different weights to different branches.

[0015] The calculation steps for the branch attention module to assign different weights to different branches are as follows:

[0016] The prediction results of each branch are concatenated together, and the concatenated result is presented in the form of a multi-channel feature map F;

[0017] Perform a 1×n convolution operation on the graph F to obtain the compressed information of multiple branches;

[0018] The obtained compressed information is fed into a fully connected layer, and then the output of the fully connected layer is fed into an activation function to obtain an attention mask W 1 ;

[0019] Attention mask W 1 Multiply with the graph F to obtain a weighted graph F 1 .

[0020] The in-task attention module is based on the branch attention module and generates different weights for the prediction results of the same subclass on different branches.

[0021] The steps for the in-task attention module to calculate the weights are as follows:

[0022] The multi-channel graph output by the branch attention module F 1 Is converted into a single-channel graph F 2 ;

[0023] Perform a k×1 convolution operation on the single-channel graph F 2 To obtain an attention mask W 2 ;

[0024] Attention mask W 2 Multiply with the single-channel graph To obtain a weighted graph R, and the final classification result of the task is obtained by adding the elements of the weighted graph R column by column.

[0025] The beneficial effects of the present invention compared with the prior art are as follows: The multi-task multi-branch attention network structure provided by the present invention has three modules: a multi-branch feature extraction network, an attention-based feature selection module, and an attention-based multi-branch prediction module. The multi-branch feature extraction and output network can extract the feature maps of the picture at different stages of the network. The feature selection module is used to weight the feature maps to provide task-related feature maps for each task, while the attention-based multi-branch prediction module can integrate the prediction results of the same task from different branches, thereby improving the accuracy of the network. Experiments show that this network structure can utilize the feature information at different stages of the network compared with the single-task network, and different tasks can also promote each other, and its effect is better than that of the single-task network. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The present invention will be further described below with reference to the accompanying drawings:

[0027] Figure 1 It is a schematic diagram of the overall network structure of the present invention;

[0028] Figure 2 It is a schematic diagram of the internal structure of the attention feature selection module based on the present invention;

[0029] Figure 3 It is a schematic diagram of the internal structure of the multi-branch prediction module based on attention of the basic invention. DETAILED DESCRIPTION OF THE INVENTION

[0030] As Figures 1 to 3 shown, the problems to be solved by the present invention are: 1. Multi-task networks usually provide the same feature map for all tasks, resulting in the problem of negative transfer in multi-task networks. The reasons for negative transfer may include: (1) Different tasks require feature maps at different stages. Some tasks require low-level image features, while some tasks require high-level image features. (2) Different tasks also focus on different regions in the same image. 2. There is also the problem of how to allocate appropriate weights to the prediction results of each branch in the multi-branch structure. For a specific task, it has different prediction results on different branches. How to allow the network to reasonably utilize the prediction results of different branches in the multi-task network, suppress the branches with low prediction accuracy, and obtain better classification results is the problem to be solved by the present invention.

[0031] The technical solutions adopted by the present invention to solve its technical problems include the following three modules, as Figure 1 shown:

[0032] Module 1: Multi-branch feature extraction network

[0033] To solve the problem that different tasks require feature maps at different stages, the present invention proposes a multi-branch feature extraction network. As Figure 2 shown, the feature extraction network extracts features from the preprocessed image and divides the network into five branches, consisting of branches 1-5. Each branch outputs the feature map extracted by the feature extraction network at this stage. Such a multi-branch structure can output the features extracted by the network at different stages to meet the needs of different tasks.

[0034] The feature extraction network uses a modified ResNet50 as the backbone network. The modified ResNet50 consists of 1 input part and 4 Blocks. The input part is composed of 3 consecutive 3×3 convolutional layers. The Blocks are composed of 3, 4, 6, and 3 Layers respectively. Each Layer contains a 1×1 convolutional layer, a batch normalization layer (Batch Normalization, BN), a rectified linear unit layer (Rectified Linear Unit, ReLu), a 3×3 convolutional layer, a batch normalization layer, a rectified linear unit layer, a 1×1 convolutional layer, a batch normalization layer, and a rectified linear unit layer.

[0035] The specific structure of the network is shown in Table 1 below:

[0036]

[0037] Table 1 Feature extraction network structure.

[0038] Module 2: Attention-based feature selection module

[0039] This module uses channel attention to perform a weighting operation on the channels of the feature maps output by each branch, generating task-related feature maps for each task to improve the accuracy of the network.

[0040] The attention-based feature selection module can be represented by the following formula:

[0041] ;

[0042] In the above formula: f i represents the original feature map output by the i-th branch of the feature extraction network; AvgPool and MaxPool represent pooling operations, and MLP represents a multi-layer perceptron with one hidden layer. represents the sigmoid function. represents the task-specific feature map generated by the attention module on the i-th branch.

[0043] Module 3: Attention-based multi-branch prediction module

[0044] As Figure 3 shown, this module inputs each weighted feature map with channel attention into a fully connected layer for multi-task prediction, and then uses a branch attention module and an intra-task attention module to integrate the prediction results of the same task on different branches. Finally, this module outputs the classification results of each task.

[0045] Branch Attention Module: The branch attention module assigns weights to each branch by generating an attention mask. By assigning different weights to different branches, the network can reasonably utilize the prediction results of different branches, so as to pay more attention to the branches with better classification results, thereby improving the classification performance of the model. It is calculated through the following steps:

[0046] (1) The prediction results of each branch are concatenated together, and the concatenated result is presented in the form of a feature map F with 5 channels.

[0047] (2) A 1×n convolution operation is performed on the graph F. This operation compresses the information of each branch and realizes the overall evaluation of the prediction results of each branch.

[0048] (3) The obtained compressed information is fed into the fully connected (FC) layer, and then the output of the FC layer is fed into the activation function to obtain the attention mask W 1 .

[0049] (4) The attention mask W 1 is multiplied by the graph F to obtain the weighted graph F 1 .

[0050] Intra-task Attention Module: The branch attention module only evaluates the overall performance of the network in classifying on that branch, ignoring the accuracy of the sub-categories of the task on that branch. To solve this problem, the present invention designs an intra-task attention module. Based on the branch attention module, different weights are generated for the prediction results of the same sub-category on different branches. This can improve the accuracy of the model. It can be calculated through the following steps:

[0051] (1) The multi-channel graph F 1 output by the branch attention module is converted into a single-channel graph F 2 .

[0052] (2) To enable the network to analyze the prediction results of the same sub-category on different branches, a k×1 convolution operation is performed on the graph F 2 to obtain the attention mask W 2 .

[0053] (3) The attention mask W 2 is multiplied by the graph F 2 to obtain the weighted graph R, and the final classification result of this task is obtained by adding the elements of the weighted graph R column by column.

[0054] Regarding the specific structure of the present invention, it should be noted that the connection relationships between the various component modules adopted by the present invention are determined and achievable. Except for the special descriptions in the embodiments, the specific connection relationships can bring corresponding technical effects and, on the premise of not relying on the execution of corresponding software programs, solve the technical problems proposed by the present invention. The models and connection methods of the components, modules, and specific components in the present invention, except for the specific descriptions, all belong to the prior art such as publicly available patents, publicly available journal papers, or common general knowledge that those skilled in the art can obtain before the filing date, and need not be elaborated. This makes the technical solution provided in this case clear, complete, and achievable, and can reproduce or obtain the corresponding physical product according to this technical means.

[0055] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A multi-task multi-branch attention network structure, Characterized in that: It includes the following modules: Multi-branch feature extraction network: used to extract features from the preprocessed image, and divide the network into multiple branches, and each branch outputs the feature map extracted by the feature extraction network at the current stage; Attention-based feature selection module: uses channel attention to perform a weighting operation on the channels of the feature maps output by each branch feature extraction network to generate task-related feature maps for each task; Attention-based multi-branch prediction module: inputs each weighted feature map with channel attention into a fully connected layer for multi-task prediction, and then uses a branch attention module and an intra-task attention module to integrate the prediction results of the same task on different branches. Finally, the classification results of each task are output; The branch attention module assigns weights to each branch by generating an attention mask, and assigns different weights to different branches; The calculation steps for the branch attention module to assign different weights to different branches are as follows: The prediction results of each branch are concatenated together, and the concatenated result is presented in the form of a multi-channel feature map F; Perform a 1×n convolution operation on the graph F to obtain the compressed information of multiple branches; The obtained compressed information is fed into the fully connected layer, and then the output of the fully connected layer is fed into the activation function to obtain the attention mask W 1 ; Attention mask W 1 Multiply with Figure F to obtain a weighted graph F 1 .

2. A multi-task multi-branch attention network structure according to claim 1, Characterized in that: The feature extraction network uses the modified ResNet50 as the backbone network. The modified ResNet50 contains a total of 1 input part and 4 Blocks. The input part consists of 3 cascaded 3×3 convolutional layers. The Blocks are composed of 3, 4, 6, and 3 Layers respectively. Each Layer contains a 1×1 convolutional layer, a batch normalization layer, a rectified linear unit layer, a 3×3 convolutional layer, a batch normalization layer, a rectified linear unit layer, a 1×1 convolutional layer, a batch normalization layer, and a rectified linear unit layer.

3. A multi-task multi-branch attention network structure according to claim 1, Characterized in that: The attention-based feature selection module is represented by the following formula: ; In the above formula: f i represents the original feature map output by the i-th branch of the feature extraction network; AvgPool and MaxPool represent pooling operations, and MLP represents a multi-layer perceptron with one hidden layer, represents the sigmoid function, represents the task-specific feature map generated by the attention module on the i-th branch.

4. A multi-task multi-branch attention network structure according to claim 1, Characterized in that: The intra-task attention module is based on the branch attention module and generates different weights for the prediction results of the same subclass on different branches.

5. A multi-task multi-branch attention network structure according to claim 4, Characterized in that: The calculation steps for the intra-task attention module to calculate weights are as follows: Multi-channel graph output by the branch attention module F 1 is converted into a single-channel graph F 2 ; For a single-channel image F 2 perform a k×1 convolution operation to obtain an attention mask W 2 ; Attention mask W 2 Multiply with the single-channel graph to obtain the weighted graph R, and the final classification result of the task is obtained by adding the elements of the weighted graph R column by column.

Citation Information

Patent Citations

  • Face attribute recognition method and device, electronic equipment and storage medium

    CN111339813A