An Image Recognition Method Based on a Fusion Attention Mechanism
By combining the fusion channel and spatial attention mechanism with the two-layer long and short-term memory network, the problem of insufficient feature extraction in image description is solved, more accurate image description results are generated, and the performance of the model is improved.
Patent Information
- Application Number
- CN202310120205.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-15
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-02-15
AI Technical Summary
In the existing image description methods, the image feature extraction is insufficient, resulting in the description generated by the decoder inaccurate enough, and the single attention mechanism and LSTM model fail to effectively combine image feature information, resulting in low correlation of generated sentence words.
The fusion attention mechanism is adopted, and the initial feature map is weighted in combination with channel and spatial attention mechanisms, and image description is performed through two-layer long and short-term memory networks and multi-head attention mechanisms to generate more accurate description results.
It improves the overall performance of the image description model, the generated description is more accurate and vivid, and improves the utilization of image information.
Smart Images

Figure CN116229234B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of computer vision and natural speech processing, and particularly relates to an image recognition method based on a fusion attention mechanism. Background Art
[0002] Many traditional fields such as computer vision have been greatly improved. However, the development of a series of emerging technologies including visual navigation and virtual reality has put forward higher requirements for computer vision, especially for image description. These technologies hope to obtain richer and more comprehensive image information. Therefore, more and more researchers have started to study computer vision, and among them, the research on image description has gradually increased. Image description involves multi-level utilization of image information. Object detection in images, the relationships between objects, and the construction of scene graphs all fall within the scope of image description research. Some progress has been made in tasks such as object detection in image description, but it still difficult to meet the requirements of our actual applications. Moreover, the research on tasks such as image description and scene graph construction is still lacking. These tasks represent a deeper understanding of images and are also more core issues in image content understanding. Therefore, generally speaking, each task in image description has great research value and great practical application value.
[0003] Traditional image description methods have problems such as being too rigid and lacking flexibility, which greatly affect their actual application effects. As deep learning has been gradually applied to other fields and, due to its fast computing power, it can obtain the most valuable information for specific tasks with the support of big data. Specifically in the field of computer vision, it can compress an image into a feature vector containing a large amount of information and continuously optimize the effect of information extraction for different tasks using a large amount of data. Such characteristics are very important for image content understanding. It can obtain a lot of potential information in images for specific tasks, greatly improve the utilization degree of image information, and thus achieve better actual effects. Therefore, the method based on deep learning is the mainstream method for current image description tasks. With the continuous in-depth research on deep learning, many new models and methods can further improve the effects of various image description tasks, can improve the utilization degree of image information from multiple aspects, promote the development of the computer vision field, and thus play a huge role in the construction of the future intelligent society.
[0004] Currently, most image descriptions adopt the encoder-decoder framework based on deep learning as the basic framework. At the same time, the attention mechanism is also widely applied to related networks and has achieved good results, improving the performance of the image description model. However, for most encoders, they simply use convolutional neural networks or a single attention mechanism to assist in extracting image features. These methods cannot fully extract and utilize image features, resulting in insufficient image information obtained by the decoder and inaccurate generated description sentences. For the decoder that generates descriptions, some models do not fully analyze the correlation between the image feature information extracted by the encoder and the information of the long short-term memory network (LSTM). For a single LSTM, the generation of sentences is predicted by the hidden state of the LSTM. If the feature information cannot be well combined, the generated words will not be accurate and clear enough. Ultimately, the correlation between the words in the predicted sentences is not high enough to achieve the effect of high-quality descriptions. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention proposes an image recognition method based on a fusion attention mechanism. The method includes: obtaining an image to be recognized, inputting the image to be recognized into a trained image description model to obtain an image description result; recognizing the image according to the image description result to obtain an image recognition result;
[0006] The process of training the image description model includes:
[0007] S1: Obtain the MSCOCO image dataset and preprocess the images in the image dataset;
[0008] S2: Input the preprocessed image into the Resnet101 network for feature extraction to obtain an initial feature map;
[0009] S3: Respectively adopt the channel attention mechanism and the spatial attention mechanism to perform weighted processing on the initial feature map, and perform parallel fusion processing on the channel attention feature and the spatial attention feature to obtain a fusion feature map;
[0010] S4: Use a two-layer long short-term memory network to recognize and decode the fusion feature map to obtain an image description result;
[0011] S5: Calculate the loss function of the model according to the recognition result;
[0012] S6: Adopt a reinforcement learning loss strategy to optimize the parameters of the model, and complete the training of the model when the loss function is minimized.
[0013] Preferably, the process of processing the initial features using the channel attention mechanism includes: performing maximum pooling and average pooling on the initial features respectively to obtain the maximum feature and the average feature of the image; inputting the maximum feature and the average feature into a multi-layer perceptron for dimensionality reduction processing, aggregating the dimensionality-reduced maximum feature and average feature, and activating them through an activation function to obtain the channel attention feature.
[0014] Preferably, the process of processing the initial features using the spatial attention mechanism includes: inputting the initial image features into a multi-layer perceptron to extract feature weights, and fusing the information on each channel of the extracted feature weights through a batch normalization layer and an average pooling layer to obtain the spatial position attention weights; calculating the spatial attention feature of the image according to the spatial position attention weights.
[0015] Preferably, the formula for parallel fusion processing of the channel attention feature and the spatial attention feature is:
[0016]
[0017] where F represents the initial input feature, F C (F) represents the channel attention feature, F S (F) represents the spatial attention feature, λ C and λ S are two hyperparameters, represents the feature after the fusion of spatial attention and channel attention.
[0018] Preferably, the process of using a two-layer long short-term memory network to identify and decode the fused feature map includes: combining a two-layer long short-term memory network with a multi-head attention mechanism to form a decoder; using the image features extracted by the encoder as the query matrix, and the output of the first long short-term memory network as the key matrix and value matrix and inputting them into the multi-head dot product attention module for attention fusion; inputting the attention image features and the hidden state of the previous moment together into the second long short-term memory network, calculating the word distribution probability on the vocabulary, and obtaining a word sequence according to the word distribution probability; generating an image description result according to the word sequence.
[0019] Preferably, the expression of the loss function of the model is:
[0020]
[0021] where L XE (θ) represents the cross-entropy loss, θ represents the learnable parameters of the model, T represents the length of the word embedding vector, p θ represents the model probability distribution, represents the true value, represents the true description sequence.
[0022] Advantages of the present invention:
[0023] The present invention combines channel attention and spatial attention to assist in extracting image features, enabling the feature map to contain more image information; the present invention proposes to use a two-layer long short-term memory network to fuse the multi-head attention mechanism to solve the problem of inaccurate feature decoding, improving the accuracy of generated words and enhancing the overall performance of the image description model. Description of the drawings
[0024] Figure 1 It is a structural diagram of the channel attention module and the spatial attention module of the present invention;
[0025] Figure 2 It is a structural diagram of the decoder of the present invention;
[0026] Figure 3 It is a flowchart of image recognition based on the fusion attention mechanism of the present invention. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] An image recognition method based on a fusion attention mechanism, the method includes: obtaining an image to be recognized, inputting the image to be recognized into a trained image description model to obtain an image description result; recognizing the image according to the image description result to obtain an image recognition result.
[0029] The process of training the image description model includes:
[0030] S1: Obtain the MSCOCO image data set and preprocess the images in the image data set;
[0031] S2: Input the preprocessed image into the Resnet101 network for feature extraction to obtain an initial feature map;
[0032] S3: Respectively use the channel attention mechanism and the spatial attention mechanism to perform weighted processing on the initial feature map, and perform parallel fusion processing on the channel attention feature and the spatial attention feature to obtain a fusion feature map;
[0033] S4: Use a two-layer long short-term memory network to recognize and decode the fusion feature map to obtain an image description result;
[0034] S5: Calculate the loss function of the model according to the recognition result;
[0035] S6: Optimize the parameters of the model using the reinforcement learning loss strategy, and complete the training of the model when the loss function is minimized.
[0036] The image encoder integrating channel and spatial attention mechanisms is used to enhance the feature information of the image, which can make the extracted features more abundant and contain as much semantic information of the image as possible. At the same time, a language decoding model combined with a dot product attention module is proposed to further fuse the image feature information with the language information in the attention module to improve the prediction of the image description and obtain a more vivid and accurate image description result.
[0037] The present invention is based on the combination of channel and spatial attention mechanisms, and embeds a fusion attention module into the Resnet101 network to assist in extracting image features. The channel and spatial attention mechanisms are respectively adopted for the initial image features to achieve the purpose of aggregating the initial image features.
[0038] A specific implementation manner of an image description method integrating attention mechanisms is as Figure 3 shown. The method includes: extracting the initial features of the image according to the input image; using channel attention and spatial attention to help understand more semantic information of the image content, and merging the image features integrating channels and spaces in a parallel manner to obtain fused image features; using a two-layer long short-term memory network combined with a multi-head attention mechanism to form a decoder; taking the image features extracted by the encoder as the query matrix, and taking the output of the first long short-term memory network as the key matrix and value matrix and inputting them into the multi-head dot product attention module for attention fusion; inputting the attention image features and the hidden state of the previous moment together into the second long short-term memory network, calculating the word distribution probability on the vocabulary, obtaining a word sequence, and generating a description.
[0039] In the design of the channel attention module in this embodiment, as Figure 1 shown, this module uses two pooling methods to process the initial features to obtain the channel attention feature weights of the image. Among them, average pooling can aggregate the spatial information of the feature map, and max pooling can obtain other important information of the object.
[0040] Specifically, first use max pooling and average pooling to obtain the max feature and average feature of the image respectively, and then input these two features into a multilayer perceptron (MLP). First reduce the dimension of the features, and then increase the dimension of the features to better aggregate the features. Then aggregate the two features by element-wise summation and activate them using an activation function to obtain the channel attention features. The calculation formula is:
[0041] F C (F) = σ(MLP(Max(F) + MLP(Avg(F)))
[0042] Among them, represents the initial input feature, σ represents the Sigmoid activation function, Max and Avg respectively represent max pooling and average pooling, MLP represents a multi-layer perceptron, and F C (F) is the channel attention feature.
[0043] Compared with channel attention, spatial attention focuses more on the spatial information of the image, the information of where the object is. This helps to describe the spatial relationship of the objects in the image. Therefore, the present invention designs a spatial attention network to extract the spatial feature information of the image. First, the initial image feature is input into a multi-layer perceptron (MLP) to refine the feature weights, and then the information on each channel is fused through a batch normalization layer and an average pooling layer to obtain the spatial position attention weights. The calculation method is as follows:
[0044] F S (F) = Avg(BN(MLP(F)))
[0045] Among them, F represents the initial input feature, MLP represents a multi-layer perceptron, BN represents batch normalization, Avg represents average pooling, and F S (F) represents the spatial attention feature.
[0046] After obtaining the channel attention feature and the spatial attention feature respectively, the present invention uses a parallel method to combine these two features as the final input image feature to the language decoder. The fusion method is:
[0047]
[0048] Among them, F represents the initial input feature, F C (F) represents the channel attention feature, F S (F) represents the spatial attention feature, λ C and λ S are two hyperparameters, represents the feature after the fusion of spatial attention and channel attention.
[0049] Optionally, the two hyperparameters are respectively taken as 0.5.
[0050] The language model of the present invention uses a two-layer long short-term memory network as the language generator. And in order to make the generated language description more accurate and reasonable, the present invention adds a multi-head attention (Multi HeadAttention) module to the language model, which can effectively improve the quality of the generated description.
[0051] As shown Figure 2 in the figure, this is the decoder structure of the present invention. The language model consists of two LSTMs and an attention module. Different from the ordinary model that uses only LSTM alone, after adding the attention module, it can improve the sensitivity of the LSTM unit to information, can amplify the weight of important information, and is beneficial to the accuracy of generating descriptions. LSTM is superior to RNN in processing long sequence information and can solve the problems of gradient disappearance and gradient explosion to a certain extent. The LSTM unit mainly includes four modules, namely the forget gate f t , the input gate i t , the output gate o t and the cell state c t ; The formula for the long short-term memory network to calculate the input data is:
[0052] f t =σ(W fh h t-1 +W fx x t +b f )
[0053] i t =σ(W ih h t-1 +W ix x t +b i )
[0054] o t =σ(W oh h t-1 +W ox x t +b o )
[0055] g t =tanh(W gh h t-1 +W gx x t +b g )
[0056] c t =f t ⊙c t-1 +i t ⊙g t
[0057] h t =o t ⊙tanh(c t )
[0058] Among them, σ represents the Sigmoid activation function, g tCandidate vector representing the cell state, h t Hidden state representing the current moment, x t Input to the LSTM at the current moment, W fh 、W fx 、W ih 、W ix 、W oh 、W ox 、W gh And W gx Are all learnable weight matrices, b f 、b i 、b o And b g Are all bias vectors, and ⊙ represents element-wise multiplication.
[0059] The present invention uses the image features extracted by the encoder as the query matrix, and the output of the first LSTM as the key matrix and value matrix, which are input into the multi-head dot product attention module for attention fusion. The multi-head dot product attention divides each Q, K, V into H = 8 channels to calculate new attention features. The calculation method of the multi-head dot product attention is as follows:
[0060] f mh-att (Q, K, V) = Concat(head1, …, head H )
[0061] head i = f dot-att (Q i , K i , V i )
[0062]
[0063] Q = v t
[0064]
[0065]
[0066]
[0067] Among them, f dot-att Represents dot product attention, Q is the query matrix, K is the key matrix, V is the value matrix, Is the average image feature, v t Is the encoder image feature, Is the attention image feature, Is the hidden state of the previous moment, w e Is the embedding matrix of the dictionary Σ, П t Is the one-hot encoding at time t.
[0068] The attention image features and the hidden state at the previous moment are input into the second LSTM together, and the word distribution probability on the vocabulary is calculated to obtain a word sequence: y = {y1, y2, …, y t}, and the formula for the second long short-term memory network to process the data is:
[0069]
[0070]
[0071] Among them, and are the hidden state of the first LSTM and the hidden state of the second LSTM respectively, is the attention image feature, y 1:t-1 is the word sequence, W y is the weight, b y is the bias.
[0072] During the training process of the present invention, cross-entropy loss optimization training is first used. For the true label sequence The present invention trains the model by optimizing the cross-entropy (XE):
[0073]
[0074] Among them, L XE (θ) represents the cross-entropy loss, θ represents the model learnable parameters, T represents the length of the word embedding vector, p θ represents the model probability distribution, represents the true value, represents the true description sequence.
[0075] After the cross-entropy training, the present invention then uses the reinforcement learning loss to further train and optimize the model, and the calculation method is as follows:
[0076]
[0077] Among them, r is the scoring function of the evaluation index, and the CIDEr evaluation index is used here. Its gradient can be approximated as:
[0078]
[0079] Among them, represents a sampling result, represents the result of greedy decoding, represents the gradient, L R(θ) represents the reinforcement learning function, θ represents the model's learnable parameters, r(.) represents the reward function, and p θ represents the model probability distribution.
[0080] After cross-entropy optimization and reinforcement learning optimization, an image description model with the best accuracy for describing the input image is obtained.
[0081] The above examples have further elaborated on the purpose, technical solution, and advantages of the present invention. It should be understood that the above examples are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An image recognition method based on a fusion attention mechanism, characterized in that, Including: Obtain the image to be recognized, input the image to be recognized into the trained image description model, and obtain the image description result; Recognize the image according to the image description result to obtain the image recognition result; The process of training the image description model includes: S1: Obtain the MSCOCO image dataset and preprocess the images in the image dataset; S2: Input the preprocessed image into the Resnet101 network for feature extraction to obtain the initial feature map; S3: Respectively use the channel attention mechanism and the spatial attention mechanism to perform weighted processing on the initial feature map, and perform parallel fusion processing on the channel attention feature and the spatial attention feature to obtain the fusion feature map; S4: Use a two-layer long short-term memory network to perform recognition and decoding on the fusion feature map to obtain the image description result; specifically including: the two-layer long short-term memory network combines the multi-head attention mechanism to form a decoder; use the image features extracted by the encoder as the query matrix, and the output of the first long short-term memory network as the key matrix and value matrix to input into the multi-head dot product attention module for attention fusion; input the attention image features and the hidden state of the previous moment into the second long short-term memory network together, calculate the word distribution probability on the vocabulary, and obtain a word sequence according to the word distribution probability; generate the image description result according to the word sequence; S5: Calculate the loss function of the model according to the recognition result; S6: Use the reinforcement learning loss strategy to optimize the parameters of the model, and complete the training of the model when the loss function is the smallest.
2. The image recognition method based on a fusion attention mechanism according to claim 1, wherein The process of using the channel attention mechanism to process the initial feature includes: using max pooling and average pooling to process the initial feature respectively to obtain the maximum feature and average feature of the image; input the maximum feature and average feature into the multi-layer perceptron for dimensionality reduction processing, aggregate the dimensionality-reduced maximum feature and average feature, and activate through the activation function to obtain the channel attention feature.
3. A method for image recognition based on a fusion attention mechanism according to claim 1, characterized in that The process of using the spatial attention mechanism to process the initial feature includes: input the initial image feature into the multi-layer perceptron to extract the feature weight, and fuse the information on each channel of the extracted feature weight through the batch normalization layer and the average pooling layer to obtain the spatial position attention weight; calculate the spatial attention feature of the image according to the spatial position attention weight.
4. A method for image recognition based on a fusion attention mechanism according to claim 1, characterized in that, The formula for parallel fusion processing of the channel attention feature and the spatial attention feature is: Among them, F represents the initial input feature, F C (F) represents the channel attention feature, F S (F) represents the spatial attention feature, λ C and λ S are two hyperparameters, represents the feature after the fusion of spatial attention and channel attention.
5. An image recognition method based on a fusion attention mechanism according to claim 1, characterized in that The long short-term memory network consists of four modules, namely the forget gate f t , the input gate i t , the output gate o t and the cell state c t ; The formula for the long short-term memory network to calculate the input data is: f t = σ(W fh h t-1 + W fx x t + b f ) i t = σ(W ih h t-1 + W ix x t + b i ) o t = σ(W ox x t-1 + W ox x t + b o ) g t = tanh(W gh h t-1 + W gx x t + b g ) c t = f t ⊙ c t-1 + i t ⊙ g t h t = o t ⊙tanh(c t ) Among them, σ represents the Sigmoid activation function, g t represents the candidate vector of the cell state, h t represents the hidden state at the current moment, x t represents the input of the LSTM at the current moment, W fh 、W fx 、W ih 、W ix 、W oh 、W ox 、W gh and W gx are all learnable weight matrices, b f 、b i 、b o and b g are all bias vectors, and ⊙ represents element-wise multiplication.
6. The image recognition method based on a fusion attention mechanism according to claim 1, characterized in that The calculation formula for attention fusion is: f mh-att (Q, K, V) = Concat(head1, …, head H ) head i = f dot-att (Q i , K i , V i ) Q = v t Among them, f dot-att represents the dot product. Note that Q is the query matrix, k is the key matrix, and V is the value matrix. is the average image feature, v t is the encoder image feature, is the attention image feature, is the hidden state at the previous moment, w e is the embedding matrix of the vocabulary Σ, П t is the one-hot encoding at time t.
7. A method for image recognition based on a fusion attention mechanism according to claim 1, characterized in that The formula for the second long short-term memory network to process the data is: Among them, and are the hidden states of the first LSTM and the second LSTM respectively, is the attention image feature, y 1:t-1 is the word sequence, W y is the weight, b y is the bias.
8. A method for image recognition based on a fusion attention mechanism according to claim 1, characterized in that, The expression of the loss function of the model is: Among them, L XE (θ) represents the cross-entropy loss, θ represents the model's learnable parameters, T represents the length of the word embedding vector, p θ represents the model's probability distribution, represents the true value, represents the true description sequence.