A Three-Channel Residual Attention Image Captioning Method
By introducing the residual attention module of three-path residual attention path and relative position in the attention mechanism model, the problem of poor connection between attention scores in different layers is solved, and a better image description effect is achieved.
Patent Information
- Application Number
- CN202210680166.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-15
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-06-15
AI Technical Summary
In the existing attention mechanism model, the connection between attention scores of different layers is not strong enough, resulting in the grid characteristics of the image losing geometric position information when entering the Transformer model.
A three-way residual attention image description method is proposed. By extracting the grid characteristics of the input image, three residual attention paths are constructed, and skip connections are added between the paths to generate attention scores. At the same time, a residual attention module of relative positions is introduced in the encoder, combining the relative position score with attention score, and a residual attention module with layer normalized query vector is introduced in the decoder.
By enhancing the connection between attention scores and effectively utilizing the relative position relationship, the problem of weak connection between attention scores is solved, and the quality and accuracy of image description are improved.
Smart Images

Figure CN114863222B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a three-channel residual attention image description method. Background Art
[0002] The attention mechanism and grid features are widely used in visual language tasks such as image description, and the attention score is a key factor for the success of the attention mechanism.
[0003] However, since Transformer (an attention mechanism model) is a hierarchical structure, the connection between the attention scores of different layers is not strong enough. When the grid features of an image enter the Transformer model, the geometric position information will be lost. Summary of the Invention
[0004] The purpose of the present invention is to provide a three-channel residual attention image description method, aiming to solve the problem that the connection between the attention scores of different layers in the existing attention mechanism model is not strong enough.
[0005] To achieve the above purpose, the present invention provides a three-channel residual attention image description method, including the following steps:
[0006] Extract the grid features of the input picture;
[0007] Construct three residual attention paths;
[0008] Generate attention scores by adding skip connections between the three residual attention paths;
[0009] Introduce a relative position residual attention module in the encoder to combine the relative position scores with the attention scores to obtain an updated encoder;
[0010] Introduce a residual attention module with a layer-normalized query vector in the decoder to obtain an updated decoder;
[0011] Construct and train an attention mechanism model based on the three residual attention paths, the updated encoder, and the updated decoder;
[0012] Input the grid features into the trained attention mechanism model for fusion and output to obtain an image text description.
[0013] Wherein, the specific method for extracting the grid features of the input picture is:
[0014] Set grid extraction parameters;
[0015] Extract the grid features of the input picture based on the grid extraction parameters using visual features.
[0016] Among them, the specific method of constructing and training the attention mechanism model based on the three residual attention paths, the update encoder, and the update decoder is as follows:
[0017] Construct a network model based on the three residual attention paths, the update encoder, and the update decoder;
[0018] Obtain an image caption benchmark dataset;
[0019] Delete the punctuation marks of all sentences in the image caption benchmark dataset, and convert all words in the image caption benchmark dataset to lowercase to obtain a training dataset;
[0020] Use the training dataset to train, validate, and test the network model to obtain an attention mechanism model.
[0021] Among them, the specific method of using the training dataset to train, validate, and test the network model to obtain an attention mechanism model is as follows:
[0022] Calculate the attention of the training dataset to obtain a calculated value;
[0023] Based on the calculated value, perform image annotation on the training dataset to obtain an annotated dataset;
[0024] Use the annotated dataset to train, validate, and test the network model to obtain an attention mechanism model.
[0025] Among them, the specific method of inputting the grid features into the trained attention mechanism model for fusion and then outputting to obtain an image text description is as follows:
[0026] Flatten the grid features to obtain flattened features;
[0027] Feed the flattened features into the trained attention mechanism model for fusion and then output to obtain an image text description.
[0028] A three-path residual attention image description method of the present invention extracts grid features of an input picture; constructs three residual attention paths; generates attention scores by adding skip connections between the three residual attention paths; introduces a relative position residual attention module in the encoder to combine the relative position scores with the attention scores to obtain an updated encoder; introduces a residual attention module with a layer-normalized query vector in the decoder to obtain an updated decoder; constructs and trains an attention mechanism model based on the three residual attention paths, the updated encoder, and the updated decoder; inputs the grid features into the trained attention mechanism model for fusion and then outputs to obtain an image text description. The relative position residual attention module makes full use of the relative position relationship, and it can combine the relative position scores with the attention scores, solving the problem that the connection between the attention scores of different layers of the existing attention mechanism model is not strong enough. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0030] Figure 1 is a flowchart of a three-path residual attention image description method provided by the present invention.
[0031] Figure 2 is an overview of the architecture of the Tri-RAT model.
[0032] Figure 3 is the standard attention mechanism in Transformer and the residual attention path of the present invention.
[0033] Figure 4 are two functions used to calculate relative coordinates.
[0034] Figure 5 is a schematic diagram of the encoder.
[0035] Figure 6 is a schematic diagram of the decoder. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The following will describe in detail the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention and should not be construed as a limitation of the present invention.
[0037] Please refer to Figures 1 to 6 , the present invention provides a three-channel residual attention image description method, including the following steps:
[0038] S1 Extract the grid features of the input picture;
[0039] Specifically:
[0040] S11 Set the grid extraction parameters;
[0041] S12 Extract the grid features of the input picture based on the grid extraction parameters using visual features.
[0042] S2 Construct three residual attention paths;
[0043] Specifically, the residual attention path (Residual Attention Path, abbreviated as RAP) propagates the previous attention score as a prior for the next layer through residual connection. It is an enhancement of the traditional attention mechanism. The traditional attention mechanism calculates the attention score independently in each layer. The performance of multi-head attention (Multi-Head Attetion, MHA) has been proven to be very excellent:
[0044] MHA(Q, K, V) = Concat(head 1 ,..., head h )W O
[0045]
[0046] where are the query, key, and value matrices respectively. For each matrix, n is the number of feature vectors and d is the dimension of each vector.
[0047] is the matrix that linearly maps the queries, keys, and values vectors to the "attention space" in each attention head. W O is a linear transformation that combines the attention results of different attention heads.
[0048] As Figure 3 shown, Figure 3On the left is the standard attention mechanism in Transformer, and on the right is the Residual Attention Path (RAP) of our Tri-RAT. The Residual Attention Path (RAP) is the direct path for propagating attention scores. The Previous Attention Score (PAS) is the attention score generated by the previous layer.
[0049] Standard Transformers usually implement using scaled dot-product attention:
[0050]
[0051] where Q′, K′, V′ are the matrices obtained by slicing the Q, K, V matrices to a certain head. The weighted sum of V′ is obtained using the Softmax function.
[0052] Different from the typical scaled dot-product attention, we create a direct path for propagating attention scores by adding a simple skip connection. Both the encoder and the decoder have their own direct paths to obtain the Previous Attention Score (PAS) of the previous layer. The attention module with the Residual Attention Path (RAP) can be defined as follows:
[0053]
[0054] where As the new attention score of this layer can be passed to the next layer. Note that since there is no Previous Attention Score (PAS) of the previous layer in the first layer,
[0055] PAS′ i the i in it must be greater than 2.
[0056] After the residual attention module, a fully-connected feedforward network (FFN) is applied:
[0057] FFN(x) = FC 2 (Dropout(RELU(FC 1 (x))))
[0058] where FC 1 and FC 2There are two linear layers. Finally, there are two Layer Normalization (LN) layers after the residual attention module and the FFN respectively:
[0059] x out = LN(x in + Sublayer(x in ))
[0060] where x in , x out are the input and output of a sublayer respectively, and the sublayer can be an attention layer or a feed-forward layer.
[0061] S3 generates attention scores by adding skip connections between the three residual attention paths;
[0062] S4 introduces a residual attention module with relative position in the encoder to combine the relative position scores with the attention scores to obtain an updated encoder;
[0063] Specifically, geometric position information should affect the attention scores. Therefore, we construct a residual attention module with relative position (Residual Attention with Relative Position, abbreviated as RARP) module in the encoder to obtain a better global feature representation. There is a big difference between grid features and region features: when grid features are flattened and fed into the Tri-RAT model, spatial information will inevitably be lost. Therefore, a new method should be developed to fully exploit the potential of grid features. We propose a residual attention module with relative position (Residual Attention with Relative Position, abbreviated as RARP) module to combine relative position weights with attention weights. To obtain the relative position score (Relative Position Score, abbreviated as RPS), we propose our relative position (Relative Position, abbreviated as RP) module. The relative position (Relative Position, abbreviated as RP) module first calculates the relative center coordinates cx i , cy i of the grid feature i as follows:
[0064]
[0065] where for grid i, and They are the relative position coordinates of the upper left corner and the lower right corner respectively. The specific calculation process of obtaining the relative position coordinates is the same as that of RSTNet. According to the extracted grid features, they have the same height and width. Considering this characteristic, we ignore the influence of height and width on the relative position. The relative geometric features can be defined as:
[0066]
[0067] Δx ij = sign(cx i - cx j )·log(1 + |cx i - cx j |),
[0068] Δy ij = sign(cy i - cx j )·log(1 + |cy i - cy j |),
[0069] where sign(·) returns a new tensor with the signs of the input elements. As Figure 4 shown, Figure 4 two functions used to calculate the relative position coordinates are shown. The V curve is the traditional log mapping function, and the S curve is the one we adopted. The new calculation method has more advantages than the traditional one. The new method replaces the old even function with a new odd function. Therefore, it can distinguish whether an object is on the left or right of another object. The old log function can only distinguish the magnitude of the distance between two features. In addition, when two features are getting closer and closer, the numerical representation of the relative position will be more continuous and stable.
[0070] Then, we apply Trigonometric Embedding (TE) to transform the relative geometric features into high-dimensional representations:
[0071] G ij = TE(r ij ),
[0072] where r ∈ R N×N×2 and are relative geometric features of different dimensions. N is the number of grid features extracted from the same image.
[0073] Finally, the Relative Position Score (RPS) is calculated as follows:
[0074] PRS = FC(G),
[0075] where FC(·) is a fully connected linear layer with the activation function ReLU. RPS ∈ R N×N can be directly added to the attention scores. Figure 5 is the Residual Attention with Relative Position (RARP) module in the encoder. The Relative Position Score (RPS) is added after the Previous Attention Score (PAS) of the previous layer and does not participate in the residual connection.
[0076] Note that the Relative Position Score (RPS) can only be used in Figure 5 the encoder shown. Therefore, the complete Residual Attention with Relative Position (RARP) module in the encoder can be described as:
[0077]
[0078] S5 introduces a residual attention module with layer-normalized query vectors in the decoder to obtain an updated decoder;
[0079] Specifically, a normalization method should be used in the attention module. Residual attention connections may exacerbate the degree of internal covariate shift. Therefore, the performance can be improved by introducing a normalization method. Layer normalization is created to reduce internal covariate shift, but it is usually used outside the attention module. Inspired by the image caption generation model NG-SAN, we perform layer normalization on Q inside the attention module to obtain a better data distribution:
[0080] Q = LNorm(Q),
[0081] where LNorm(·) is the layer normalization of the multi-head queries. Considering that every 8 attention heads are mapped from the same feature vector, our normalization should be performed along the last two dimensions (h × dk). Note that our Residual Attention Path (RAP) must be used to utilize layer normalization. Otherwise, it will lead to very poor training results. After normalization, the attention module can obtain a more reasonable data distribution to calculate the attention scores.
[0082] A module with a Residual Attention Path (RAP) and layer normalization is called a Residual Attention with Layer Normalization on Query Vectors (RALQ) module. Then we construct a decoder combined with the RALQ module:
[0083]
[0084] S6 constructs and trains an attention mechanism model based on the three residual attention paths, the updated encoder, and the updated decoder;
[0085] The specific method is as follows:
[0086] S61 constructs a network model based on the three residual attention paths, the updated encoder, and the updated decoder;
[0087] S62 obtains an image caption benchmark dataset;
[0088] S63 deletes the punctuation marks of all sentences in the image caption benchmark dataset and converts all words in the image caption benchmark dataset to lowercase to obtain a training dataset;
[0089] S64 uses the training dataset to train, validate, and test the network model to obtain an attention mechanism model.
[0090] The specific method is as follows:
[0091] S641 calculates the attention of the training dataset to obtain a calculated value;
[0092] S642 performs image annotation on the training dataset based on the calculated value to obtain an annotated dataset;
[0093] The training dataset is image-annotated by the Up-Down model and a template-based method. The template-based method is based on the generation of simple templates, which can be filled with the outputs of object detectors or attribute predictors. Inspired by the development of neural machine translation, the encoder-decoder paradigm is used in the caption model. For example, a CNN encodes an image into a feature vector, and an LSTM decodes it into a caption. Subsequently, the attention mechanism has become the focus due to its excellent ability to fuse visual and language contexts.
[0094] S643 uses the annotated dataset to train, validate, and test the network model to obtain an attention mechanism model.
[0095] Specifically, first, we train our Tri-RAT (network model) by optimizing the Cross Entropy (XE) loss L XE :
[0096]
[0097] where θ are the parameters to be trained, is the true text sequence.
[0098] Then, we directly use reinforcement training to optimize the non-differentiable metric:
[0099]
[0100] where the reward r(·) is the CIDEr-D score.
[0101] We also use the gradient expression in the Meshed-memory Transformer, where the average value of the rewards instead of greedy decoding is used as the baseline of the rewards. The gradient expression for a sample is:
[0102]
[0103] where k is the number of sample sequences, is the i-th sample sequence, and b is the average value of the rewards obtained from the sampled sequences.
[0104] The labeled dataset is the MS-COCO dataset. We evaluate our Tri-RAT on the MS-COCO dataset, which is the most popular image captioning benchmark dataset. The MS-COCO dataset contains 123,287 images. These images can be divided into 82,783 training images, 40,504 validation images, and 40,775 test images. Each of them has 5 different caption annotations. We adopt the splitting method provided by Karpathy et al. Among them, 5,000 images are used for validation, 5,000 images are used for testing, and the remaining images are used for training. Remove the punctuation marks from all sentences. All words are converted to lowercase, and if a word appears less than 5 times, it will be discarded.
[0105] Evaluation metrics. The evaluation metrics follow the standard evaluation protocol. We use a full set of caption metrics to evaluate the quality of image captions, including BLEU, METEOR, ROUGR, CIDEr, and SPICE.
[0106] Implementation details. Set hyperparameters according to Transformer and facilitate the training of our Tri-RAT. Different from that, our input image I is represented as grid features, just like visual features. The grid size is 7×7, and the image feature dimension is 2048. In our model, the dropout probability is 0.1, the d_model of Transformer is 512, the number of heads is 8, and the internal dimension of the fully-connected feedforward network (FFN) module is 2048.
[0107] The Adam optimizer is used to train our model. The learning rates of cross-entropy training lambda_lr and self-critical sequence training rl_lambda_lr are defined as follows:
[0108]
[0109] where base_lr and rl_base_lr are set to 0.0001 and 5×10 -6 respectively. e is the current epoch number starting from 0, and refine_epoch is set to 28.
[0110] S7 inputs the grid features into the trained attention mechanism model for fusion and then outputs an image text description.
[0111] The specific method is as follows:
[0112] S71 flattens the grid features to obtain flattened features;
[0113] S72 feeds the flattened features into the trained attention mechanism model for fusion and then outputs an image text description.
[0114] Ablation experiment
[0115] To comprehensively examine the impact of our proposed Residual Attention with Relative Position (RARP) and Residual Attention with Layer Normalization on Query Vectors (RALQ), we conducted a set of ablation experiments. First, we started with a base model that only uses the basic Transformer. Then, we incorporated the Residual Attention Path (RAP) and Layer Normalization on Query Vectors (LQ) into the base model separately. As shown in Table 1, the Residual Attention Path (RAP) can perform better in some metrics. LQ will reduce the performance. However, when we combine these two modules, excellent performance emerges.
[0116] Table 1 Ablation Study of RAP and LQ
[0117]
[0118] To fully utilize the positional information of grid features, we adopted different methods to obtain relative geometric positions. As shown in Table 2, Tri-RAT without any geometric positional relationship can obtain a CIDEr score of 133.2. Then we added a Grid-Augmented (GA) module, and the score increased to 133.7. We used new Log Space (LS) coordinates to replace the old coordinates in GA. LE and TE are learning embedding and triangular embedding respectively. They can achieve some increases in metrics. Finally, we proposed a new method to generate RPS and obtained the best performance.
[0119] Table 2 Comparison of Different Methods for Obtaining Geometric Information
[0120]
[0121] In the encoder, we can apply two different methods to add residual attention scores. The second method, RARPv2, which calculates the Previous Attention Score (PAS) of the previous layer, can be described as:
[0122]
[0123] Both of these two methods can make some improvements, but the first one can have better performance. The reason may be that the relative position information used in all layers should be the same. Therefore, the RPS should not be put into the Residual Attention Path (RAP) for propagation.
[0124] Table 3 Comparison of Different Residual Connection Methods
[0125]
[0126] Note that the number of layers plays an important role in our Tri-RAT. Therefore, we conducted a set of experiments to find the best choice. As shown in Table 4, when the number of attention layers is set to 3, we can obtain the best performance.
[0127] Table 4 Comparison of Different Numbers of Layers L
[0128]
[0129] Quantitative Analysis
[0130] Table 5 Performance Comparison with the Prior Art
[0131]
[0132]
[0133] Offline Evaluation
[0134] We report the performance of the offline test split of our Tri-RAT and the comparison models in Table 5. The models we compared include SCST, Up-Down, ORT, NG-SAN, AoANet,
[0135] M 2Transformer, X-Transformer, RSTNet, and DLCT. SCST proposed a new sequence training method called self-critical sequence training. Up-Down proposed a combined bottom-up and top-down attention mechanism that can compute attention at the level of objects and other salient image regions. ORT introduced the Transformer architecture into image captioning and incorporated information about the spatial relationships between regional features. NG-SAN proposed an effective reparameterization inside self-attention called Normalized Self-Attention (NSA) and introduced a class of Geometry-aware Self-Attention (GSA) to consider the relative geometric relationships between objects in the image. AoANet optimizes the attention results by computing the correlation between the attention results and the queries, further extending the traditional attention mechanism. M 2 Transformer constructs a fully connected architecture between each encoder layer and each decoder layer.
[0136] X-Transformer introduced bilinear pooling into the attention module of the base Transformer. RSTNet completely ignores regional features and directly applies self-attention to grid features, incorporating their relative geometric relationships into the self-attention calculation. DLCT proposed a hybrid method that combines regional and grid features to leverage their complementary advantages. As shown in Table 5, our Tri-RAT has better performance than other methods in terms of most evaluation metrics.
[0137] In addition to the usual single-model experiments, we also constructed an ensemble of 4 single models to achieve the image captioning task. Using two different features respectively, the ensemble model can obtain a CIDEr-D score of 137.6 using ResNext101 grid features and a CIDEr-D score of 139.3 using ResNext152 grid features.
[0138] Comparison with strong baselines. For a fair comparison, we also compared our Tri-RAT with other SOTA models on the same ResNext101 grid features. In addition to using the same features, these models are also under the same architecture configuration shown in RSTNet. As shown in Table 7, our model still outperforms other models in terms of most evaluation metrics.
[0139] Table 6 Comparison with SOTA on ResNext101 grid features
[0140]
[0141]
[0142] Online evaluation
[0143] We also evaluate our model on the online COCO test server. The model we use is an ensemble of 4 Tri-RAT models trained on the Karpathy training split. Table 7 shows the performance of our model compared to other SOTA models on the image captioning task.
[0144] Table 7. Leaderboard of state-of-the-art image annotation models released on the COCO online test server, where B@N, M, R, and C are abbreviations for BLEU@N, METEOR, ROUGE-L, and CIDEr scores.
[0145]
[0146] The above-disclosed is only a preferred embodiment of a three-channel residual attention image description method of the present invention. Of course, it cannot be used to limit the scope of the rights of the present invention. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. A three-path residual attention image description method, characterized in that, it includes the following steps: Extract the grid features of the input image; Construct three residual attention paths; Generate attention scores by adding skip connections between the three residual attention paths; Introduce a residual attention module with relative positions in the encoder to combine the relative position scores with the attention scores to obtain an updated encoder; Introduce a residual attention module with layer-normalized query vectors in the decoder to obtain an updated decoder; Construct and train an attention mechanism model based on the three residual attention paths, the updated encoder, and the updated decoder; Input the grid features into the trained attention mechanism model for fusion and output to obtain an image text description; The attention module of the residual attention path is defined as follows: Among them, Q′, K′, V′ are the matrices obtained by splitting the Q, K, V matrices to a certain head. Secondly, and are the query, key, and value matrices respectively. For each matrix, n is the number of feature vectors and d is the dimension of each vector. PAS refers to the attention score. As the new attention score of this layer, it is passed to the next layer. Note that since there is no attention score from the previous layer in the first layer, i in PAS′ i must be greater than 2; The residual attention module with relative positions refers to combining the relative position weights with the attention weights, and the relative position score, abbreviated as RPS, is calculated as follows: RPS = FC(G), where FC(·) is a fully connected linear layer with an activation function, and G is relative geometric features of different dimensions, and RELU; The complete residual attention module with relative positions in the encoder is described as: The residual attention module with layer-normalized query vectors refers to a module with a residual attention path and layer normalization, and then construct a decoder combined with the residual attention module with layer-normalized query vectors: where LNorm(·) is the layer normalization of the multi-head query.
2. The three-path residual attention image description method according to claim 1, characterized in that, The specific method for extracting the grid features of the input image is: Set grid extraction parameters; Extract the grid features of the input image based on the grid extraction parameters using visual features.
3. The three-path residual attention image description method according to claim 2, characterized in that, The specific method for constructing and training an attention mechanism model based on the three residual attention paths, the updated encoder, and the updated decoder is: Construct a network model based on the three residual attention paths, the updated encoder, and the updated decoder; Obtain an image caption benchmark dataset; Delete all punctuation marks in all sentences in the image caption benchmark dataset and convert all words in the image caption benchmark dataset to lowercase to obtain a training dataset; Use the training dataset to train, validate, and test the network model to obtain an attention mechanism model.
4. The three-path residual attention image description method according to claim 3, characterized in that, The specific method for using the training dataset to train, validate, and test the network model to obtain an attention mechanism model is: Calculate the attention of the training dataset to obtain a calculated value; Perform image annotation on the training dataset based on the calculated value to obtain an annotated dataset; Use the annotated dataset to train, validate, and test the network model to obtain an attention mechanism model.
5. The three-path residual attention image description method according to claim 4, characterized in that, The specific way of inputting the grid features into the trained attention mechanism model for fusion and then outputting to obtain the image text description is as follows: Flatten the grid features to obtain flattened features; Feed the flattened features into the trained attention mechanism model for fusion and then output to obtain the image text description.