A visual question answering method optimized using position information
By introducing a position attention mechanism in the visual question-and-answer model, the Transformer architecture is used to fuse visual and language information, and the existing model is solved inadequate attention to position information, realizing more accurate object relationship reasoning and more efficient training process.
Patent Information
- Application Number
- CN202210327078.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-30
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-03-30
AI Technical Summary
Existing visual question-and-answer models lack attention to position information, are difficult to answer position-related questions, and are difficult to establish more accurate inter-object relationships through position information.
Using a Transformer architecture with strong information extraction and information fusion capabilities, the information flow of visual modes and language modes is guided through position information, and the guiding role of the position attention mechanism in visual reasoning is realized. Specific steps include building an encoder and decoder module, performing multi-head position self-attention operation and position joint attention operation to blend visual and verbal information.
The model's ability to answer position-related questions is improved, the accuracy of the object relationship in the figure is enhanced, and the ability to reason is better. The training time is not much different than that of the previous model, and the efficiency is higher.
Smart Images

Figure CN114818739B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual question answering, and in particular to a visual question answering method optimized by using position information. Background Art
[0002] Visual question answering has been a hot topic in the field of artificial intelligence in recent years, attracting widespread attention from scholars. Given an image and questions related to the content in the image, the task requires the machine to correctly understand the image and the question and give an appropriate answer to the question. The task requires the model to have a fine-grained understanding of visual information and language information, and to be able to perform cross-modal reasoning like a human (connecting objects in the image with words in the question, understanding the logical relationship in the question, and being able to infer the corresponding answer based on the content in the image). The research results in this field have been widely used in the industry: tens of thousands of Taobao merchants have opened the visual question answering function of the Xiaomi customer service. AI has helped improve the question resolution rate, optimized the buyer experience, and reduced the configuration workload of merchants. The customer service scenarios of Hema and Koala, and the picture and text matching scenarios of Xianyu have also been connected to the visual question answering capability. In the future, visual question answering technology also has a broad imagination space. It can be used in picture and text reading, cross-modal search, visual question and answer for the blind, medical consultation, intelligent driving, virtual anchors and other fields, and may change the way of human-computer interaction. Moreover, vision and language are the main modalities in human real life. Solving this multimodal problem can accelerate the development of artificial intelligence.
[0003] Visual question answering is a very challenging task, but with people's continuous exploration, some progress has been made. The popular processing pipeline has evolved from directly using convolutional neural networks and recurrent neural networks to extract features from images and questions respectively when the task was first proposed, to using the trained object detection network to extract high-level semantic features of objects in the image, which better overcomes the "semantic gap" between the image information expressed by pixels and the question information conveyed by human language. In recent years, with the prosperity of the image and natural language communities, the cross-field of visual question answering has also seen a variety of methods flourish: 1) Multimodal fusion methods based on attention mechanisms can allow the model to focus on key objects in the image and related nouns in the text, but the model lacks a non-selection mechanism, and the normalization operation in the attention will make it impossible to introduce quantitative information into the model. 2) Methods based on modular networks that imitate the human reasoning process. This method performs grammatical parsing on questions, designs several types of basic modular networks, and inputs the questions into the corresponding modules for processing according to the parsed results. It has good interpretability, but the effect on real image datasets is not good. 3) Methods based on graph learning. This method regards objects in the image as nodes and the relationships between objects as edges to build the graph, but the effect is not much improved. 4) Methods that introduce external knowledge. There are many questions that require external knowledge to answer, but how to make the model better learn external knowledge is still a difficulty. 5) Optimization of the counting problem of the attention mechanism, etc., to a certain extent, improves the problem that the aforementioned attention mechanism cannot introduce counting information into the model, but the scope of application is small.
[0004] Recently, inspired by the Transformer structure and the great success of BERT in the field of natural language processing, the best performing new visual question answering models all have a Transformer-like structure, such as the 2019 Visual Question Answering Challenge champion model MCAN, and some large pre-trained general representation models such as VL-BERT. In particular, these large pre-trained general representation models can be applied to multiple visual question answering tasks, such as standard visual question answering tasks, tasks that describe images in text, and tasks that require retrieval of related images, without having to customize a model for each task. They are gradually becoming a research hotspot in the cross-field of visual question answering. However, this type of state-of-the-art model lacks attention to position information. For example, MCAN does not use the position of objects in the image, and large pre-trained general representation models simply add the position encoding and semantic encoding directly before inputting features, without making full use of position information. Therefore, this type of model has difficulty answering some questions related to position, and it is difficult to establish more accurate relationships between objects through position information.
[0005] In the prior art, "Visual Question Answering Method, Device, Medium, and Equipment Based on Deep Learning Model" has made improvements to the way of generating answers for the Transformer architecture. The academic community generally regards visual question answering as a classification problem, pre-setting thousands of answers, allowing the model to generate a predicted distribution of these answers in the end, and selecting the one with the highest probability as the answer. This method uses the result obtained by this predictive method as the first answer, and uses the Transformer decoder combined with a common word list to construct a generative answer to the question as the second answer. Finally, according to the predicted probabilities of the two answers, the first answer and / or the second answer are selected as the target answer corresponding to the question data and output. However, this method only focuses on the application of the Transformer architecture in the visual question answering industry. The answer table preset by the academic community is too small and not suitable for real scenarios, while an overly large answer table (some companies set a size of hundreds of thousands) will affect the training efficiency and deployment difficulty. This method uses generative answers to make a compromise. However, the basic structure of this method is still the unchanged Transformer architecture, which does not fully utilize the position information of objects in the picture, and the established relationship between objects is not accurate enough.
[0006] In the prior art, “A Visual Question Answering Method Based on Multimodal Fusion and Structural Control” inputs the semantic feature vector of the visual modality and the semantic feature vector of the language modality into a network based on a collaborative attention mechanism, calculates the multimodal information fusion feature vector, and then calculates the answer semantic feature vector based on the answer sample data set, and uses structural control to narrow the probability distribution gap between the multimodal information fusion feature vector and the answer semantic feature vector. However, this method uses an earlier collaborative attention network, and the basic structure of the model has not been changed, and the network's ability to extract and fuse information is poor. Before the model generates the answer prediction distribution, the invention adds the distance between the multimodal information fusion feature vector and the answer semantic feature vector as a loss to the model training, but the original training process already uses the gap between the model's predicted distribution and the true answer distribution as a loss, and this operation has a certain redundancy.
[0007] In the prior art, "A Visual Question Answering Method and System Based on Multiple Attentions" simultaneously inputs the image feature data and question information into the first long short-term memory network based on the attention mechanism, obtains the key image information by allocating attention weights, and then completes the combination of the question and the image through two bidirectional long short-term memory networks to give the answer to the problem to be processed. At the same time, a memory module is added to the pre-processing process of the image (target detection network) to expand the memory source of the model during the training process. However, this method uses a long short-term memory network to complete the reasoning of the problem, but this network is good at processing sequence-related data, and the object features in the picture are not sequence-related, but graph-related. Furthermore, this type of network cannot be calculated in parallel, and the training efficiency is low. And when you want to deepen the network depth, the problem of gradient disappearance will still occur.
[0008] In the prior art, "A Visual Question Answering Method and System Based on Matching Algorithm" constructs a structured scene graph for the input image, and constructs a structured text graph for the input question through natural language processing methods. Then the scene graph and the text graph are matched using a matching algorithm to get the answer to the question. However, this method uses a matching algorithm between the visual modality scene graph and the language modality text graph to predict the answer, which is similar to the module network that was popular in academia in the past few years and has good interpretability. However, due to the complexity of the real scene, the performance of this method is inferior to the various popular end-to-end methods today, and this method only has matching operations between the visual modality and the language modality, and lacks the step of reasoning about the problem. Summary of the invention
[0009] Most of the existing inventions adopt the existing model structure, and add some operations adapted to the actual implementation in the industry in the pre-processing stage or the post-processing part, and there are few improvements to the basic structure of the model. Some inventions that make improvements to the basic structure either adopt an architecture that is not powerful enough, or the improved method has lower training efficiency than the previous one. The present invention proposes a visual question answering method based on a Transformer network architecture with powerful information extraction and information fusion capabilities, which uses position information to guide the information flow within and between the visual mode and language mode, so that the position attention mechanism can better guide the relationship inference during visual reasoning. The training time required for this method is not much different from the visual question answering method based on the Transformer architecture before the modification, and the efficiency is higher. Since the position information is fully utilized, it is particularly good at some position-related problems, and the relationship between objects in the picture is generated more accurately, and it has better reasoning ability.
[0010] The purpose of the present invention is achieved by at least one of the following technical solutions.
[0011] A visual question answering method optimized by using position information comprises the following steps:
[0012] S1. Collect training data, including pictures and questions related to given pictures, and then manually calibrate the answers to the questions;
[0013] S2. Construct a question preprocessing module to preprocess the input question and obtain the semantic feature vector input and position feature vector of the question;
[0014] S3, construct an image preprocessing module to pre-process the input image to obtain the bounding box of the object in the image and the semantic feature vector of the image;
[0015] S4. Construct an encoder module and perform multi-head position self-attention operations to obtain the feature vector after information fusion of each word in the input question and the remaining words in the question, and obtain the fused feature vector of the words in the question:
[0016] S5. Construct a decoder module, perform position self-attention operation, and use the position joint attention mechanism to fuse the visual modality and language modality to obtain the fused feature vector of the object in the image;
[0017] S6, constructing a compression fusion module to compress and fuse the fusion feature vector of the object and the fusion feature vector of the word;
[0018] S7. Construct a prediction module to form a visual question answering model. Use the prediction module to predict the answer to the question, calculate the difference between the answer and the true value, and train the visual question answering model through back propagation. Input data into the trained visual question answering model to perform visual question answering.
[0019] Further, step S2 includes the following steps:
[0020] S2.1. Calculate the semantic feature vector of each word in the input question: initialize each word with GLoVe word embedding, and then input it into the long short-term memory network (LSTM) to obtain the semantic feature vector of a single word; since the length of each question is different, use zero vectors to pad or reduce it to obtain a semantic feature vector of dimension N×d1 representing the question, where N is the number of words in the question and d1 is the dimension of the semantic feature vector of a single word. Then use linear transformation to transform the feature dimension of the semantic feature vector of the question to N×d, where d is a uniform dimension. For ease of processing, all input vector dimensions are transformed to d;
[0021] S2.2. Calculate the positional feature vector of each word in the input question: The positional feature vector of a single word is determined by the position of each word in the question. It is calculated directly using word embedding, and a matrix of dimension N×d is used to represent the positional features of each word in the N positions in the question. The matrix is updated during back propagation.
[0022] Furthermore, in step S3, a target detection algorithm is used to extract semantic feature vectors of objects included in the image and a bounding box representing the position of the object in the image from the input image; this method can better narrow the "semantic gap" between the language modality and the visual modality than directly using a convolutional neural network to extract some structured features;
[0023] Since the number of objects extracted from each picture is different, each picture is padded with a zero vector or reduced to M objects to form a semantic feature vector of the picture with a dimension of M×d2, where M is the number of objects in the picture and d2 is the feature dimension obtained by the target detection algorithm. The semantic feature vector of the picture is then linearly transformed to M×d.
[0024] Furthermore, in step S4, the multi-head attention mechanism is performed on the basis of the normalized dot product attention operation, and the normalized dot product attention operation calculates the attention map in the following manner:
[0025]
[0026] Assume that q is of dimension 1×n, K is of dimension Y×n, and the output of f is a 1×Y-dimensional attention map. f generates an attention map of one item relative to the remaining Y items, where n represents the feature dimension of each item and Y is the number of items in K. Then the multi-head attention operation Multihead adds several parallel "heads" to the f function:
[0027] Multihead(Q i ,K,s)=Concate[head i1 ,…,head ij ,…,head is ]
[0028]
[0029] Among them, Q i represents the i-th dimension of Q, Q i The dimension of is 1×n, the dimension of Q is X×n, X represents the number of items in Q, s is the number of heads in the multi-head attention operation, head ij represents the attention map corresponding to the jth head, where j = 1~s, It is the jth head Q i The transformation matrix, W j K is the transformation matrix of K in the jth head, Multihead(Q i The output of ,K,h) is an attention map of dimension s×Y;
[0030] The fusion process of multi-head attention graphs can be abstractly represented as:
[0031] Fuse(map,V)=Concate[O1,…,O j ,…,O s ]
[0032] O j =map j (VW j V )
[0033] The Fuse function has two inputs, where map represents an attention map with dimensions s×Y and s heads, V represents a value vector of Y×n, and the output of the Fuse function is a 1×n fusion vector. j Indicates the fusion vector corresponding to the jth head, map j represents the jth dimension of the map, W j V represents the transformation matrix corresponding to V in the jth dimension;
[0034] Finally, the obtained fusion vector needs to be operated to enhance the stability of the network. The process can be abstractly represented as:
[0035] Stablize(Q,O)=LayerNorm(Z′+FFN(Z′))
[0036] Z′=Concate(Z′1,…,Z′ Y )
[0037] Z′ i =LayerNorm(Q i +O i W O )
[0038] The Stablize function has two inputs, where O is a vector of dimension Y×n, LayerNorm represents the layer normalization operation, which is used to maintain network stability. The addition in the middle of LayerNorm is actually a residual connection, which is often used when a deeper network needs to be trained; Z′ is a vector of dimension Y×n, Z′ i The dimension of is 1×n, which is the i-th dimension of Z′; W O Indicates O i The transformation matrix, O i is the i-th dimension of O, FFN is a feed-forward network layer, including two fully connected layers.
[0039] Further, in step S4, an encoder module with T layers of encoding layers stacked is constructed, and each encoding layer of the encoder module performs a position self-attention operation within the language modality; each encoding layer has two inputs and one output; the output of the last encoding layer is the output of the encoder module; each encoding layer has one input which is the position feature vector of the question, another input of the first encoding layer is the semantic feature vector of the question, and another input of each other encoding layer is the output of the previous encoding layer;
[0040] In each encoding layer, two inputs are converted into one output, including the following steps:
[0041] S4.1. Apply a multi-head attention operation to the input semantic feature vector of the question or the output of the previous encoding layer to obtain an attention map of each word in the question relative to other words;
[0042] To calculate the rough semantic attention map of each word relative to the rest of the words, first perform three linear transformations on the semantic feature vector of the question to obtain the semantic query vector Q of the question. Lsem , the semantic key vector K of the question Lsem , the semantic value vector V of the question Lsem , and then the coarse semantic attention map corresponding to each word relative to the rest of the words Calculated by the following formula:
[0043]
[0044] in, It's Q Lsem The i′th dimension of , i.e., the semantic query vector of the i′th word in the question, is a vector of dimension 1×d. h is the number of heads of the multi-head attention operation set in the actual experiment. The dimension is h×N, Represents the rough semantic attention map of the i′th word relative to the rest of the words;
[0045] S4.2, similar to step S4.1, perform a multi-head attention operation on the position feature vector of the language modality question, and calculate a rough position attention map of each word relative to the other words:
[0046]
[0047] in, represents the rough position attention map of the i′th word relative to the rest of the words, is the location query vector Q of the problem Lpos The i′th dimension of , represents the position query vector of the i′th word, K Lpos A position key vector representing the problem;
[0048] S4.3. Fuse the coarse position attention map and coarse semantic attention map of each word with the rest of the words to obtain the fused attention map of each word with the rest of the words:
[0049]
[0050] Among them, A i′ represents the fused attention map of the i′th word and the rest of the words, and softmax represents the normalized exponential function, so that the ratio of the i′th word to the vector of each other word is greater than 0;
[0051] The fusion vector L of the i′th word in the language modality i′ It is obtained by the following formula:
[0052] L i′ =Fuse(A i′ ,V Lsem );
[0053] S4.4. Output L of the positional self-attention operation within the language modality of each encoding layer O for:
[0054] L O =Stablize(Q Lsem ,L)
[0055] Among them, L is the fusion vector L corresponding to all words in the question i′ Cascade obtained.
[0056] Furthermore, in step S5, a decoder module with T layers of decoding layers stacked is constructed, and each decoding layer in the decoder module first performs a position self-attention operation within the visual modality, and then performs a position joint attention operation;
[0057] The position self-attention operation includes two inputs and one output. One input of the position self-attention operation in each decoding layer is the bounding box of all objects in the image. The other input of the position self-attention operation in the first decoding layer is the semantic feature vector of the image. The other input of the position self-attention operation in each other decoding layer is the output of the previous layer.
[0058] The position joint attention operation includes four inputs and one output. The output of the position joint attention operation is also the output of each decoding layer. The output of the last layer of the position joint attention operation is the output of the entire decoder module. The four inputs of the position joint attention operation are the output of the position self-attention operation, the output of the encoder module, the position feature vector of the problem, and the bounding boxes of all objects in the image.
[0059] Furthermore, in each decoding layer in the decoder module, a position self-attention operation is first performed within the visual modality, and then a position joint attention operation is performed, including the following steps:
[0060] S5.1. Perform position self-attention operation on the semantic feature vector of the input image or the output of the previous decoding layer and the bounding boxes of all objects in the image: In order to obtain a rough position attention map of each object relative to the rest of the objects, geometric processing is required. The position attention weight of the sth object in the image relative to the tth object Calculated by the following formula:
[0061]
[0062] b s and b t Respectively represent the bounding boxes of the sth object and the tth object, W G represents the d×1-dimensional transformation matrix, E G The input of the function is the bounding box of the sth object and the tth object, and the output is a d-dimensional vector, E G The function first calculates the relative position feature R of the sth object and the tth object st :
[0063]
[0064] Among them, x s Represents the horizontal coordinate of the center point of the bounding box of the sth object, x t Represents the horizontal coordinate of the center point of the bounding box of the tth object, y s Represents the ordinate of the center point of the bounding box of the sth object, y t represents the ordinate of the center point of the bounding box of the tth object, w s Indicates the width of the bounding box of the sth object, h s represents the height of the bounding box of the sth object, w t represents the width of the bounding box of the tth object, h t represents the height of the bounding box of the t-th object; the above method of calculating the relative position features of the s-th object and the t-th object can ensure that the relative position feature R st Translation and scaling invariance of ;
[0065] Then, the relative position feature R of the sth object and the tth object is calculated by using the method of calculating position encoding in the 'Transformer' model (calculating sine and cosine values at different wavelengths). st Mapped to a d-dimensional vector, which is also E G The output of the function is finally used by W GMap the vector into an attention weight to get the position attention weight of the sth object in the image relative to the tth object
[0066] Then calculate the rough position attention map of the sth object of the multi-head attention operation relative to the rest of the objects:
[0067]
[0068]
[0069] in, represents the rough position attention map of the sth object relative to other objects, express The kth head of , h refers to the number of heads of the multi-head attention operation; and They represent the coarse position attention weights of the sth object in the kth head relative to the 1st and Mth objects respectively;
[0070] In order to obtain a rough semantic attention map of the sth object in the visual modality relative to other objects, the semantic feature vector of the image is first linearly transformed three times to obtain the semantic query vector Q of the image: Isem , the semantic key vector K of the image Isem and the semantic value vector V of the image Isem , and then use the same operation as step S4.1 to process the semantic feature vector of the object using a multi-head attention operation:
[0071]
[0072] represents the rough semantic attention map of the sth object to the rest of the objects, with a dimension of h×M. represents the semantic query vector of the sth object, K Isem A semantic key vector representing the image;
[0073] Then, the same method as step S4.3 is used to fuse the coarse position attention map and coarse semantic attention map of the sth object in the visual modality relative to other objects:
[0074]
[0075] Among them, B s It represents the attention map of the fused sth object to the other objects, and then the information of other object features is fused according to the weight given by the attention map:
[0076] I s =Fuse(B s ,VIsem )
[0077] Among them I s represents the fusion vector of the sth object in the visual modality;
[0078] Finally, similar to step S4.4, it is still necessary to perform operations on the obtained feature vector to enhance the network stability:
[0079] I O =Stablize(Q Isem ,I)
[0080] Where I is the fusion vector I corresponding to all objects in the picture s Cascaded, I O is the output of the visual modality position self-attention operation at each encoding layer;
[0081] S5.2. In the position joint attention operation, the fused feature vector of the words in the question output by the encoder module is first linearly transformed twice to obtain the fused semantic key vector K sem and the fused semantic value vector V sem Output of the position self-attention operation I O Apply linear transformation to obtain the fused semantic value vector Q sem , and then use the multi-head attention operation to get a rough semantic joint attention map:
[0082]
[0083] in, represents the coarse semantic joint attention map of the p-th object to all words, is the fused semantic value vector of the pth object, i.e., Q sem The pth dimension of
[0084] Then calculate the coarse position joint attention map of the pth object for all words, input the bounding box of all objects in the picture and the position feature vector of the sentence;
[0085] The bounding box encoding is linearly transformed into an object position feature vector with the same dimension as the word position feature vector, and then the position feature vector of the problem is linearly transformed twice to obtain the fused position key vector K pos and the fused position value vector V pos , linearly transform the position feature vector of the image to the fused position query vector Q pos , the coarse position joint attention map of the p-th object for all words Calculated by the following formula:
[0086]
[0087] in, Represents the position query vector of the pth object, i.e. Q pos The pth dimension of
[0088] Next, fusion and
[0089]
[0090] Among them, C p Represents the joint attention map of the p-th object to all words;
[0091] Then the semantic features of the sentence are fused according to the joint attention map:
[0092] U p =Fuse(C p ,V sem )
[0093] Among them, U p represents the fusion vector of the pth object in the position joint attention operation;
[0094] Finally, the obtained feature vector is used to enhance the stability of the network:
[0095] U O =Stablize(Q sem ,U)
[0096] Among them, U is the fusion vector U obtained by the position joint attention operation of all objects in the picture p Cascaded, U O is the output of the positional joint attention operation and is also the output of each decoding layer in the decoder module.
[0097] Furthermore, the bounding box of the object is represented by the coordinates of the upper left corner and the lower right corner of the object in the figure. First, the four-dimensional bounding box is converted into a bounding box code (x_min / W, y_min / H, x_max / W, y_max / H, Area), where x_min represents the minimum horizontal coordinate in the bounding box, y_min represents the minimum vertical coordinate in the bounding box, x_max represents the maximum horizontal coordinate in the bounding box, y_max represents the maximum vertical coordinate in the bounding box, W represents the width of the bounding box, H represents the height of the bounding box, and Area represents the area of the bounding box. That is, for stability, the values of x_min, y_min, x_max, and y_max are first reduced to the range of 0 to 1, and then the area is added as a new feature.
[0098] Furthermore, in step S6, in the compression fusion module, the fused feature vector of the object and the fused feature vector of the word are compressed and fused, as follows:
[0099]
[0100]
[0101] E s′ is the s′th dimension of the output vector of the encoder module, MLP is a multilayer perceptron, and E′ is N E s′ The compressed vector, s′=1~N; D t′ is the t′th dimension of the decoder module output vector, and D′ is the M D t′ The compressed vector, t′=1~M; the above formula means calculating an attention weight for each feature vector, normalizing it with softmax and then taking the weighted average to get a feature vector; finally, the following formula is used to fuse the two:
[0102]
[0103] is the transformation matrix of E′, is the transformation matrix of D′; F is the compressed fusion vector of objects and words output by the compression fusion module.
[0104] Further, in step 7, the visual question answering model includes a question preprocessing module, an image preprocessing module, an encoder module, a decoder module, a compression fusion module and a prediction module;
[0105] In the prediction module, the score of each candidate answer obtained by F is calculated as follows:
[0106] s′=σ(ω o F)
[0107] σ ensures that each score is in the range of 0 to 1, s′ is the score of each candidate answer predicted by the visual question answering model, and ω o is the transformation matrix of F;
[0108] The loss function of the visual question answering model is as follows:
[0109]
[0110] Where b is the number of all questions, r is the number of all alternative answers, and s is p′q′ represents the true score of the q′th alternative answer in the p′th question, s′ p′q′ It represents the score predicted by the model for the q′th alternative answer in the p′th question, and then the Adam optimizer is used to train the visual question answering model.
[0111] Compared with the prior art, the advantages of the present invention are:
[0112] 1. The visual question answering method proposed in the present invention can better understand the problem. The present invention tested the situation where the position attention mechanism is only applied to the language modality in the model, and the ordinary attention mechanism is applied to other parts. It is found that the performance of the model in answering a type of question with a yes or no answer is significantly improved compared with the basic model. This is because the present invention amplifies the importance of the position of the word in the sentence to the model, which helps the model understand the semantics of the sentence.
[0113] 2. The visual question-answering method proposed in the present invention can better establish the connection between objects in the picture. The present invention tests the situation where the position attention mechanism is only applied to the self-attention part of the visual modality in the model, and the ordinary attention mechanism is applied to other parts. It is found that the model has a significant improvement in performance when answering a type of questions with numerical answers compared with the basic model. This is because many counting problems involve positional relationships. For example, "How many cats are there on the sofa?" The model needs to clearly know the position of the object in the picture.
[0114] 3. The present invention requires about 30% more training time than the original model, and is more efficient. In addition, the present invention can sacrifice a little accuracy to further optimize time efficiency, that is, only the position-related attention map is calculated in the first layer, and the results of this layer are used in the subsequent layers. After this operation, there is almost no need to increase training time compared to the original model. BRIEF DESCRIPTION OF THE DRAWINGS
[0115] Figure 1 is an overall flow chart of a visual question answering method optimized by using position information in an embodiment of the present invention;
[0116] Figure 2 It is a structural diagram of a position self-attention mechanism and a position joint attention mechanism in an embodiment of the present invention;
[0117] Figure 3 It is a schematic diagram of the connection relationship between each layer in the decoder and the encoder in an embodiment of the present invention. DETAILED DESCRIPTION
[0118] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the specific implementation of the present invention is described in detail below with reference to the accompanying drawings and examples.
[0119] Embodiment 1:
[0120] A visual question answering method optimized using position information, such as Figure 1 As shown, the following steps are included:
[0121] S1. Collect training data, including pictures and questions related to given pictures, and then manually calibrate the answers to the questions; this embodiment directly uses the VQA v2.0 dataset.
[0122] S2. Construct a question preprocessing module to preprocess the input question to obtain the semantic feature vector input and position feature vector of the question, including the following steps:
[0123] S2.1. Calculate the semantic feature vector of each word in the input question: initialize each word with GLoVe word embedding, and then input it into the long short-term memory network (LSTM) to obtain the semantic feature vector of a single word; since the length of each question is different, use zero vectors to pad or reduce it to obtain a semantic feature vector of dimension N×d1 representing the question, where N is the number of words in the question and d1 is the dimension of the semantic feature vector of a single word. Then use linear transformation to transform the feature dimension of the semantic feature vector of the question to N×d, where d is a uniform dimension. For ease of processing, all input vector dimensions are transformed to d;
[0124] In this embodiment, taking the VQA v2.0 dataset as an example, each question is unified to a length of 14, that is, N is 14, d1 is 300, and d is 512, so that a vector representing the semantic features of the question with a dimension of 14×512 is obtained.
[0125] S2.2. Calculate the positional feature vector of each word in the input question: The positional feature vector of a single word is determined by the position of each word in the question. It is calculated directly using word embedding, and a matrix of dimension N×d is used to represent the positional features of each word in the N positions in the question. The matrix is updated during back propagation.
[0126] S3, construct an image preprocessing module to pre-process the input image to obtain the bounding box of the object in the image and the semantic feature vector of the image;
[0127] The object detection algorithm is used to extract the semantic feature vector of the object included in the image and the bounding box representing the location of the object in the image from the input image; this method can better narrow the "semantic gap" between the language modality and the visual modality than directly using convolutional neural networks to extract some structured features;
[0128] Since the number of objects extracted from each picture is different, each picture is padded with a zero vector or reduced to M objects to form a semantic feature vector of the picture with a dimension of M×d2, where M is the number of objects in the picture and d2 is the feature dimension obtained by the target detection algorithm. The semantic feature vector of the picture is then linearly transformed to M×d.
[0129] In this embodiment, the Faster RCNN network model (a target detection algorithm) based on the ResNet-101 convolutional network is used to complete this step. First, the Faster RCNN network model is trained on the Visual Genome data set. In order to enhance its feature representation ability, it is not only allowed to predict the category of the object, but also the attribute information of the object needs to be inferred. Then the pictures collected by the visual question-answering task are input into the trained Faster RCNN network model. The model uses bounding boxes of different shapes and sizes to obtain pre-selected objects with each pixel as the center. After position adjustment, it determines whether it contains objects and other operations, and outputs a series of bounding boxes, as well as the feature vector of each bounding box corresponding to the object. Finally, in order to obtain the feature vector of the visual modality for the visual question-answering task, in this embodiment, the non-maximum suppression algorithm is used to control the number of objects finally obtained for the large number of candidate objects output by the Faster RCNN network model, reduce the overlap between the bounding boxes, and obtain the semantic feature vector dimension d2 of 2048, and the bounding box is composed of the coordinates of the upper left pixel and the lower right pixel. Similar to the aforementioned processing of the input problem, the number of objects obtained in each image is different and still needs to be unified into one number. In this embodiment, 100 is used as the unified number, and each image is supplemented with zero vectors or deleted to 100 objects.
[0130] S4. Construct an encoder module and perform multi-head position self-attention operations to obtain the feature vector after information fusion of each word in the input question and the remaining words in the question, and obtain the fused feature vector of the words in the question:
[0131] like Figure 2 As shown in the figure, the multi-head attention mechanism is based on the normalized dot product attention operation. The normalized dot product attention operation calculates the attention map as follows:
[0132]
[0133] Assume that q is of dimension 1×n, K is of dimension Y×n, and the output of f is a 1×Y-dimensional attention map. f generates an attention map of one item relative to the remaining Y items, where n represents the feature dimension of each item and Y is the number of items in K. Then the multi-head attention operation Multihead adds several parallel "heads" to the f function:
[0134] Multihead(Q i ,K,s)=Concate[head i1 ,…,head ij ,…,head is ]
[0135]
[0136] Among them, Q i represents the i-th dimension of Q, Q i The dimension of is 1×n, the dimension of Q is X×n, X represents the number of items in Q, s is the number of heads in the multi-head attention operation, head ij represents the attention map corresponding to the jth head, where j = 1~s, It is the jth head Q i The transformation matrix, W j K is the transformation matrix of K in the jth head, Multihead(Q i The output of ,K,h) is an attention map of dimension s×Y;
[0137] The fusion process of multi-head attention graphs can be abstractly represented as:
[0138] Fuse(map,V)=Concate[O1,…,O j ,…,O s ]
[0139] O j =map j (VW j V )
[0140] The Fuse function has two inputs, where map represents an attention map with dimensions s×Y and s heads, V represents a value vector of Y×n, and the output of the Fuse function is a 1×n fusion vector. j Indicates the fusion vector corresponding to the jth head, map j represents the jth dimension of the map, W j V represents the transformation matrix corresponding to V in the jth dimension;
[0141] Finally, the obtained fusion vector needs to be operated to enhance the stability of the network. The process can be abstractly represented as:
[0142] Stablize(Q,O)=LayerNorm(Z′+FFN(Z′))
[0143] Z′=Concate(Z′1,…,Z′ Y )
[0144] Z′ i =LayerNorm(Q i +O i W O )
[0145] The Stablize function has two inputs, where O is a vector of dimension Y×n, LayerNorm represents the layer normalization operation, which is used to maintain network stability. The addition in the middle of LayerNorm is actually a residual connection, which is often used when a deeper network needs to be trained; Z′ is a vector of dimension Y×n, Z′ i The dimension of is 1×n, which is the i-th dimension of Z′; W O Indicates O i The transformation matrix, O i is the i-th dimension of O, FFN is a feed-forward network layer, including two fully connected layers.
[0146] Construct an encoder module with T layers of encoding layers stacked together. Each encoding layer of the encoder module performs position self-attention operations within the language modality. Each encoding layer has two inputs and one output. The output of the last encoding layer is the output of the encoder module. Each encoding layer has one input, which is the position feature vector of the question. Another input of the first encoding layer is the semantic feature vector of the question. Another input of each other encoding layer is the output of the previous encoding layer.
[0147] In each encoding layer, two inputs are converted into one output, including the following steps:
[0148] S4.1. Apply a multi-head attention operation to the input semantic feature vector of the question or the output of the previous encoding layer to obtain an attention map of each word in the question relative to other words;
[0149] To calculate the rough semantic attention map of each word relative to the rest of the words, first perform three linear transformations on the semantic feature vector of the question to obtain the semantic query vector Q of the question. Lsem , the semantic key vector K of the question Lsem , the semantic value vector V of the question Lsem , and then the coarse semantic attention map corresponding to each word relative to the rest of the words Calculated by the following formula:
[0150]
[0151] in, It's Q Lsem The i′th dimension of , i.e., the semantic query vector of the i′th word in the question, is a vector of dimension 1×d. h is the number of heads of the multi-head attention operation set in the actual experiment. The dimension is h×N, Represents the rough semantic attention map of the i′th word relative to the rest of the words;
[0152] In this embodiment, h is set to 8. Query vector Q Lsem , the key vector KLsem , value vector V Lsem The dimensions are all 14×512.
[0153] S4.2, similar to step S4.1, perform a multi-head attention operation on the position feature vector of the language modality question, and calculate a rough position attention map of each word relative to the other words:
[0154]
[0155] in, represents the rough position attention map of the i′th word relative to the rest of the words, is the location query vector Q of the problem Lpos The i′th dimension of , represents the position query vector of the i′th word, K Lpos A position key vector representing the problem;
[0156] In this embodiment, The dimension is 1×512, K Lpos The dimensions are 14×512.
[0157] S4.3. Fuse the coarse position attention map and coarse semantic attention map of each word with the rest of the words to obtain the fused attention map of each word with the rest of the words:
[0158]
[0159] Among them, A i′ represents the fused attention map of the i′th word and the rest of the words, and softmax represents the normalized exponential function, so that the ratio of the i′th word to the vector of each other word is greater than 0;
[0160] The fusion vector L of the i′th word in the language modality i′ It is obtained by the following formula:
[0161] L i′ =Fuse(A i′ ,V Lsem );
[0162] In this embodiment, V Lsem The dimension is 14×512, A i , and The dimensions are all 8×14, L i The dimension is 1×512.
[0163] S4.4. Output L of the positional self-attention operation within the language modality of each convolutional layer O for:
[0164] LO =Stablize(Q Lsem ,L)
[0165] Among them, L is the fusion vector L corresponding to all words in the question i′ Cascade obtained.
[0166] In this embodiment, the dimension of L is 14×512, and Q Lsem The dimension is 14×512, L O The dimensions are 14×512.
[0167] S5. Construct a decoder module, perform position self-attention operation, and use the position joint attention mechanism to fuse the visual modality and language modality to obtain the fused feature vector of the object in the image;
[0168] Construct a decoder module with T layers of decoding layers stacked together. Each decoding layer in the decoder module first performs a position self-attention operation within the visual modality, and then performs a position joint attention operation;
[0169] The position self-attention operation includes two inputs and one output. One input of the position self-attention operation in each decoding layer is the bounding box of all objects in the image. The other input of the position self-attention operation in the first decoding layer is the semantic feature vector of the image. The other input of the position self-attention operation in each other decoding layer is the output of the previous layer.
[0170] The position joint attention operation includes four inputs and one output. The output of the position joint attention operation is also the output of each decoding layer. The output of the last layer of the position joint attention operation is the output of the entire decoder module. The four inputs of the position joint attention operation are the output of the position self-attention operation, the output of the encoder module, the position feature vector of the question, and the bounding boxes of all objects in the image.
[0171] In each decoding layer, the position self-attention operation is performed within the visual modality first, and then the position joint attention operation is performed, including the following steps:
[0172] S5.1. Perform position self-attention operation on the semantic feature vector of the input image or the output of the previous decoding layer and the bounding boxes of all objects in the image: In order to obtain a rough position attention map of each object relative to the rest of the objects, geometric processing is required. The position attention weight of the sth object in the image relative to the tth object Calculated by the following formula:
[0173]
[0174] b s and b tRespectively represent the bounding boxes of the sth object and the tth object, W G represents the d×1-dimensional transformation matrix, E G The input of the function is the bounding box of the sth object and the tth object, and the output is a d-dimensional vector, E G The function first calculates the relative position feature R of the sth object and the tth object st :
[0175]
[0176] Among them, x s Represents the horizontal coordinate of the center point of the bounding box of the sth object, x t Represents the horizontal coordinate of the center point of the bounding box of the tth object, y s Represents the ordinate of the center point of the bounding box of the sth object, y t represents the ordinate of the center point of the bounding box of the tth object, w s Indicates the width of the bounding box of the sth object, h s represents the height of the bounding box of the sth object, w t represents the width of the bounding box of the tth object, h t represents the height of the bounding box of the t-th object; the above method of calculating the relative position features of the s-th object and the t-th object can ensure that the relative position feature R st Translation and scaling invariance of ;
[0177] Then, the relative position feature R of the sth object and the tth object is calculated by using the method of calculating position encoding in the 'Transformer' model (calculating sine and cosine values at different wavelengths). st Mapped to a d-dimensional vector, which is also E G The output of the function is finally used by W G Map the vector into an attention weight to get the position attention weight of the sth object in the image relative to the tth object
[0178] Then calculate the rough position attention map of the sth object of the multi-head attention operation relative to the rest of the objects:
[0179]
[0180]
[0181] in, represents the rough position attention map of the sth object relative to other objects, express The kth head of , h refers to the number of heads of the multi-head attention operation; and They represent the coarse position attention weights of the sth object in the kth head relative to the 1st and Mth objects respectively;
[0182] In order to obtain a rough semantic attention map of the sth object in the visual modality relative to other objects, the semantic feature vector of the image is first linearly transformed three times to obtain the semantic query vector Q of the image: Isem , the semantic key vector K of the image Isem and the semantic value vector V of the image Isem , and then use the same operation as step S4.1 to process the semantic feature vector of the object using a multi-head attention operation:
[0183]
[0184] represents the rough semantic attention map of the sth object to the rest of the objects, with a dimension of h×M. represents the semantic query vector of the sth object, K Isem A semantic key vector representing the image;
[0185] Then, the same method as step S4.3 is used to fuse the coarse position attention map and coarse semantic attention map of the sth object in the visual modality relative to other objects:
[0186]
[0187] Among them, B s It represents the attention map of the fused sth object to the other objects, and then the information of other object features is fused according to the weight given by the attention map:
[0188] I s =Fuse(B s ,V Isem )
[0189] Among them I s represents the fusion vector of the sth object in the visual modality;
[0190] Finally, similar to step S4.4, it is still necessary to perform operations on the obtained feature vector to enhance the network stability:
[0191] I O =Stablize(Q Isem ,I)
[0192] Where I is the fusion vector I corresponding to all objects in the picture s Cascaded, I O is the output of the visual modality position self-attention operation at each encoding layer;
[0193] In this embodiment, B s The dimensions are all 8×100, Q Isem , K Isem , V Isem The dimensions of are 100×512, and the dimensions of I are 100×512.
[0194] S5.2. In the position joint attention operation, the fused feature vector of the words in the question output by the encoder module is first linearly transformed twice to obtain the fused semantic key vector K sem and the fused semantic value vector V sem Output of the position self-attention operation I O Apply linear transformation to obtain the fused semantic value vector Q sem , and then use the multi-head attention operation to get a rough semantic joint attention map:
[0195]
[0196] in, represents the coarse semantic joint attention map of the p-th object to all words, is the fused semantic value vector of the pth object, i.e., Q sem The pth dimension of
[0197] Then calculate the coarse position joint attention map of the pth object for all words, input the bounding box of all objects in the picture and the position feature vector of the sentence;
[0198] The bounding box of an object is represented by the coordinates of the upper left corner and the lower right corner of the object in the figure. First, the four-dimensional bounding box is converted into a bounding box code (x_min / W, y_min / H, x_max / W, y_max / H, Area), where x_min represents the smallest horizontal coordinate in the bounding box, y_min represents the smallest vertical coordinate in the bounding box, x_max represents the largest horizontal coordinate in the bounding box, y_max represents the largest vertical coordinate in the bounding box, W represents the width of the bounding box, H represents the height of the bounding box, and Area represents the area of the bounding box. That is, for stability, the values of x_min, y_min, x_max, and y_max are first reduced to the range of 0 to 1, and the area is added as a new feature; then the bounding box code is linearly transformed into an object position feature vector with the same dimension as the word position feature vector, and then the position feature vector of the problem is linearly transformed twice to obtain the fused position key vector K pos and the fused position value vector V pos , linearly transform the position feature vector of the image to the fused position query vector Q pos, the coarse position joint attention map of the p-th object for all words Calculated by the following formula:
[0199]
[0200] in, Represents the position query vector of the pth object, i.e. Q pos The pth dimension of
[0201] Next, fusion and
[0202]
[0203] Among them, C p Represents the joint attention map of the p-th object to all words;
[0204] Then the semantic features of the sentence are fused according to the joint attention map:
[0205] U p =Fuse(C p ,V sem )
[0206] Among them, U p represents the fusion vector of the pth object in the position joint attention operation;
[0207] Finally, the obtained feature vector is used to enhance the stability of the network:
[0208] U O =Stablize(Q sem ,U)
[0209] Among them, U is the fusion vector U obtained by the position joint attention operation of all objects in the picture p Cascaded, U O is the output of the positional joint attention operation and is also the output of each decoding layer in the decoder module.
[0210] In this embodiment, C p The dimensions are all 8×14, K sem , V sem , K pos , V pos The dimensions of Q are all 14×512. sem , Q pos The dimensions of U, U are all 100×512. O The dimensions are all 100×512.
[0211] S6, constructing a compression fusion module to compress and fuse the fusion feature vector of the object and the fusion feature vector of the word;
[0212] In the compression fusion module, the fused feature vector of the object and the fused feature vector of the word are compressed and fused as follows:
[0213]
[0214]
[0215] E s′ is the s′th dimension of the output vector of the encoder module, MLP is a multilayer perceptron, and E′ is N E s′ The compressed vector, s′=1~N; D t′ is the t′th dimension of the decoder module output vector, and D′ is the M D t′ The compressed vector, t′=1~M; the above formula means calculating an attention weight for each feature vector, normalizing it with softmax and then taking the weighted average to get a feature vector; finally, the following formula is used to fuse the two:
[0216]
[0217] is the transformation matrix of E′, is the transformation matrix of D′; F is the compressed fusion vector of objects and words output by the compression fusion module.
[0218] S7. Construct a prediction module to form a visual question answering model. Use the prediction module to predict the answer to the question, calculate the difference between the answer and the true value, and train the visual question answering model through back propagation. Input data into the trained visual question answering model to perform visual question answering.
[0219] like Figure 3 As shown, the visual question answering model includes a question preprocessing module, an image preprocessing module, an encoder module, a decoder module, a compression fusion module and a prediction module;
[0220] In the prediction module, the score of each candidate answer obtained by F is calculated as follows:
[0221] s′=σ(ω o F)
[0222] σ ensures that each score is in the range of 0 to 1, s′ is the score of each candidate answer predicted by the visual question answering model, and ω o is the transformation matrix of F;
[0223] The loss function of the visual question answering model is as follows:
[0224]
[0225] Wherein, b represents the number of all questions. In this example, b is 443757, that is, the size of the VQA v2.0 training set is 443757. r represents the number of all alternative answers. In this example, r is 2847, that is, the number of alternative answers for VQA v2.0 is 2847. s p′q′ represents the true score of the q′th alternative answer in the p′th question, s′ p′q′ It represents the score predicted by the model for the q′th alternative answer in the p′th question, and then the Adam optimizer is used to train the visual question answering model.
[0226] In this embodiment, the learning rate is set to 1e-4, and the warm-up training technique and the learning rate decay technique are used, that is, the learning rates of 2.5e-5, 5e-5, and 7.5e-5 are used in the first three rounds of model training, the learning rate of 1e-4 is used in the next seven rounds, and the learning rates of 2e-5, 2e-5, and 4e-6 are used in the last three rounds. A total of 13 rounds of training are performed, and the parameters are saved after completion for subsequent practical applications.
[0227] Embodiment 2:
[0228] In this embodiment, the difference from Embodiment 1 is that this embodiment uses the VQA v1.0 data set, b in step S7 is 248349, that is, the size of the VQA v1.0 training set is 248349, and r in step 7 is 2410, that is, the number of alternative answers of VQA v1.0 is 2410.
[0229] Embodiment 3:
[0230] In this embodiment, the difference from Embodiment 1 is that this embodiment uses the COCO-QA dataset, b in step S7 is 78736, that is, the size of the COCO-QA training set is 78736, and r in step 7 is 435, that is, the number of alternative answers of COCO-QA is 435.
Claims
1. A visual question answering method optimized using position information, characterized in that: The following steps are involved: S1. Collect training data, including pictures and questions related to given pictures, and then manually calibrate the answers to the questions; S2. Construct a question preprocessing module to preprocess the input question and obtain the semantic feature vector input and position feature vector of the question; S3, construct an image preprocessing module to pre-process the input image to obtain the bounding box of the object in the image and the semantic feature vector of the image; S4. Construct an encoder module and perform multi-head position self-attention operations to obtain the feature vector after information fusion of each word in the input question and the remaining words in the question, and obtain the fused feature vector of the words in the question: S5. Construct a decoder module, perform position self-attention operation, and use the position joint attention mechanism to fuse the visual modality and language modality to obtain the fused feature vector of the object in the image; S6, constructing a compression fusion module to compress and fuse the fusion feature vector of the object and the fusion feature vector of the word; S7. Construct a prediction module to form a visual question answering model. Use the prediction module to predict the answer to the question, calculate the difference between the answer and the true value, and train the visual question answering model through back propagation. Input data into the trained visual question answering model to perform visual question answering.
2. The visual question answering method using position information optimization according to claim 1, characterized in that: Step S2 includes the following steps: S2.
1. Calculate the semantic feature vector of each word in the input question: initialize each word with GLoVe word embedding, and then input it into the long short-term memory network to obtain the semantic feature vector of a single word; since the length of each question is different, use zero vectors to fill or reduce it to obtain a semantic feature vector of dimension N×d1 representing the question, where N is the number of words in the question and d1 is the dimension of the semantic feature vector of a single word. Then use linear transformation to transform the feature dimension of the semantic feature vector of the question to N×d, where d is a uniform dimension and all input vector dimensions are transformed to d; S2.
2. Calculate the positional feature vector of each word in the input question: The positional feature vector of a single word is determined by the position of each word in the question. It is calculated directly using word embedding, and a matrix of dimension N×d is used to represent the positional features of each word in the N positions in the question. The matrix is updated during back propagation.
3. The visual question answering method using position information optimization according to claim 2, characterized in that: In step S3, a target detection algorithm is used to extract a semantic feature vector of an object included in the image and a bounding box representing the position of the object in the image from the input image; Since the number of objects extracted from each picture is different, each picture is padded with a zero vector or reduced to M objects to form a semantic feature vector of the picture with a dimension of M×d2, where M is the number of objects in the picture and d2 is the feature dimension obtained by the target detection algorithm. The semantic feature vector of the picture is then linearly transformed to M×d.
4. The visual question answering method using position information optimization according to claim 1, characterized in that: In step S4, the multi-head attention mechanism is performed on the basis of the normalized dot product attention operation. The normalized dot product attention operation calculates the attention map in the following way: Assume that q is of dimension 1×n, K is of dimension Y×n, and the output of f is a 1×Y-dimensional attention map. f generates an attention map of one item relative to the remaining Y items, where n represents the feature dimension of each item and Y is the number of items in K. Then the multi-head attention operation Multihead adds several parallel "heads" to the f function: Multihead(Q i ,K,s)=Concate[head i1 ,…,head ij ,…,head is ] Among them, Q i represents the i-th dimension of Q, Q i The dimension of is 1×n, the dimension of Q is X×n, X represents the number of items in Q, s is the number of heads in the multi-head attention operation, head ij represents the attention map corresponding to the jth head, where j = 1~s, It is the jth head Q i The transformation matrix, is the transformation matrix of K in the jth head, Multihead(Q i The output of ,K,h) is an attention map of dimension s×Y; The fusion process of multi-head attention graph is abstractly expressed as: Fuse(map,V)=Concate[O1,…,O j ,…,O s ] The Fuse function has two inputs, where map represents an attention map with dimensions s×Y and s heads, V represents a value vector of Y×n, and the output of the Fuse function is a 1×n fusion vector. j Indicates the fusion vector corresponding to the jth head, map j represents the j-th dimension of the map, represents the transformation matrix corresponding to V in the jth dimension; Finally, the obtained fusion vector needs to be operated to enhance the stability of the network. The process is abstractly expressed as follows: Stablize(Q,O)=LayerNorm(Z′+FFN(Z′)) Z′=Concate(Z′1,…,Z′ Y ) Z′ i =LayerNorm(Q i +O i W O ) The Stabilize function has two inputs, where O is a vector of dimension Y×n, and LayerNorm represents the layer normalization operation; Z′ is a vector of dimension Y×n, and Z i The dimension of ′ is 1×n, which is the i-th dimension of Z′; W O Indicates O i The transformation matrix, O i is the i-th dimension of O, FFN is a feed-forward network layer, including two fully connected layers.
5. The visual question answering method using position information optimization according to claim 4, characterized in that: In step S4, an encoder module with T layers of encoding layers stacked is constructed, and each encoding layer of the encoder module performs a position self-attention operation within the language modality; each encoding layer has two inputs and one output; the output of the last encoding layer is the output of the encoder module; each encoding layer has one input which is the position feature vector of the question, another input of the first encoding layer is the semantic feature vector of the question, and another input of each other encoding layer is the output of the previous encoding layer; In each encoding layer, two inputs are converted into one output, including the following steps: S4.
1. Apply a multi-head attention operation to the input semantic feature vector of the question or the output of the previous encoding layer to obtain an attention map of each word in the question relative to other words; To calculate the rough semantic attention map of each word relative to the rest of the words, first perform three linear transformations on the semantic feature vector of the question to obtain the semantic query vector Q of the question. Lsem , the semantic key vector K of the question Lsem , the semantic value vector V of the question Lsem , and then the coarse semantic attention map corresponding to each word relative to the rest of the words Calculated by the following formula: in, It's Q Lsem The i′th dimension of , i.e., the semantic query vector of the i′th word in the question, is a vector of dimension 1×d. h is the number of heads of the multi-head attention operation set in the actual experiment. The dimension is h×N, Represents the rough semantic attention map of the i′th word relative to the rest of the words; S4.2, similar to step S4.1, perform a multi-head attention operation on the position feature vector of the language modality question, and calculate a rough position attention map of each word relative to the other words: in, represents the rough position attention map of the i′th word relative to the rest of the words, is the location query vector Q of the problem Lpos The i′th dimension of , represents the position query vector of the i′th word, K Lpos A position key vector representing the problem; S4.
3. Fuse the coarse position attention map and coarse semantic attention map of each word with the rest of the words to obtain the fused attention map of each word with the rest of the words: Among them, A i′ represents the fused attention map of the i′th word and the rest of the words, and softmax represents the normalized exponential function, so that the ratio of the i′th word to the vector of each other word is greater than 0; The fusion vector L of the i′th word in the language modality i′ It is obtained by the following formula: L i′ =Fuse(A i′ ,V Lsem ); S4.
4. Output L of the positional self-attention operation within the language modality of each encoding layer O for: L O =Stablize(Q lsem ,L) Among them, L is the fusion vector L corresponding to all words in the question i′ Cascade obtained.
6. The visual question answering method using position information optimization according to claim 5, characterized in that: In step S5, a decoder module with T layers of decoding layers stacked is constructed, and each decoding layer in the decoder module first performs a position self-attention operation within the visual modality, and then performs a position joint attention operation; The position self-attention operation includes two inputs and one output. One input of the position self-attention operation in each decoding layer is the bounding box of all objects in the image. The other input of the position self-attention operation in the first decoding layer is the semantic feature vector of the image. The other input of the position self-attention operation in each other decoding layer is the output of the previous layer. The position joint attention operation includes four inputs and one output. The output of the position joint attention operation is also the output of each decoding layer. The output of the last layer of the position joint attention operation is the output of the entire decoder module. The four inputs of the position joint attention operation are the output of the position self-attention operation, the output of the encoder module, the position feature vector of the problem, and the bounding boxes of all objects in the image.
7. The visual question answering method using position information optimization according to claim 6, characterized in that: In each decoding layer of the decoder module, the position self-attention operation within the visual modality is performed first, and then the position joint attention operation is performed, including the following steps: S5.
1. Perform position self-attention operation on the semantic feature vector of the input image or the output of the previous decoding layer and the bounding boxes of all objects in the image: In order to obtain a rough position attention map of each object relative to the rest of the objects, geometric processing is required. The position attention weight of the sth object in the image relative to the tth object Calculated by the following formula: b s and b t Respectively represent the bounding boxes of the sth object and the tth object, W G represents the d×1-dimensional transformation matrix, E G The input of the function is the bounding box of the sth object and the tth object, and the output is a d-dimensional vector, E G The function first calculates the relative position feature R of the sth object and the tth object st : Among them, x s Represents the horizontal coordinate of the center point of the bounding box of the sth object, x t Represents the horizontal coordinate of the center point of the bounding box of the tth object, y s Represents the ordinate of the center point of the bounding box of the sth object, y t represents the ordinate of the center point of the bounding box of the tth object, w s Indicates the width of the bounding box of the sth object, h s represents the height of the bounding box of the sth object, w t represents the width of the bounding box of the tth object, h t represents the height of the bounding box of the t-th object; the above method of calculating the relative position features of the s-th object and the t-th object can ensure that the relative position feature R st Translation and scaling invariance of ; Then, the relative position feature R of the sth object and the tth object is calculated using the position encoding method in the 'transformer' model st Mapped to a d-dimensional vector, which is also E G The output of the function is finally used by W G Map the vector into an attention weight to get the position attention weight of the sth object in the image relative to the tth object Then calculate the rough position attention map of the sth object of the multi-head attention operation relative to the rest of the objects: in, represents the rough position attention map of the sth object relative to other objects, express The kth head of , h refers to the number of heads of the multi-head attention operation; and They represent the coarse position attention weights of the sth object in the kth head relative to the 1st and Mth objects respectively; In order to obtain a rough semantic attention map of the sth object in the visual modality relative to other objects, the semantic feature vector of the image is first linearly transformed three times to obtain the semantic query vector Q of the image: Isem , the semantic key vector K of the image Isem and the semantic value vector V of the image Isem , and then use the same operation as step S4.1 to process the semantic feature vector of the object using a multi-head attention operation: represents the rough semantic attention map of the sth object to the rest of the objects, with a dimension of h×M. represents the semantic query vector of the sth object, K Isem A semantic key vector representing the image; Then, the same method as step S4.3 is used to fuse the coarse position attention map and coarse semantic attention map of the sth object in the visual modality relative to other objects: Among them, B s It represents the attention map of the fused sth object to the other objects, and then the information of other object features is fused according to the weight given by the attention map: I s =Fuse(B s ,V Isem ) Among them I s represents the fusion vector of the sth object in the visual modality; Finally, similar to step S4.4, it is still necessary to perform operations on the obtained feature vector to enhance the network stability: I O =Stablize(Q Isem ,I) Where I is the fusion vector I corresponding to all objects in the picture s Cascaded, I O is the output of the visual modality position self-attention operation at each encoding layer; S5.
2. In the position joint attention operation, the fused feature vector of the words in the question output by the encoder module is first linearly transformed twice to obtain the fused semantic key vector K sem and the fused semantic value vector V sem Output of the position self-attention operation I O Apply linear transformation to obtain the fused semantic value vector Q sem , and then use the multi-head attention operation to get a rough semantic joint attention map: in, represents the coarse semantic joint attention map of the p-th object to all words, is the fused semantic value vector of the pth object, i.e., Q sem The pth dimension of Then calculate the coarse position joint attention map of the pth object for all words, input the bounding box of all objects in the picture and the position feature vector of the sentence; The bounding box encoding is linearly transformed into an object position feature vector with the same dimension as the word position feature vector, and then the position feature vector of the problem is linearly transformed twice to obtain the fused position key vector K pos and the fused position value vector V pos , linearly transform the position feature vector of the image to the fused position query vector Q pos , the coarse position joint attention map of the p-th object for all words Calculated by the following formula: in, Represents the position query vector of the pth object, i.e. Q pos The pth dimension of Next, fusion and Among them, C p Represents the joint attention map of the p-th object to all words; Then the semantic features of the sentence are fused according to the joint attention map: U p =Fuse(C p ,V sem ) Among them, U p represents the fusion vector of the pth object in the position joint attention operation; Finally, the obtained feature vector is operated to enhance the stability of the network: IN O =Stabilize(Q sem ,IN) Among them, U is the fusion vector U obtained by the position joint attention operation of all objects in the picture p Cascaded, U O is the output of the positional joint attention operation and is also the output of each decoding layer in the decoder module.
8. The visual question answering method using position information optimization according to claim 7, characterized in that: The bounding box of an object is represented by the coordinates of the upper left corner and the lower right corner of the object in the figure. First, the four-dimensional bounding box is converted into a bounding box code (x_min / W, y_min / H, x_max / W, y_max / H, Area), where x_min represents the minimum horizontal coordinate in the bounding box, y_min represents the minimum vertical coordinate in the bounding box, x_max represents the maximum horizontal coordinate in the bounding box, y_max represents the maximum vertical coordinate in the bounding box, W represents the width of the bounding box, H represents the height of the bounding box, and Area represents the area of the bounding box. That is, for stability, the values of x_min, y_min, x_max, and y_max are first reduced to the range of 0 to 1.
9. A visual question answering method using position information optimization according to any one of claims 1 to 8, characterized in that: In step S6, in the compression fusion module, the fused feature vector of the object and the fused feature vector of the word are compressed and fused, as follows: E s′ is the s′th dimension of the output vector of the encoder module, MLP is a multilayer perceptron, and E′ is N E s′ The compressed vector, s′=1~N; D t′ is the t′th dimension of the decoder module output vector, and Dt is the M D t′ The compressed vector, t′=1~M; the above formula means calculating an attention weight for each feature vector, normalizing it with softmax and then taking the weighted average to get a feature vector; finally, the following formula is used to fuse the two: is the transformation matrix of E′, is the transformation matrix of D′; F is the compressed fusion vector of objects and words output by the compressed fusion module.
10. The visual question answering method using position information optimization according to claim 9, characterized in that: In step 7, the visual question answering model includes a question preprocessing module, an image preprocessing module, an encoder module, a decoder module, a compression fusion module and a prediction module; In the prediction module, the score of each candidate answer obtained by F is calculated as follows: s′=σ(ε o F) σ ensures that each score is in the range of 0 to 1, s′ is the score of each candidate answer predicted by the visual question answering model, and ω o is the transformation matrix of F; The loss function of the visual question answering model is as follows: Where b is the number of all questions, r is the number of all alternative answers, and s is p′q′ represents the true score of the p′th alternative answer in the q′th question, s p ' ′q′ It represents the score predicted by the model for the q′th alternative answer in the p′th question, and then the Adam optimizer is used to train the visual question answering model.
Citation Information
Patent Citations
Visual dialogue method, visual dialogue model training method, device and equipment
CN111897940A
Knowledge base question-answering method fusing multi-head-attention mechanism and relative position encoding
CN113704437A