Visual Question Answering Method Based on Nested Attention Network with Selective Graph Structure

By introducing the select graph structure module and nested attention network module in the visual question-and-answer system, the problem of existing systems ignoring the relationship between goals and targets when dealing with visual features is solved, and clearer answers and more comprehensive semantic understanding are achieved.

CN115248872BActive Publication Date: 2025-05-27CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111068174.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-13
Publication Date
2025-05-27
Estimated Expiration
2041-09-13

AI Technical Summary

Technical Problem

When processing visual features, existing visual question and answer systems tend to ignore the relationship between the target and the target, and when generating weighted average vectors, the method based on the stacking attention mechanism may output irrelevant information, resulting in misleading.

Method used

The graph structure module is used to construct the graph structure for visual features, and select edges and nodes of the graph structure according to the problem features to find out the relationship between the targets and the targets in the visual features. At the same time, a nested attention network module is built to enable linear operation of the features generated by common attention and the original features, eliminate irrelevant information, and effectively integrate visual and problem features.

Benefits of technology

By selecting the graph structure module to clarify the relationship between goals and goals in the visual characteristics, the answer is clearer. The nested attention network module effectively eliminates irrelevant information, improves the fusion effect of visual features and problem features, and achieves a more comprehensive semantic understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115248872B_ABST
    Figure CN115248872B_ABST
Patent Text Reader

Abstract

The present invention discloses a visual question answering method based on a nested attention network with a selection graph structure. Previous methods mainly used convolutional neural networks or attention mechanisms for processing visual features, which would ignore the relationships between objects in visual features. In addition, in co-attention, regardless of whether the visual features are relevant to the question features, the attention will output weighted averages for both visual features and question features. The present invention for the first time proposes a nested attention network based on a selection graph structure to study the correspondence between images and questions. A nested attention reinforcement network is designed to effectively fuse visual features and question features by considering the relationships between regions and eliminating irrelevant information in vectors. A selection graph structure module is proposed to find the relationships between visual feature objects based on the question, making the answer clearer. The present invention has conducted a large number of experiments on VQA 2.0 to prove the effectiveness of the proposed model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the method of visual question answering, and relates to the technical fields of computer vision and natural language processing. Background Art

[0002] Visual question answering is formulated as a classification problem in most studies, with an image and a question as inputs and an answer as the output class (due to the limited number of possible answers). Since the visual question answering task was proposed after the widespread popularity of deep learning methods, almost all current visual question answering solutions use CNNs to model the image input and RNNs to model the question. Attention mechanisms have been widely studied in visual question answering. This includes visual attention, which focuses on dealing with the problem of where to look, and question attention, which focuses on solving the problem of where to read. Since images and questions are two different modalities, it is straightforward to jointly embed the two modalities together to uniformly describe the image / question pair.

[0003] The common practice of existing models is to separately extract visual and language features and then merge them into a common space. Then, based on these fused bimodal features, the answer to the input question is predicted. In early studies, researchers adopted some relatively simple fusion methods, such as feature concatenation, multiplication, and dot product of feature vectors. Fukui et al. demonstrated that more complex fusion methods can indeed improve the prediction accuracy, so they introduced a bilinear (merging) method. In their work, the outer product of two vectors of visual and language features was used for fusion. Since the external output has a very high-dimensional feature, they adopted the concept of Gao et al. Gao et al. compressed the fused features and named it the MCB merging method. However, to ensure stable performance, the compressed features of MCB still tend to be high-dimensional. Kim et al. used the Hadamard product of two feature vectors to propose a low-rank bilinear pool, called the multimodal low-rank bilinear pool (MLB). Yu et al. proposed a multimodal factorization bilinear pool (MFB), which uses matrix factorization techniques to calculate the fused features, thereby reducing the number of parameters and improving the convergence speed.

[0004] Attention mechanisms are effective in many visual and language processing tasks, such as caption generation, action recognition, natural language processing, and so on. Without exception, it has been introduced into visual question answering and has been proven to be helpful for answer prediction. So far, many methods have been developed, and the commonly used one is to guide attention in image regions. According to the type of image features, the methods are divided into two categories. On the one hand, visual features of region proposals are used to focus on objects, and these attention objects are generated by bounding boxes or region proposal networks. Another type of visual feature is extracted from convolutional features.

[0005] There are several ways to create and use attention maps. Yang et al. developed a stacked attention network that generates multiple attention maps on an image in a sequential manner, aiming to perform multiple inference steps. Kim et al. extended this idea by incorporating it into the residual architecture to produce better attention information. Chen et al. proposed a structured attention model that can encode relationships across regions, aiming to correctly answer questions involving relationships between complex regions. Duy-Kien Nguyen et al. proposed the well-known co-attention mechanism to better fuse the representations of images and question words. However, existing attention models mainly consider the possible interactions between image regions and question words, while ignoring the self-correlation information of the image regions themselves. Additionally, some network structures are multi-layer iterative, usually causing some valuable but unattended original image edge information to be completely forgotten after multiple bilateral co-attention operations. Summary of the Invention

[0006] The object of the present invention is to solve the problem that current visual question answering systems generally use convolutional neural networks or attention mechanisms to process visual features, which will ignore the relationships between objects in visual features. Moreover, in the visual question answering method based on the stacked attention mechanism, regardless of whether the visual features are relevant to the question features, the attention module will output a weighted average of the visual features and the question features. Even without a relevant vector, the attention module will still generate a weighted average vector, which may be irrelevant and even misleading.

[0007] The technical solutions adopted by the present invention to solve the above technical problems are as follows:

[0008] S1. Construct a selection graph structure module to construct a graph structure for visual features, then select edges and nodes for the graph structure according to the question features, and find the relationships between visual feature objects based on the question, making the answer clearer.

[0009] S2. Construct a nested attention network module to perform a linear operation on the features generated by co-attention with the original features of another modality, eliminating vector-irrelevant information and effectively fusing visual features and question features.

[0010] S3. Combine the network in S2 and the network in S3 to construct a nested attention network architecture based on the selection graph structure.

[0011] S4. Training and visual question answering of the nested attention network based on the selection graph structure.

[0012] The selection graph structure module of the present invention generates a graph structure according to image features, and then finds the relationships between visual feature objects based on the question, making the answer clearer. We will describe the detailed operations below:

[0013] The graph structure is a directed multi - graph, where each node corresponds to a scene entity, which can be an object associated with a bounding box or an attribute of an object. Each scene entity has a type corresponding to the predicted object or attribute label. Typed edges specify how the scene entities are related to each other. More formally, let E denote the set of scene entities and consider the set of binary relations R. Then, the graph structure is the set of ordered triples (s, p, o) of subject, predicate, and object. In the remainder of this work, integrity is imposed on the inverse relation because for every (s, p -1 , o) ∈ SG there is implied (s, p

[0014] The state of the mapping space S is given by E × Q, where E is the nodes of the scene graph SG and Q represents the set of all questions. The mapping space S at the t - th node represents the current entity e t and question Q, denoted as St. Thus, St=(e t , Q) represents the state S t ∈ S for t ∈ N. The set of available actions from the mapping space St is denoted by A St . It contains all the outgoing edges of the node e t and their corresponding object nodes, represented by the following formula:

[0015] A st ={(r, e) ∈ R × E: S t =(et, Q) ∧ (et, r, e) ∈ SG} (1)

[0016] Also, let A t ∈ A St represent the action performed by the mapping space at the t - th node. When creating the graph structure, node embeddings are passed through a multi - layer graph attention network (GAT). GAT extends it from the graph convolutional network through a self - attention mechanism, mimicking the convolutional operator on a regular grid by embedding the node features of its neighbors to form entity embeddings. Thus, the generated embeddings are context - aware, which makes nodes of the same type but with different graph neighborhoods distinguishable.

[0017] Extract features from the input image and question respectively, and then generate the graph structure from the image features. We use the tuple H t =(H t-1 , A t-1 ) to represent the history of the mapping space, where t ≥ 1, H 0 =hub, t = 0. The feature h is encoded through a multi - layer LSTM, as shown in the following formula:

[0018] h t= LSTM(a t-1 ) (2)

[0019] where a t-1 = [r t-1 , e t ∈ R 2d corresponds to using rt-1 to embed the previous action, and e t etc. respectively represent embedding the edge and the target node into R d .

[0020] h and the question feature Q are simultaneously input into the ReLU activation function, and after passing through the action set operation At(), the visual feature I is generated through the softmax function, as shown in the following formula:

[0021] I = softmax(A t (W 2 ReLU(W 1 [h t Q]))) (3)

[0022] where the rows of contain the latent representations of all admissible actions. In addition, Q ∈ R d encodes the question Q. The action A t = (r, e) ∈ A St is drawn according to the feature I. Formulas (2) and (3) yield a stochastic policy π θ , where θ represents the set of trainable parameters.

[0023] The nested attention network module mines more complete visual and text features based on the correlations between input image regions and between question words, eliminates vector-irrelevant information, and effectively fuses visual and question features. By considering the association degree between regions in the image, the overall semantic understanding bias is reduced. We describe the detailed operations below:

[0024] Given the image feature I and the question feature T, they are input into the nested attention model, and after training of the model, the corresponding modal features are finally generated.

[0025] The middle part of the model is a classical co-attention architecture. First, for the given image feature I k and the question feature T k , the attention matrix A is obtained through cross multiplication, and then A is passed through a double-layer softmax to generate the attention matrix A T about the question and the attention matrix A I about the image. Finally, they are multiplied by the image feature and the question feature respectively to obtain the image feature I R containing bilateral information and the question feature T RThis process can be represented by the following five formulas:

[0026] A = T k ×W R ×I k T (4)

[0027] A I = softmax(A) (5)

[0028] A T = softmax(A T ) (6)

[0029] T′ k = T k ×A T (7)

[0030] I′ k = I k ×A I (8)

[0031] Among them, in formula (1), represents the weight matrix, and the dimensions of the attention matrix are both d×d.

[0032] The overall structure of the model enhances the correlation between feature regions in bilateral information. For the question feature, the question feature T′ k generated in the previous step is k concatenated with the image feature I k through a concat operation, and then respectively passed through two linear operations (with different parameters). One of them is then input into a Sigmoid activation function, and the two features are multiplied to obtain the final question feature T

[0033] T ck = Concat(T′ k , I k ) (9)

[0034] T lk = Linear(T ck ) (10)

[0035] T lsk = Sigmoid(Linear(T ck )) (11)

[0036] T K = T lk ×T lsk (12)

[0037] Among them, linear is a Linear function that has 1024 hidden units with ReLU non-linearity and dropout.

[0038] Similarly, for the image features, the image features I' generated in the previous step k and the question features T k are subjected to a concat connection operation, and then respectively pass through two linear operations (with different parameters), one of which is then input into a Sigmoid activation function, and the two features are multiplied to obtain the final image features I k (k = k + 1). This process can be represented by the following two formulas:

[0039] I ck = Concat(I' k , T k ) (13)

[0040] I lk = Linear(I ck ) (14)

[0041] I lsk = Sigmoid(Linear(I ck )) (15)

[0042] I K = I lk ×I lsk (16)

[0043] Among them, linear is a Linear function that has 1024 hidden units with ReLU non-linearity and dropout.

[0044] The image features I k (k = k + 1) and the question features T k (k = k + 1) are the outputs of the nested attention model, and their dimensions are the same as those of the input features. Then, the above operations are repeated with them as the input of the nested attention model. When k = n, it is the final output, that is, the image features I n and the question features T n .

[0045] The visual question answering method based on the nested attention network with a selection graph structure includes a nested attention module and a selection graph structure module.

[0046] The training method of the nested attention network based on the selection graph structure is as follows:

[0047] In our implementation, all experiments were conducted using the PyTorch framework with Python 3.6, and the experiments were carried out on a computer equipped with an Nvidia Tesla P100 GPU.

[0048] Before feeding into the CNN, the sizes of all images were resized to 448*448. All questions were tokenized using the Python Natural Language Toolkit (nltk). We used the vocabulary provided by the CommonCrawl-840B Glove model for English word vectors. We limited the maximum length of questions to 14 words and then dynamically padded each question to allow questions of different lengths. Throughout the experiment, we used an eight-layer network, namely a network with eight layers of compound attention and visual feature enhancement mechanisms (L = 8). This number of layers was selected based on our preliminary experiments. During training, the ADAM optimizer was used to train our model in VQA 2.0 with 400 batches respectively. The weight decay was 0.01. We gradually decreased the learning rate using exponential decay:

[0049]

[0050] where the initial learning rate α was set to α = 0.001, and the decay epochs for VQA2.0 were set to 7 epochs in sequence; we set the parameters to β 1 = 0.9, β 2 = 0.99. To prevent overfitting, Dropout was used, with a dropout rate of ρ = 0.3 for each fully connected layer and a dropout rate of ρ = 0.1 for the LSTM.

[0051] Compared with the existing technologies, the beneficial effects of the present invention are:

[0052] 1. This paper proposes an image question answering model with a nested attention mechanism, which can linearly operate on the features generated by co-attention with the original features of another modality to eliminate vector-unrelated information and effectively fuse visual features and question features, achieving a more comprehensive semantic understanding and fusion.

[0053] 2. The present invention designs a selection graph structure module to find the relationships between visual feature targets based on questions, making the answers clearer. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 FIG. is a schematic structural diagram of a visual question answering method based on a nested attention network with a selection graph structure.

[0055] Figure 2 FIG. is a schematic model diagram of a selection graph structure network.

[0056] Figure 3 It is a schematic diagram of the model for the nested attention network.

[0057] Figure 4 It is a comparison graph of the results of the visual question answering model of the nested attention network based on the selection graph structure and the visual question answering models of other networks on the dataset.

[0058] Figure 5 It is a visualization result graph of visual question answering. Detailed implementation manners

[0059] The accompanying drawings are only for illustrative purposes and should not be construed as limitations of this patent.

[0060] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0061] Figure 1 It is a schematic diagram of the nested attention network structure based on the selection graph structure. As Figure 1 shown, the entire framework of visual question answering mainly consists of two parts: the selection graph structure (Selected Graph Network) and the nested attention (Nested attention mechanism).

[0062] Figure 2 It is a schematic diagram of the selection graph structure module. As Figure 2 shown, the graph structure is a directed multigraph, where each node corresponds to a scene entity, and the scene entity can be an object associated with a bounding box or an attribute of an object. Each scene entity has a type corresponding to the predicted object or attribute label. The typed edges specify how the scene entities are related to each other. More formally, let E denote the set of scene entities and consider the set of binary relations R. Then, the graph structure is the set of ordered triples (s, p, o) of subject, predicate, and object. In the remainder of this work, integrity is imposed for the inverse relation because for each (s, p -1 , o) ∈ SG, there is implicitly (s, p

[0063] The state of the mapping space S is given by E × Q, where E is the node of the scene graph SG and Q represents the set of all questions. The mapping space S at the t-th node represents the current entity e t and the question Q, denoted as St. Therefore, St = (e t , Q) represents the state S of t ∈ N t ∈ S. The set of available actions from the mapping space St is denoted by A St . It contains the node e tAll the outgoing edges and their corresponding object nodes are represented by the following formula:

[0064] A st ={(r, e) ∈ R × E: S t =(et, Q) ∧ (et, r, e) ∈ SG} (1)

[0065] At the same time, use A t ∈ A St to represent the action performed by the mapping space at the t-th node. When creating the graph structure, node embeddings are passed through a multi-layer graph attention network (GAT). GAT extends it from the graph convolutional network through the self-attention mechanism, mimicking the convolutional operator on a regular grid and forming entity embeddings by embedding the node features of its neighbors. Therefore, the generated embeddings are context-aware, which makes nodes of the same type but different graph neighborhoods distinguishable.

[0066] Extract features from the input image and question respectively, and then generate a graph structure from the image features. We use the tuple H t =(H t-1 , A t-1 ) to represent the history of the mapping space, where t ≥ 1, H 0 = hub, t = 0. The feature h is encoded through a multi-layer LSTM, as shown in the following formula:

[0067] h t = LSTM(a t-1 ) (2)

[0068] where a t-1 =[r t-1 , e t ∈ R 2d corresponds to embedding the previous action using rt-1, and e t etc. respectively represent embedding the edge and the target node into R d .

[0069] h and the question feature Q are simultaneously put into the ReLU activation function, and then after the action set operation At(), the visual feature I is generated through the softmax function, as shown in the following formula:

[0070] I = softmax(A t (W 2 ReLU(W 1 [h t Q]))) (3)

[0071] where the rows of contain the latent representations of all admissible actions. In addition, Q ∈ R dEncode the question Q. Draw the action A according to the feature I t =(r, e) ∈ A st . Formulas (2) and (3) yield a stochastic policy π θ , where θ represents the set of trainable parameters.

[0072] Figure 3 Schematic diagram of the nested attention module. As Figure 3 shown, the middle part of the model is a classical co-attention architecture. First, for the given image feature I k and the question feature T k , the attention matrix A is obtained through cross multiplication, and then A is passed through a double-layer softmax to generate the attention matrix A T about the question and the attention matrix A I about the image. Finally, multiply them with the image feature and the question feature respectively to obtain the image feature I R and the question feature T R containing bilateral information. This process can be represented by the following five formulas:

[0073] A = T k × W R × I k T (4)

[0074] A I = softmax(A) (5)

[0075] A T = softmax(A T ) (6)

[0076] T′ k = T k × A T (7)

[0077] I′ k = I k × A I (8)

[0078] where in formula (1) represents the weight matrix, and the dimensions of the attention matrices are both d × d.

[0079] The overall structure of the model enhances the correlation between feature regions in the bilateral information. For the question feature, concatenate the question feature T′ k generated in the previous step with the image feature I k , then pass through two linear operations (with different parameters) respectively. One of them is then input into a Sigmoid activation function, and finally multiply the two features to obtain the final question feature Tk (k = k + 1). This process can be represented by the following two formulas:

[0080] T ck = Concat(T′ k , I k ) (9)

[0081] T lk = Linear(T ck ) (10)

[0082] T lsk = Sigmoid(Linear(T ck )) (11)

[0083] T K = T lk × T lsk (12)

[0084] Among them, linear is a Linear function that has 1024 hidden units with ReLU non-linearity and dropout.

[0085] Similarly, for image features, the image feature I′ k generated in the previous step is concatenated with the question feature T k through a concat operation, and then respectively passes through two linear operations (with different parameters), one of which is then input into a Sigmoid activation function, and the two features are multiplied to obtain the final image feature I k (k = k + 1). This process can be represented by the following two formulas:

[0086] I ck = Concat(I′ k , T k ) (13)

[0087] I lk = Linear(I ck ) (14)

[0088] I lsk = Sigmoid(Linear(I ck )) (15)

[0089] I K = I lk × I lsk (16)

[0090] Among them, linear is a Linear function that has 1024 hidden units with ReLU non-linearity and dropout.

[0091] Image Features I k (k=k+1) and problem feature T k (k=k+1) is the output of the nested attention model, and its dimension is consistent with the dimension of the input feature. Then use it as the input of the nested attention model and repeat the above operation. When k=n, it is the final output, that is, the image feature I n With problem feature T n .

[0092] Figure 4 The following is a comparison of the visual question answering results of the nested attention network based on the selection graph structure and other network visual question answering models on the VQA2.0 dataset. Figure 4 As shown in Figure 2, the visual question answering results of the nested attention network based on the selection graph structure are more accurate than other models.

[0093] Figure 5 This is the visualization result of the visual question answering model. Figure 4 As shown, given an image and a question, the nested attention network model based on the selection graph structure can generate the corresponding answer.

[0094] This invention proposes a nested attention network with a selection graph structure for visual question answering. The selection graph structure is introduced to find the relationship between visual feature targets and targets based on the question, making the answer clearer. In addition, a nested attention module is added to mine the correlation between image regions. With the help of the proposed nested attention mechanism, the features generated by common attention are linearly operated with the original features of another modality to eliminate vector irrelevant information and effectively fuse visual features and question features. Extensive experiments on the VQA2.0 database show that the model has achieved good results in visual question answering. In future work, we will continue to explore how to better learn the semantics of images and question texts and effectively integrate them for answer reasoning.

[0095] Finally, the details of the above examples of the present invention are only examples for explaining the present invention. For those skilled in the art, any modification, improvement and replacement of the above embodiments should be included in the protection scope of the claims of the present invention.

Claims

1. A visual question answering method based on a nested attention network with a selective graph structure, characterized in that, the method includes the following steps: S1. Construct a nested attention network module to mine complete features according to the correlation between input image regions and between words; S2. Construct a selective graph structure module to construct a graph structure for image features, and then select edges and nodes for the graph structure based on question features; S3. Combine the network in S2 and the network in S3 to construct a nested attention network architecture based on the selective graph structure; The specific process of S1 is: Given image feature I and question feature T, input them into the nested attention model, and finally generate corresponding modal features after the training of the model; The middle part of the model is a classic co-attention architecture. First, for the given image feature I k and the question feature T k , the attention matrix A is obtained through cross multiplication, and then A is passed through a two-layer softmax to generate the attention matrix A T about the question and the attention matrix A I about the image. Finally, they are multiplied by the image feature and the question feature respectively to obtain the image feature I′ k containing bilateral information and the question feature T′ k ; this process can be represented by the following five formulas: A = T k × W R × I k T (1) A I = softmax(A) (2) A T = softmax(A T ) (3) T′ k = T k × A T (4) I′ k = I k × A I (5) where \(W\) in formula (1) R ∈R N×K represents the weight matrix, and the dimensions of the attention matrices are both \(d\times d\); The overall structure of the model enhances the correlation between feature regions in bilateral information; For Problem features, the problem features T' generated in the previous step k and the image features I k are concatenated, and then passed through two linear operations respectively. One of them is then input into a Sigmoid activation function, and the two features are multiplied to obtain the final problem feature T k ; This process is represented by the following four formulas: T ck = Concat(T′ k , I k ) (6) T lk = Linear(T ck ) (7) T lsk = Sigmoid(Linear(T ck )) (8) T K = T lk × T lsk (9) where linear is a Linear function with 1024 hidden units with ReLU non-linearity and dropout; Similarly, for the image feature, the image feature I' generated in the previous step k is concatenated with the problem feature T k through a concat operation, and then respectively passes through two linear operations. One of them is then input into a Sigmoid activation function, and the two features are multiplied to obtain the final image feature I k ; this process is represented by the following four formulas: I ck = Concat(I′ k , T k ) (10) I lk = Linear(I ck ) (11) I lsk = Sigmoid(Linear(I ck )) (12) I K = I lk × I lsk (13) where linear is a Linear function with 1024 hidden units with ReLU non-linearity and dropout; Image feature I k and problem feature T k are the outputs of the nested attention model, whose dimensions are consistent with those of the input features; then repeat the above operations with them as the inputs of the nested attention model. When k = n, it is the final output, i.e., image feature I n and problem feature T n .

2. The visual question answering method based on a nested attention network with a selective graph structure according to claim 1, characterized in that, the specific process of S2 is: Select the graph structure module, generate a graph structure based on the image features, and then find the relationship between the image feature targets and the targets according to the question to make the answer clearer; the graph structure is a directed multigraph, where each node corresponds to a scene entity, and the scene entity is an object associated with a bounding box and also an attribute of the object; each scene entity has a type corresponding to the predicted object or attribute label; the typed edges specify how the scene entities are related to each other; then, the graph structure is a set of ordered triples (s, p, o) of subject, predicate, and object; impose integrity for inverse relationships because for each (s, p -1 , o) ∈ SG; The state of the mapping space S is given by E×Q, where E is the node of the scene graph SG and Q represents the set of all questions; The mapping space S at the t-th node represents the current entity e t and the problem Q, denoted as St; thus, St = (e t , Q) represents the state for t ∈ N, where N is all nodes and S t ∈ S; The available action set from the mapping space St is represented by A St ; it contains all the outbound edges of node e t and their corresponding object nodes, represented by the following formula: A St = {(r,e) ∈ R × E: S t = (et, Q) ∧ (et, r, e) ∈ SG} (14) Simultaneously use A t ∈ A St represents the action performed by the mapping space at the t-th node; when creating the graph structure, node embeddings are passed through a multi-layer graph attention network; the multi-layer graph attention network extends it from the graph convolutional network through the self-attention mechanism, mimicking the convolutional operator on a regular grid, and forms entity embeddings by embedding the node features of its neighbors; therefore, the generated embeddings are context-aware, which makes nodes of the same type but different graph neighborhoods distinguishable; Extract features from the input image and question respectively, and then generate a graph structure from the image features. Use the frequency tuple H t =(H t-1 , A t-1 ) to represent the history of the mapping space, where t≥1, H 0 =hub, t = 0; The feature h is encoded by a multi-layer LSTM, as shown in the following formula: h t = LSTM(a t-1 ) (15) where a t-1 = [r t-1 , e t ∈ R 2d corresponding to using r t-1 to embed the previous action, e t represents embedding the edge and the target node into R d ; h and the question feature Q are simultaneously put into the ReLU activation function, and then after the operation of taking the action set At(), the image feature I is generated through the softmax function, as shown in the following formula: I = softmax(A t (W 2 ReLU (W 1 [h t Q]))) (16) Among them The rows of contain potential representations of all admissible actions; in addition, Q ∈ R d Encode the question Q; draw the action A according to the image feature I t =(r,e) ∈ A St ; Equations (2) and (3) yield a stochastic policy π θ , where θ represents the set of trainable parameters.

3. The visual question answering method based on a nested attention network with a selective graph structure according to claim 1, characterized in that, the specific process of S3 is: The visual question answering method based on the nested attention network with a selective graph structure includes a nested attention network module and a selective graph structure module.

Citation Information

Patent Citations

  • Visual question and answer method of original feature injection network based on composite attention

    CN112905819A