A knowledge graph visual question answering method based on dual-process cognitive theory
By introducing dual-process cognitive theory in the visual question-and-answer system, using multimodal Transformer network and graph neural network to reason on the knowledge graph, the existing VQA methods are solved to solve the problems that require common sense and low efficiency problems, and achieve more efficient and accurate visual question-and-answer performance.
Patent Information
- Application Number
- CN202110374169.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-07
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-04-07
AI Technical Summary
Existing visual question-and-answer (VQA) methods are difficult to answer questions that need to be answered in combination with common sense, and existing methods are inefficient in dealing with synonyms and homonyms, and cannot effectively utilize structural information from external knowledge bases.
Using a knowledge graph visual question-and-answer method based on dual-process cognition theory, a multimodal Transformer network (System 1) is used to capture the complex relationship between the problem and the image, and output joint representations. Then, use graph neural networks (System 2) to perform inference on fact graphs and semantic graphs, combining the attention mechanisms at the node level and path level to conduct evidence aggregation and answer prediction.
It improves the performance of visual question-and-answer, can more effectively deal with common sense questions, reduce noise information, and improve the accuracy and efficiency of answer prediction.
Smart Images

Figure CN115186072B_ABST
Abstract
Description
Technical Field
[0001] The present invention designs a knowledge graph visual question answering method based on dual-process cognitive theory, which belongs to the intersection of natural language processing and computer vision fields. Background Art
[0002] In recent years, enabling intelligent agents to understand the world by analyzing visual and language information has been a hot research topic in the combination of computer vision and natural language processing technology. Related research has promoted the development of many applications, such as visual question answering (VQA), image indexing, and image description. Among them, VQA is a challenging task that requires the model to answer arbitrary questions based on a given image. In order to promote the development of VQA technology, relevant scholars have done a lot of preliminary research in recent years and have made great progress. However, existing VQA methods focus on answering questions based on the content of the picture, and cannot answer some questions that require common sense to answer.
[0003] To promote the development of this field, Wang et al. proposed the fact-based visual question answering (FVQA) task (P.Wang, Q.Wu, C.Shen, A.Dick and A.van denHengel, "FVQA:Fact-Based Visual Question Answering," in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol.40, no.10, pp.2413-2427, 1Oct.2018, doi:10.1109 / TPAMI.2017.2754246). At the same time, they also released a new dataset that provides additional supporting facts for each question-answer pair and requires the model to answer questions by jointly analyzing images and external knowledge. Wang et al.'s study first parses the sentence and then maps it to the knowledge graph, and then uses keyword matching to find the correct answer. This method has obvious flaws and becomes ineffective when there is no obvious visual concept mentioned in the question or when there are synonyms and homographs. Therefore, some scholars subsequently proposed a retrieval method based on semantic learning, which projects the image-question-visual concept and candidate facts into a learned embedding space and finds the supporting facts by calculating the corresponding distance. However, this method evaluates one fact node at a time, so it is inefficient. In addition, this method cannot utilize the structural information of the external knowledge base. To solve this problem, Narasimhan et al. proposed a graph inference-based method to select the answer by reasoning on the entire graph (Narasimhan M, Schwing AG. Straight to the facts: Learning knowledge base retrieval for factual visual question answering [C] / / Proceedings of the European conference on computer vision (ECCV). 2018: 451-468.). This method constructs an entity graph, in which each node is represented by a connection between the entity, image, and question representation, and then uses a graph convolutional network to aggregate the messages to obtain the corresponding node update features, and finally predicts the answer based on the updated node features. Since the question only focuses on part of the visual content, this method inevitably introduces noise information.
[0004] For humans, when given a picture and a question, inferring the answer can be divided into two steps: (1) the brain quickly obtains the content presented in the image and the visual information concerned by the question by analyzing the image and the question; (2) based on the output of the first step and combined with the knowledge stored in the brain, a deep analysis is performed to find the correct answer. This process is called the dual-process theory in cognitive science. The dual-process cognitive theory holds that the human brain first perceives external input information through an implicit, unconscious process performed by a system called System 1. This information is then sent to a system called System 2 to perform an explicit, conscious, and controllable reasoning process. This reasoning process performs sequential inferences in working memory, which is a slow process but is a unique attribute of humans as advanced intelligent agents. From this perspective, the FVQA problem can be solved using the dual-process cognitive theory—System 1 quickly retrieves information from images and questions, and System 2 combines external knowledge for deep reasoning to find the correct answer.
[0005] Inspired by the dual-process cognitive theory, the present invention proposes a new framework based on a dual-process cognitive system to solve the FVQA problem. Specifically, system 1 in the framework of the present invention is implemented by a multimodal Transformer network, which uses a cross-attention mechanism to capture the complex relationship between the question and the image. System 1 outputs a joint representation of the image and the question. For system 2, the present invention uses a graph neural network (GNN) to reason on two external knowledge graphs (factual graph and semantic graph) to find the correct answer. System 2 first performs intra-modal evidence selection and aggregation, and aggregates evidence information related to the question from the two knowledge graphs respectively; then it performs cross-modal selection and aggregates the evidence in the factual graph into the semantic graph to assist in better inference of the answer. In the reasoning process, the present invention proposes a dual attention mechanism of node-level attention and path-level attention, which captures valuable information from key nodes and paths, makes the reasoning process more reasonable, and further improves the performance of visual question answering based on knowledge graphs. Summary of the invention
[0006] Inspired by the dual-process cognitive theory, this paper proposes a new framework for solving FVQA problems using a dual-process cognitive system. Specifically, system 1 in the framework of the present invention is implemented by a multimodal Transformer network, which uses a cross-attention mechanism to capture the complex relationship between questions and images and outputs a joint representation of a question-image. For system 2, a GCN network is used to reason on two external knowledge graphs (factual graph and semantic graph) to find the correct answer.
[0007] The present invention achieves the above-mentioned purpose through the following technical solutions:
[0008] 1. The knowledge graph visual question answering framework based on dual-process cognitive theory described in the invention is as follows Figure 1 As shown, it consists of two parts: a coordinated perception module (system 1) and an explicit reasoning module (system 2). The specific reasoning process of this framework includes the following steps:
[0009] (1) The text pre-training model BERT and the image pre-training model Faster-RCNN are used to extract features from the input text and image, respectively. For each question, the [CLS] and [SEP] flags are added at the beginning and end of the question, respectively, and then sent to the BERT model for feature extraction. For each image, 36 target regions are extracted, each of which contains the visual features of a target and the spatiotemporal position feature information of the target.
[0010] (2) The image and text features extracted in step (1) are sent to a two-stream Transformer network to learn the joint representation of image and text. One single-stream Transformer network is used to learn the image-guided question representation, while the other single-stream Transformer network is used to learn the question-guided image representation. Finally, the outputs of the two two-stream Transformer networks are average-pooled and further multiplied to obtain the joint representation.
[0011] (3) Construction of factual graph and semantic graph. For each question-image pair, a factual graph and a semantic graph are constructed respectively. The factual graph is selected from the external knowledge base based on sentence-level semantic matching, while the semantic graph is obtained by first semantically describing the image and then semantically parsing the generated sentence.
[0012] (4) Evidence aggregation based on graph reasoning: First, for the fact graph and semantic graph, two graph reasoning networks are used based on the attention mechanism to aggregate evidence information from the knowledge graph and fact graph respectively. Then, a cross-modal reasoning network is used to aggregate evidence information related to the question from the semantic graph into the knowledge graph.
[0013] (5) Answer prediction: The joint representation of the question and the image is multiplied by the feature vector of each node in the fact graph to obtain the semantic relevance score between each node and the question. Finally, the relevance score is sent to a Sigmoid layer to predict the corresponding answer.
[0014] Specifically, in step (1), the BERT pre-trained model is first used to initialize the vector of the word in the input question, where BERT uses the bert-base-uncased version. After the vector initialization, the feature vector C = [c0, c1, ..., c n], the feature vector dimension of each word is 768. Then the Faster-RCNN pre-trained model is used to extract 36 target regions of the input image, each of which contains the appearance features of a target. Spatial location characteristics of the target In order to capture both visual features and spatial features, the present invention first projects the appearance features and spatial position features into the same dimension (768 dimensions in the present invention), and then averages the appearance visual features and spatial position features to obtain the feature representation of each target object.
[0015] In step (2), a cross-modal Transformer is used to align the relationship between the two modalities and learn the question representation of the image and the image representation of the question, as shown in Figure 1.
[0016] The question feature vector C and the image feature vector V are obtained through step (1). The question feature vector C and the image feature vector V are sent to a two-stream Transformer network to learn the complex interaction between the question and the image. One single-stream Transformer network is used to learn the question representation guided by the image, and the other single-stream Transformer network is used to learn the image representation guided by the question. In the learning of the question representation guided by the image, the question feature vector is used as the query vector, and the image feature vector is used as the key and value vectors. The calculation formula of the dependency relationship between the image and the question is as follows:
[0017]
[0018] W′ in formula (1) * Represents the parameter matrix of the model to be learned; in the picture representation learning guided by the question, the picture feature vector is used as the query vector, and the question feature vector is used as the key and value vector. The calculation formula of the dependency relationship between the two is as follows:
[0019]
[0020] W″ in formula (1) * The parameter matrix to be learned by the representation model; after multiple layers (specifically 9 layers in the present invention) of joint representation of images and texts, the features of the [CLS] bit of the text sequence are used as the joint representation features of the entire question-image.
[0021] In step (3), the construction of the fact graph and semantic graph is as follows: first, for each question-image pair, the input question and the label of the object detected in the image are concatenated in sequence to obtain a question-image instance set; each triple in the external knowledge base is converted (the conversion method is to concatenate the head entity, the relationship and the tail entity in sequence) into a natural language processing sentence to obtain the corresponding fact instance set. Second, the question-image instance set and the fact instance set are encoded using the pre-trained sentence encoder Universal-sentence-encoder to obtain the corresponding instance representation, and then the feature representation of each instance object in the question-image instance set is calculated in sequence with the feature representation of the fact instance object to obtain the corresponding association score. Finally, all instance objects are sorted according to the cosine similarity score, and the top 10 instances with the highest scores are obtained as candidate supporting facts. The third step is to construct a fact graph based on the alternative supporting facts retrieved. The nodes in the graph are entities in the knowledge base, and the edges are the relationships between two entities. After the fact graph is constructed, BERT is used to initialize the nodes and edges accordingly, where the initialization representation of the nodes and edges is the average of the word embeddings of all words in the nodes and edges.
[0022] In step (4), the evidence aggregation based on graph reasoning includes evidence aggregation within the modality and evidence aggregation between modalities:
[0023] When aggregating evidence within a modality, the dual-level attention mechanism proposed in this invention is first used to select and aggregate features: node-level attention and path-level attention. In the node-level attention calculation process, the attention score of each node in the graph and the joint representation of the question-image is first calculated. The calculation process is as follows:
[0024]
[0025] Then the attention score is multiplied by the initial feature vector of each node in the graph to obtain the node feature vector guided by the picture-question. In the process of calculating the path node attention, the main focus is on which path is more important to the reasoning process, where each path is defined as the path composed of all nodes and edges directly connected to the target node, which is defined as follows:
[0026] φ ij =(v i ,r ij ,v j ) (4)
[0027] where v i , r ij , v jThe separate tables show the representation of the head node, the representation of the relationship, and the feature representation of the tail node in the fact graph. After obtaining the path representation, the attention calculation process at the path level is as follows:
[0028]
[0029] Then, the features are aggregated from neighboring nodes according to the message propagation network. The calculation formula for the feature aggregation process of neighboring nodes is as follows:
[0030]
[0031] Finally, the features of the neighboring nodes are fused with the features of the target node to further update the features of the target node. In order to prevent the features of the neighboring nodes from over-updating the initial features of the node, a gating mechanism is designed to control the proportion of the features of the neighboring nodes and the original features of the target node. The feature update process of the target node is calculated as follows:
[0032]
[0033]
[0034] The process of evidence selection in semantic graph is the same as that in fact graph, so it will not be repeated here;
[0035] Inter-modal evidence aggregation, when performing inter-modal evidence aggregation, first, under the guidance of the question, the attention weight coefficient of each node in the fact graph and each node in the semantic graph is calculated, and finally, the relevant features in the semantic graph are obtained by weighted summing of each node in the semantic graph according to the attention weight coefficient. The relevant process calculation process is as follows:
[0036]
[0037]
[0038] Finally, the feature vector in the semantic graph is fused with the feature vector of the original node in the fact graph to obtain the updated features after cross-modal fusion. In order to prevent the features from the semantic graph from over-updating the features of the fact node, a corresponding gating mechanism is designed to control the proportion of the two different modal features. The specific process calculation formula is as follows:
[0039]
[0040]
[0041] Finally, the updated features are used to infer the answer. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1It is the main framework of the network model proposed in this invention.
[0043] Figure 2 It is the structure of the cross-modal Transformer network. DETAILED DESCRIPTION
[0044] The present invention will be further described below in conjunction with the accompanying drawings:
[0045] Figure 1 is the structure of the entire network, which consists of two parts, namely the coordinated perception module (system 1) and the explicit reasoning module (system 2). System 1 in the framework is implemented by a multimodal Transformer network, which uses a cross-attention mechanism to capture the complex relationship between questions and images. System 1 outputs a joint representation of an image and a question. For System 2, the present invention uses a Graph Neural Network (GNN) to reason on two external knowledge graphs (factual graph and semantic graph) to find the correct answer.
[0046] Figure 2 It is a cross-modal Transformer framework that feeds the extracted image and text features into a two-stream Transformer network to learn the joint representation of images and texts. One single-stream Transformer network is used to learn the question representation guided by images, while the other single-stream Transformer model is used to learn the image representation guided by questions. Finally, the outputs of the two two-stream Transformer networks are average pooled and further multiplied to obtain the joint representation.
[0047] Table 1 and Table 2 are the experimental results of the present invention on the public datasets FVQA and OK-VQA. The experiments show that the proposed model achieves the best result in terms of the comprehensive evaluation index F1 value compared with the best existing model.
[0048] Table 1 Experimental comparison results of the network model of the present invention on the FVQA dataset and other existing models
[0049]
[0050] Table 2 Experimental comparison results of the network model of the present invention on the OK-VQA dataset and other existing models
[0051]
[0052] The above embodiments are only preferred embodiments of the present invention and are not limitations of the technical solutions of the present invention. Any technical solution that can be implemented on the basis of the above embodiments without creative work should be deemed to fall within the scope of protection of the patent of the present invention.
Claims
1. A knowledge graph visual question answering method based on dual-process cognitive theory, characterized in that: The following steps are involved: (1) Use the text pre-training model BERT and the object detection model Faster-RCNN to extract features from the input text and images. For text, add [CLS] and [SEP] tags at the beginning and end of each sentence, and then send them to the BERT model for feature extraction. For each image, extract 36 target regions, each of which contains the appearance visual features of an object and the spatial position feature information of the object in the image. (2) The extracted image and text features are fed into a two-stream Transformer network to learn the joint representation of image and text. One single-stream Transformer network is used to learn the question representation guided by the image, while the other single-stream Transformer model is used to learn the image representation guided by the question. Finally, the outputs of the two two-stream Transformer networks are average-pooled and further multiplied to obtain the joint representation of question and image. (3) Construction of factual graph and semantic graph. For each question-image pair, a factual graph and a semantic graph are constructed respectively. The factual graph is constructed by retrieving alternative supporting facts from an external knowledge base based on sentence-level semantic matching, while the semantic graph is constructed by first semantically describing the image and then semantically parsing the generated sentences. (4) Evidence aggregation based on graph reasoning: First, evidence aggregation is performed on the fact graph and semantic graph within the modality. The attention score of each node in the graph and the joint representation of the question-image is calculated through node-level attention. Then, the attention score is multiplied by the initial feature vector of each node in the graph to obtain the node feature vector guided by the image-question. Then, the path node attention is used to calculate which path is more important to the reasoning process. Then, features are aggregated from neighbor nodes according to the message propagation network. Finally, the features of the neighbor nodes are fused with the features of the target node to further update the features of the target node. In order to prevent the features of the neighbor nodes from over-updating the initial features of the node, a gating mechanism is designed to control the proportion of the features of the neighbor nodes to the original features of the target node. Then, evidence aggregation between modalities is performed. First, under the guidance of the question, the attention weight coefficient of each node in the factual graph and each node in the semantic graph is calculated. Finally, the weighted sum of each node in the semantic graph is performed according to the attention weight coefficient to obtain the relevant features in the semantic graph. Finally, the feature vector in the semantic graph is fused with the feature vector of the original node in the factual graph to obtain the updated features after cross-modal fusion. In order to prevent the features from the semantic graph from over-updating the features of the factual nodes, a corresponding gating mechanism is designed to control the proportion of the two different modal features, and the updated features are used to infer the answer. (5) Answer prediction: The joint representation of the question and the image is multiplied by the feature vector of each node in the fact graph to obtain the semantic matching score between each node and the question. Finally, the matching score is sent to a Sigmoid layer to predict the corresponding answer.
2. The method according to claim 1, characterized in that (2) The joint representation learning method of the question and the image is as follows: Given a question feature vector C and an image feature vector V, the question feature vector C and the image feature vector V are fed into a two-stream Transformer network to learn the complex interaction between the question and the image. One single-stream Transformer network is used to learn the question representation guided by the image, and the other single-stream Transformer network is used to learn the image representation guided by the question. In the learning of the question representation guided by the image, the question feature vector is used as the query vector, and the image feature vector is used as the key and value vector. The calculation formula of the dependency relationship between the image and the question is as follows: W in formula (1) * ' represents the parameter matrix of the model to be learned, d k Represents feature dimension; In the problem-guided image representation learning, the image feature vector is used as the query vector, and the question feature vector is used as the key and value vector. The calculation formula of the dependency relationship between the two is as follows: W in formula (2) * ” also represents the parameter matrix to be learned by the model; after multiple layers of joint representation of images and texts, the features of the [CLS] bit of the text sequence are used as the joint representation features of the entire question-image.
3. The method according to claim 1, characterized in that The specific process of constructing the fact graph in step (3) is as follows: (a) For each question-image pair, we first concatenate the input question and the labels of the objects detected in the image in sequence to obtain a set of question-image instances. Then, we convert each triple in the external knowledge base into a natural language sentence to obtain a set of corresponding fact instances. The conversion method is to concatenate the head entity, the relation, and the tail entity in sequence. (b) Encode the question-image instance set and the fact instance set using the pre-trained sentence encoder Universal-sentence-encoder to obtain the corresponding instance representation; (c) Finally, the feature representation of each instance object in the question-image instance set is sequentially compared with the feature representation of the fact instance object to calculate the cosine similarity to obtain the corresponding association score. Finally, all instance objects are sorted according to the cosine similarity score, and the top 10 instances with the highest scores are obtained as candidate supporting facts. (d) Construct a fact graph based on the retrieved alternative supporting facts. The nodes in the graph are entities in the knowledge base, and the edges are the relationships between two entities. After the fact graph is constructed, the BERT model is used to initialize the nodes and edges accordingly, where the initialization representation of the nodes and edges is the average of the word embeddings of all words in the nodes and edges.
4. The method according to claim 1, characterized in that The evidence aggregation step in (4) is divided into evidence aggregation within a modality and evidence aggregation between modalities: (a) When aggregating evidence within a modality, we first use a dual-level attention network: node-level attention and path-level attention for feature selection and aggregation. In the node-level attention calculation process, we first calculate the attention score of each node in the graph and the joint representation of the question and image. The calculation process is as follows: W and b in the above formula represent the parameter matrix and bias of the model to be learned. The different subscripts of W and b in the following formula represent the parameter matrix and bias of the model to be learned in the calculation process, which will not be repeated. Problem - Joint feature representation H of images vl , the feature representation of each ROI is V i , H in the following formula vl 、V i They all have the same meaning and will not be repeated. After obtaining the attention weight coefficient, the attention score is multiplied by the initial feature vector of each node in the graph to obtain the node feature representation guided by the picture-question. In the process of calculating the path node attention, we mainly focus on which path is more important to the reasoning process, where each path is defined as the path composed of all nodes and edges directly connected to the target node, which is defined as follows: φ ij =(v i ,r ij ,v j ) (4) where v i , r ij , v j The separate tables show the feature representation of the head node, the feature representation of the relationship, and the feature representation of the tail node in the fact graph. After obtaining the path representation, the attention calculation process at the path level is as follows: Where o represents the Hadamard product, and then the message propagation mechanism is used to aggregate features from neighboring nodes. The feature aggregation process of neighboring nodes is shown in the following formula: v i '=where β ij (Table W4 shows [v question' j ; Title ij -] Figure + b 5) The attention score of the joint feature representation, and finally the features of the neighbor nodes are fused with the features of the target node to further update the features of the target node. In order to prevent the excessive update of the node features by the features of the neighbor nodes, a gating mechanism is designed to control the proportion of the neighbor node features and the original features of the target node. The feature update process of the entire target node is shown in the following formula: Where σ represents the sigmoid function, is the feature of the node before updating, It is the updated feature of the node and also serves as the input of the next layer of graph convolution. The process of evidence aggregation in the semantic graph is the same as that in the fact graph, so it will not be repeated here. (b) Inter-modal evidence aggregation. When performing inter-modal evidence aggregation, first, under the guidance of the question, the attention weight coefficient of each node in the fact graph and each node in the semantic graph is calculated. Finally, the relevant features in the semantic graph are obtained by weighted summing the features of each node in the semantic graph according to the attention weight coefficient. The relevant process calculation process is as follows: Where T is the number of nodes in the semantic graph; is the node feature obtained after the intra-modal feature aggregation in the semantic graph, It is the node feature obtained after the intra-modal feature aggregation in the fact graph; are complementary features obtained from the semantic graph; finally, the feature vector aggregated from the semantic graph is fused with the feature vector of the node in the fact graph to obtain the new feature after cross-modal fusion. In order to prevent the excessive update of the feature of the fact graph node from the feature from the semantic graph, a corresponding gating mechanism is designed to control the proportion of the two different modal features. The specific process calculation formula is as follows: Where V i FS It is the final feature representation of the nodes in the fact graph, and finally the updated features are used to infer the answer.
Citation Information
Patent Citations
End to end network model for high resolution image segmentation
CN110809784A
Structured Knowledge Modeling, Extraction and Localization from Images
US20170132498A1