Visual question answering method based on knowledge graph
Patent Information
- Application Number
- CN202310316933.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-28
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-03-28
AI Technical Summary
[0044]本发明结合知识图谱和知识特征嵌入的相关技术,提出了一种基于知识图谱的视觉问答方案,极大地提高了视觉问答技术落实到实际应用中的可能性。
Smart Images

Figure CN116303969B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and natural language processing, primarily focusing on enhancing image feature representation by fusing external knowledge related to image content, thereby improving the performance of question-answering models. This technology has significant commercial value and can be applied to early childhood education, assistive devices for the blind, and other fields. Background Technology
[0002] With the continuous development of deep learning technology in computer vision and natural language processing, visual question answering has gradually emerged. The concept of visual question answering was first proposed by Antol et al. in 2015. Visual question answering can be defined as: given an image and a natural language question related to the image, the model needs to output a correct answer. Visual question answering can be applied to various fields such as online education, assisted navigation for the blind, and automated video surveillance querying.
[0003] Clearly, this is a multimodal problem combining computer vision and natural language processing techniques. Visual question answering tasks can be implemented in various ways, and the general algorithm can be divided into three steps: extracting features from the image, extracting features from the question, and combining the image and text to generate the answer. The main difference between algorithms lies in the third step, namely the method of combining the two input features. Simple methods for directly fusing image and text features include concatenation, tensors, inner product, and outer product. After integrating the features, a simple classifier, such as a linear classifier or a multilayer perceptron, is used.
[0004] However, simply fusing text and image features for classification cannot answer questions that require prior knowledge. Previous visual question answering methods that utilize human knowledge to predict answers have primarily focused on enhancing the representation of question features. In 2018, Narasimhan et al. used Long Short-Term Memory networks to predict factual relationship types from questions. In 2020, Garderes et al. used ConceptNet as a knowledge source and embedded entity information into the language representation. However, their methods all neglected the implicit knowledge related to image features.
[0005] Several methods utilize graph neural networks (GNNs) for reasoning in visual question-answering tasks. In 2020, Zhu et al. used a multimodal heterogeneous graph to describe a graph structure for reasoning and outputting answers, containing multiple layers of information corresponding to visual, semantic, and factual features. Also in 2020, Yu et al. decomposed their model into a series of memory-based reasoning steps, each executed by a graph-based reading, updating, and control module that performs parallel reasoning on visual and semantic information. However, while these methods consider the implicit external knowledge within the input image, they are inherently limited by GNNs, such as the computational complexity becoming extremely high as the number of nodes in the GNN increases.
[0006] To address the aforementioned issues, we consider detecting entity objects from the input image using a graph, retrieving implicit external information about these objects from a knowledge graph, and then using our proposed fusion method to enhance the image information in the next step, thereby embedding external knowledge related to the input image during the implicit reasoning process. Summary of the Invention
[0007] Purpose of the invention: To address the problem that conventional visual question answering methods based on external knowledge do not fully exploit the information of entities hidden in the input image, this invention enhances the feature representation of the image by fusing the features of the nodes corresponding to the entities in the knowledge graph and the entity features in the image through a multimodal bilinear pooling module.
[0008] 1. A visual question answering method based on knowledge graphs, characterized by comprising the following steps:
[0009] Step 1.1: Input the image data into a pre-trained fast object detection network to obtain region features and bounding box features in the image. Segment the input question text into words, obtaining words of length equal to the number of words in the text, and feed them into a pre-trained bidirectional encoder of the transformer to obtain the feature representation of the sentence.
[0010] Step 1.2: Construct the entity and attribute relationships in the external knowledge information into triples of the specified relationship type to build a knowledge graph containing 26,000 edges and 6,000 nodes.
[0011] Step 1.3: Process the knowledge graph data using a graph convolutional neural network. Initialize the features of each entity node using the bidirectional encoder feature representation of a pre-trained transformer. Then, after processing the graph data structure of the knowledge graph by the graph convolutional network, obtain the updated entity node representation. Use cosine similarity to calculate the similarity between entity nodes in the knowledge graph and entity nodes detected by the object detection network to filter out knowledge graph entity nodes related to the input image. Similarly, use cosine similarity to filter out entity nodes in the knowledge graph most relevant to the keywords involved in the question.
[0012] Step 1.4: The image feature representation of knowledge embedding is obtained by fusing features from entity nodes in the knowledge graph and image features through a multimodal compact bilinear pooling module.
[0013] Step 1.5: Input the extracted text features of the question and the entity features in the knowledge graph into multiple stacked transformer blocks to generate a text feature representation of knowledge embedding.
[0014] Step 1.6: Concatenate the image features embedded with the text features embedded with the knowledge to obtain a joint feature representation.
[0015] Step 1.7: Input the features obtained from Steps 1.1 and 1.2 into a parallel transformer module to align the image and text feature representations in the high-level semantic space, thereby obtaining text features for image attention and image features for text attention.
[0016] Step 1.8: Input the three feature streams from Step 1.6 and Step 1.7 into the feature aggregator to obtain a joint representation of the image, text and external knowledge, and then input the joint representation into the classifier for classification.
[0017] 2. The method for constructing the knowledge graph in step 1.2 is as follows:
[0018] Step 2.1: Filter the entity and attribute relationships from the scene graphs in the ConceptNet, WebChild, and VisualGenome datasets, constructing triples with the structure (entity, relation, attribute or entity) to build the knowledge graph. The triplet format, where Represents entities in a knowledge graph. To represent another entity or attribute, Represents the relationship between entities. The types of relationships between entities are represented, and entity features are represented using the same word embedding method as the question input.
[0019] Step 2.2: Select several frequently used entity relationship types in the visual question answering domain, specifying 8 relationship types from the ConceptNet dataset: "at...location", "used to...", "is...", "related to...", "owns...", "created by...", "can...", "has...property", 4 relationship types from the WebChild dataset: "has...material", "has...member", "below...", "at...location", and 5 relationship types from VisualGenome: "near...", "in...", "above...", "made by...", "owns...".
[0020] 3. The method for obtaining the knowledge graph node feature representation and filtering out nodes related to the input image and text in step 1.3 is as follows:
[0021] Step 3.1: Construct the knowledge graph as The triplet format, where Represents entities in a knowledge graph. Represents the relationship between entities. This represents the types of relationships between entities. For the first... Layer nodes Hidden state ,here Here, represents the dimension of each node. The propagation model for updating nodes can be defined as follows:
[0022]
[0023] in ( () is an element-level activation function. Representing relation type Next node neighborhood index set, It is a standardized constant. For a learnable weight matrix, This indicates that for the l-th layer node The hidden state, This is a learnable multilayer perceptron. After the node information in the neural network is propagated through multiple layers, the final node expression is obtained:
[0024]
[0025] In the backpropagation process of the network, cross-entropy loss is used as the loss function for nodes. The loss function is defined as follows:
[0026]
[0027] in It is the k-th entry in the network output of the i-th label node. Each represents its own truth label.
[0028] Step 3.2: By calculating the cosine similarity, the angular distance between words in the entity in the knowledge graph and words in the question and visual objects is obtained. The formula for calculating the similarity is as follows:
[0029]
[0030] Here, A represents the word vector representation of an entity, and B represents the word vector representation of an entity from an image or question. We select the entity node with the highest similarity score for the downstream feature fusion task.
[0031] 4. The method for obtaining the image features of knowledge embedding in step 1.4 is as follows:
[0032] Step 4.1: Obtain the node expression Then, the image features The i-th 2048-dimensional entity feature detected in A 512-dimensional feature with the same dimension as the text feature is generated by a transposed convolution matrix. .
[0033] Step 4.2: Put and By placing the two features into the Fast Fourier Transform space, they are fused through element-wise dot product to generate image features with embedded knowledge. The generation method is as follows:
[0034]
[0035] 5. The feature aggregation method in step 1.8 is as follows:
[0036] Step 5.1: Obtain image features embedded with external knowledge through steps 1.4 and 1.5. Sentence features embedded with external knowledge Then, the two features containing external knowledge information are concatenated to obtain the final external knowledge feature. .
[0037] Step 5.2: Obtain image features for text attention through step 1.7. Text features and image attention .
[0038] Step 5.3: Aggregation , and Then we generate a joint representation of the three features. The generation method is as follows:
[0039]
[0040] in, It is a learnable factor matrix. This represents the Hadamard product. The coefficients and weights in the network model are continuously updated, and the loss function is:
[0041]
[0042] Where N and L are the number of training samples and the number of candidate answers, respectively. For the true answer, To predict the answer.
[0043] Beneficial results of the present invention:
[0044] This invention combines knowledge graph and knowledge feature embedding technologies to propose a knowledge graph-based visual question answering scheme, which greatly improves the possibility of implementing visual question answering technology in practical applications. Attached Figure Description
[0045] Figure 1 This is a flowchart of the knowledge graph-based visual question answering model described in this invention;
[0046] Figure 2 This is a schematic diagram of the overall structure of the knowledge graph-based visual question answering model described in this invention.
[0047] Figure 3 This is a schematic diagram of the bilinear pooling module that integrates knowledge information and image information in this invention. Detailed Implementation
[0048] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.
[0049] like Figure 1 and Figure 2 As shown, this invention provides a knowledge graph-based visual question answering method, including object detection, text feature embedding, knowledge graph, graph neural network, bilinear pooling, and feature aggregation. The implementation method of this invention will be described in detail below.
[0050] Step 1: Feed the input image into the object detection network to obtain image features, and feed the text of the input question into the bidirectional encoder of the pre-trained transformer to obtain the question feature representation.
[0051] Step 1.1: Input the image data into a pre-trained fast object detection network to obtain region features and bounding box features in the image. Specifically, the object detection network detects 36 regions in the image, obtaining 36 object-level features, each with a dimension of 2048. The object detection network also outputs bounding box features, which contain the position and size information of the bounding box and have a dimension of 5.
[0052] Step 1.2: Segment the input question text to obtain a data structure of length n words. Then, use a pre-trained bidirectional encoder with a transformer to obtain the word embedding for each word. Each labeled word has a dimension of 768, and each question is divided into 16 labeled blocks, including BERT special labels such as [start], [segment], and [mask]. Finally, concatenate the features of each word and combine them with the word position vector encoding to obtain the final feature representation of the sentence.
[0053] Step 2: Construct and process the knowledge graph data, and select relevant entity nodes for downstream knowledge embedding.
[0054] Step 2.1: Filter the entity and attribute relationships from the scene graphs in the ConceptNet, WebChild, and VisualGenome datasets, constructing triples with the structure (entity, relationship, attribute or entity). Specify 8 relationship types from the ConceptNet dataset: "located in...", "used to...", "is...", "related to...", "owns...", "created by...", "can...", "has..." property; 4 relationship types from the WebChild dataset: "has..." substance, "has..." member, "below...", "located in..."; and 5 relationship types from VisualGenome: "near...", "in...", "above...", "made by...", "owns...". The final knowledge graph contains 26,000 edges and 6,000 nodes. After processing the knowledge graph through a graph convolutional neural network, updated entity node information is obtained. The propagation model for updating nodes in the knowledge graph can be defined as follows:
[0055]
[0056] Step 2.2: By calculating the cosine similarity, the angular distance between the words in the entity in the knowledge graph and the words in the question and the corresponding words of the visual object is obtained. The corresponding entity node with the highest score in the knowledge graph is selected for downstream feature fusion.
[0057] Step 3: The entity nodes obtained from the knowledge graph are used as external knowledge features to embed into the image feature representation and text feature representation. Then, these two features are concatenated to obtain the final overall feature of the knowledge embedding.
[0058] Step 3.1: As Figure 3 As shown, knowledge features and image features are fused using a bilinear pooling module. Specifically, this involves obtaining image representations from a pre-trained object detection network. and obtaining node information from knowledge graphs Subsequently, in order to utilize the implicit prior information in the image, a multimodal bilinear pooling module is used to merge the node information in the knowledge graph into the image to obtain the feature representation of the external knowledge embedding. .
[0059] Step 3.2: Input the extracted text features of the question and the entity features from the knowledge graph into multiple stacked transformer blocks to generate a text feature representation of knowledge embedding. .
[0060] Step 3.3: Concatenate image features embedded with knowledge and just embedded text features The final knowledge embedding representation is obtained. .
[0061] Step 4: Input the features obtained from Steps 1.1 and 1.2 into a parallel transformer module to align the image and text feature representations in a high-level semantic space, thereby obtaining the image features for text attention. Text features and image attention
[0062] Step 5: Aggregate the three features from Step 3 and Step 4, respectively , , The final joint feature representation is obtained and used for the classification task. The final joint feature representation is:
[0063]
[0064] Overall, the neural network used for visual object detection was pre-trained on the COCO dataset before formal training, with the prediction confidence of the detection boxes set to 0.5. The overall neural network model was trained for 20 epochs, with a batch size of 512. The BertAdam optimizer was used during training, with an initial learning rate of 4. .
[0065] The detailed descriptions listed above are merely specific descriptions of feasible embodiments of the present invention, and are not intended to limit the scope of protection of the present invention. All equivalent embodiments or modifications made without departing from the spirit of the present invention should be included within the scope of protection of the present invention.
Claims
1. A visual question answering method based on knowledge graphs, characterized in that, Includes the following steps: Step 1.1: Input the image data into the pre-trained fast object detection network to obtain the region features and detection box features in the image. Segment the text of the input question to obtain words with a length equal to the number of words in the text. Then, feed the words into the bidirectional encoder of the pre-trained transformer to obtain the feature representation of the sentence. Step 1.2: Construct the entity and attribute relationships in the external knowledge information into triples of the specified relationship type to build a knowledge graph containing 26,000 edges and 6,000 nodes; Step 1.3: Process the knowledge graph data through a graph convolutional neural network. Initialize the features of each entity node with the bidirectional encoder feature representation of the pre-trained transformer. Then, after the graph convolutional network processes the graph data structure of the knowledge graph, the updated entity node representation is obtained. The cosine similarity is used to calculate the similarity between the entity nodes in the knowledge graph and the entity nodes detected by the object detection network to filter out the knowledge graph entity nodes related to the input image. Similarly, the cosine similarity is used to filter out the entity nodes in the knowledge graph that are most relevant to the keywords involved in the question. Step 1.4: Fuse features from entity nodes in the knowledge graph and image features using a multimodal compact bilinear pooling module to obtain the image feature representation of knowledge embedding; Step 1.5: Input the extracted text features of the question and the entity features in the knowledge graph into multiple stacked transformer blocks to generate a text feature representation of knowledge embedding; Step 1.6: Concatenate the image features embedded with the text features embedded with the knowledge to obtain a joint feature representation; Step 1.7: Input the features obtained from Steps 1.1 and 1.2 into a parallel transformer module to align the image and text feature representations in the high-level semantic space, thereby obtaining text features for image attention and image features for text attention. Step 1.8: Input the three feature streams from Step 1.6 and Step 1.7 into the feature aggregator to obtain a joint representation of the image, text and external knowledge, and then input the joint representation into the classifier for classification.
2. The visual question answering method based on knowledge graphs according to claim 1, characterized in that, The method for constructing the knowledge graph in step 1.2 is as follows: Step 2.1: Filter the entity and attribute relationships of scene graphs from the ConceptNet, WebChild, and VisualGenome datasets, constructing them into triples with the structure entity, relation, attribute, or entity, thus building the knowledge graph as follows. The triplet format, where Represents entities in a knowledge graph. To represent another entity or attribute, Represents the relationship between entities. The types of relationships between entities are represented, and entity features are represented using the same word embedding method as the question input. Step 2.2: Select several frequently used entity relationship types in the visual question answering domain, specifying 8 relationship types from the ConceptNet dataset: "at...location", "used to...", "is...", "related to...", "owns...", "created by...", "can...", "has...property", 4 relationship types from the WebChild dataset: "has...substance", "has...member", "below...", "at...location", and 5 relationship types from VisualGenome: "near...", "in...", "above...", "made by...", "owns...".
3. The visual question answering method based on knowledge graphs according to claim 1, characterized in that, The method for obtaining the knowledge graph node feature representation and filtering out nodes related to the input image and text in step 1.3 is as follows: Step 3.1: Construct the knowledge graph as The triplet format, where Represents entities in a knowledge graph. Represents the relationship between entities. Represents the types of relationships between entities, for the first Layer nodes Hidden state ,here This refers to the dimension of each node. The propagation model for updating nodes is defined as follows: in ( () is an element-level activation function. Representing relation type Next node neighborhood index set, It is a standardized constant. For a learnable weight matrix, This indicates that for the l-th layer node The hidden state, For a learnable multilayer perceptron, after the node information in the neural network is propagated through multiple layers, the final node expression is obtained: In the backpropagation process of the network, cross-entropy loss is used as the loss function for nodes. The loss function is defined as follows: in It is the k-th entry in the network output of the i-th label node. Indicate their respective truth labels; Step 3.2: By calculating the cosine similarity, the angular distance between words in the entity in the knowledge graph and words in the question and visual objects is obtained. The formula for calculating the similarity is as follows: in, A word vector representation of an entity. A word vector representation of an entity from an image or question. Word vectors The Components of each dimension Word vectors The Components of each dimension The dimension of the word vector is used to select the entity node with the highest similarity score for the downstream feature fusion task.
4. The visual question answering method based on knowledge graphs according to claim 1, characterized in that, The method for obtaining the image features of knowledge embedding in step 1.4 is as follows: Step 4.1: Obtain the node expression Then, the image features The i-th 2048-dimensional entity feature detected in A 512-dimensional feature with the same dimension as the text feature is generated by a transposed convolution matrix. ,in For the updated node feature sequence, For the corresponding node feature vector, Dimensions of node features The image representation feature sequence obtained by the object detection network. For entity feature vectors in the image, The dimension of the image features; Step 4.2: Put and By placing the two features into the Fast Fourier Transform space, they are fused through element-wise dot product to generate image features with embedded knowledge. The generation method is as follows: = in, Represents the Fast Fourier Transform. This represents the inverse fast Fourier transform. This represents element-wise dot product operations.
5. The visual question answering method based on knowledge graphs according to claim 1, characterized in that, The feature aggregation method in step 1.8 is as follows: Step 5.1: Obtain image features embedded with external knowledge through steps 1.4 and 1.
5. Sentence features embedded with external knowledge Then, the two features containing external knowledge information are concatenated to obtain the final external knowledge feature. ; Step 5.2: Obtain image features for text attention through step 1.
7. Text features and image attention ; Step 5.3: Aggregation , and A joint representation of the three features is then generated. ,in Image features representing text attention Textual features representing image attention Represents the final external knowledge characteristics after fusion. , , These represent the lengths of the three corresponding feature sequences mentioned above. and The dimension representing the feature is generated as follows: in, It is a learnable factor matrix. Representing the Hadamard product, and continuously updating the coefficients and weights in the network model, the loss function is: Where N and L are the number of training samples and the number of candidate answers, respectively. For the true answer, To predict the answer.
Citation Information
Patent Citations
Visual question-answering method based on fusion of fine-grained image features and external knowledge
CN112100346A
Visual question and answer method based on cross-modal pre-training feature enhancement
CN114663677A