A Visual Question Answering Method and Device Based on Bipartite Graph Matching

By constructing keyword diagrams and scene diagrams, using graph attention networks and Hungarian algorithms, combined with multimodal fusion strategy, the problems of focus and language bias in the key area of image in the visual question-and-answer method are solved, and the robustness and accuracy of the model are improved.

CN120198761BActive Publication Date: 2025-07-18YANAN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510676797.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-07-18
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

Existing visual question-and-answer methods are difficult to accurately focus key areas in images, and are affected by language bias, resulting in insufficient robustness of the model.

Method used

Using a two-part graph matching method, by constructing keyword graphs and scene graphs, using graph attention network to encode interaction information between nodes, combining Hungarian algorithms to determine the main effect and negative node characteristics of the scene graph, and generating a predicted probability distribution through a multimodal fusion strategy.

Benefits of technology

It improves the accuracy and applicability of the visual question-and-answer model, accurately focuses on key areas of the image, reduces the influence of language bias, and enhances the robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198761B_ABST
    Figure CN120198761B_ABST
Patent Text Reader

Abstract

The present application discloses a visual question answering method and device based on bipartite graph matching, which relates to the technical field of image processing. The method includes: constructing a question feature vector, a keyword graph, and a scene graph; using a graph attention network to determine keyword graph node features and scene graph node features; based on the keyword graph node features and the scene graph node features, using the Hungarian algorithm to determine scene graph main effect node features and scene graph negative effect node features; using a multimodal fusion strategy to fuse the question feature vector with the scene graph node features, the scene graph main effect node features, and the scene graph negative effect node features respectively to obtain multiple joint features; inputting the multiple question feature vectors into a classifier to obtain a predicted probability distribution of the question. The present application improves the robustness of the visual question answering model by determining the scene graph main effect node features and the scene graph negative effect node features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technologies, and particularly to a visual question answering method and device based on bipartite graph matching. Background Art

[0002] Visual Question Answering (VQA) technology combines computer vision and natural language processing, aiming to enable a computer to output an answer that not only conforms to natural language rules but also has reasonable content based on a given picture and question. This technology requires the machine to have a deep understanding of the picture content and the meaning and intention of the question, and its application scope is wide, including online education, assisting visually impaired people in navigation, video surveillance query, etc. In the field of visual question answering research, most system frameworks consist of an image encoder, a question encoder, multimodal fusion, and an answer prediction module.

[0003] Currently, the research methods in the field of visual question answering mainly focus on constructing a cross-attention mechanism between images and questions to improve the complexity and processing ability of the model. In addition, some research has begun to explore high-level representations of images, especially using object detectors and graph-based structures to better understand the relationships between objects in the image. However, these methods mostly use neural network encoding technologies such as Gate Recurrent Unit (GRU), Long Short-Term Memory (LSTM), or Transformer models when processing questions. These neural network encoding technologies are difficult to distinguish the keywords in the question and the semantic biases in the scene graph, resulting in the inability to accurately focus on the effective regions of the image for the question. This limits the universality of these methods, indicating that there are still limitations in the existing technologies for understanding and processing visual question answering tasks. Summary of the Invention

[0004] The purpose of the present application is to provide a visual question answering method and device based on bipartite graph matching, which can focus on the most relevant regions of the image for the question itself and weaken the negative impact caused by language bias, thereby improving the robustness of the visual question answering model.

[0005] To achieve the above purpose, the present application provides the following solutions:

[0006] In a first aspect, the present application provides a visual question answering method based on bipartite graph matching, including:

[0007] Obtain a question and a picture;

[0008] Construct a keyword graph based on the question;

[0009] Use a graph attention network to encode the interaction information between nodes in the keyword graph to obtain keyword graph node features;

[0010] Construct a scene graph based on the said picture;

[0011] Use a graph attention network to encode the interaction information between nodes in the scene graph to obtain scene graph node features;

[0012] Based on the keyword graph node features and the scene graph node features, use the Hungarian algorithm to determine the scene graph main effect node features and the scene graph negative effect node features;

[0013] Construct a problem feature vector based on the said problem;

[0014] Use a multimodal fusion strategy to fuse the problem feature vector with the scene graph node features, the scene graph main effect node features, and the scene graph negative effect node features respectively to obtain multiple joint features;

[0015] Input the multiple joint features into a classifier to obtain the predicted probability distribution of the said problem.

[0016] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above-mentioned visual question answering method based on bipartite graph matching.

[0017] According to the specific embodiments provided by the present application, the present application discloses the following technical effects:

[0018] The present application provides a visual question answering method and device based on bipartite graph matching, which uses a pre-trained KeyBERT model to extract question keywords and uses a Glove model for vectorization; constructs a keyword graph and a scene graph, and encodes the interaction information between nodes through a graph attention network (GAT); uses bilinear bipartite graph matching combined with the Hungarian algorithm to accurately match keywords with scene graph nodes; generates the main effect nodes and negative effect nodes of the scene graph through a keyword-guided attention module; uses a Transformer Encoder to extract the problem feature vector; finally, integrates the problem and scene graph features through a multimodal fusion strategy and predicts the answer. This technology optimizes the matching of keywords and image regions, improves the accuracy and applicability of the VQA task, enhances the robustness of the system, precisely focuses on the key regions of the image, and reduces the impact of language bias, achieving significant technological progress. Description of the Drawings

[0019] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0020] Figure 1 It is an application environment diagram of a visual question answering method based on bipartite graph matching in an embodiment of the present application;

[0021] Figure 2 It is a schematic diagram of the principle of a visual question answering method based on bipartite graph matching in an embodiment of the present application;

[0022] Figure 3 It is a schematic diagram of constructing a keyword graph using a pre-trained KeyBERT model in an embodiment of the present application;

[0023] Figure 4 It is a schematic diagram of constructing a scene graph using the Faster R CNN combined with the ResNet 101 model in an embodiment of the present application;

[0024] Figure 5 It is a schematic diagram of using the Hungarian algorithm to achieve the alignment between the keyword graph and the heterogeneous graph of the scene graph in an embodiment of the present application. Detailed implementation manners

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.

[0026] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0027] In an exemplary embodiment, as Figure 1 shown, a visual question answering method based on bipartite graph matching is provided, including:

[0028] Step 101: Obtain a question and an image.

[0029] Step 102: Construct a keyword graph based on the question.

[0030] Step 102 includes:

[0031] Step 102-1: Delete the punctuation marks in the text of the question to obtain a de-punctuated question text.

[0032] Step 102-2: Input the problem text after label removal into a pre-trained KeyBERT model, and extract k keywords according to the cosine similarity.

[0033] Step 102-3: Use the Glove word vector model to determine the word vectors of each keyword.

[0034] Step 102-4: Determine the word vectors of the k keywords as keyword vectors.

[0035] Step 102-5: Construct a keyword graph based on the keyword vectors. The keyword graph is an undirected fully connected graph.

[0036] Step 103: Use the graph attention network to encode the interaction information between nodes in the keyword graph to obtain the keyword graph node features.

[0037] Step 103 includes:

[0038] Step 103-1: Use the graph attention network to encode the interaction information between nodes in the keyword graph to obtain an updated keyword graph.

[0039] Step 103-2: Obtain all node features in the updated keyword graph as the keyword graph node features.

[0040] The graph attention network includes multiple attention heads.

[0041] The i-th node after the update of the keyword graph at the L-th layer is:

[0042] ;

[0043] Where ;

[0044] In the formula, is the i-th node after the update of the keyword graph at the L-th layer. is the activation function. is the number of graph nodes. is the keyword graph node at the L-th layer sending a message to node when the attention score. is the first parameter to be learned. is the normalization function. is the multi-layer perceptron. is the i -th node after the normalization process of the keyword graph at the L-1 layer; is the j -th node after the normalization process of the keyword graph at the L-1 layer.

[0045] Step 104: Construct a scene graph based on the picture.

[0046] Step 104 includes:

[0047] Step 104-1: Use an attribute detector to identify the attributes of objects in the picture.

[0048] Step 104-2: Use a relationship decoder to determine the interaction relationships between object pairs in the picture.

[0049] Step 104-3: Construct a scene graph with objects as nodes and the interaction relationships between object pairs as edges.

[0050] Step 105: Use a graph attention network to encode the interaction information between nodes in the scene graph to obtain scene graph node features.

[0051] Step 105 includes:

[0052] Step 105-1: Use a graph attention network to encode the interaction information between nodes in the scene graph to obtain an updated scene graph.

[0053] Step 105-2: Obtain all node features in the updated scene graph as scene graph node features.

[0054] The graph attention network includes multiple attention heads.

[0055] The i-th node after the update of the L-th layer scene graph is:

[0056] .

[0057] Among them, .

[0058] In the formula, is the i-th node after the update of the L-th layer scene graph; is the state of the i-th node of the L-th layer scene graph when sending a message to the node the attention score; is the second parameter to be learned; is the state of the i-th node numbered in the (L-1)-th layer scene graph; is the state of the neighbor nodes of the i-th node in the (L-1)-th layer scene graph; is the state of the relationship between the i-th node and the j-th node in the (L-1)-th layer scene graph.

[0059] Step 106: Based on the keyword graph node features and the scene graph node features, use the Hungarian algorithm to determine the main effect node features and the negative effect node features of the scene graph.

[0060] Step 106 includes:

[0061] Step 106-1: Use bilinear pooling to map the keyword graph node features and the scene graph node features to the same space to obtain a bipartite graph weight matrix.

[0062] The bipartite graph weight matrix is: .

[0063] where F is the bipartite graph weight matrix; Softmax(*) represents the matrix normalization function; represents the Hadamard product; represents a fully connected network; represents a fully connected network; represents a fully connected network. The principle of the fully connected network is: , is the output vector, is the input vector, Lin(*) is the linear layer, Relu(*) is the activation function, and Dropout(*) is the Dropout layer; is the keyword graph node feature; is the scene graph node feature;

[0064] Step 106-2: Based on the bipartite graph weight matrix, use the Hungarian algorithm to determine the K node pairs with the highest contribution to the task and the K node pairs with the highest deviation from the task. A node pair includes a keyword graph node and a scene graph node.

[0065] Step 106-3: Input the K node pairs with the highest contribution to the task into the attention module guided by the first keyword to obtain the main effect node features of the scene graph.

[0066] Step 106-4: Input the K node pairs with the highest deviation from the task into the attention module guided by the second keyword to obtain the negative effect node features of the scene graph.

[0067] Step 107: Construct a problem feature vector based on the problem.

[0068] Step 107 includes:

[0069] Step 107-1: Divide the text of the problem into an array of words according to the punctuation marks and spaces in the text of the problem.

[0070] Step 107-2: Use the Glove word vector model to determine the word vector of each word in the array of words.

[0071] Step 107-3: Input the word vector of each word in the array of words into the Transformer Encoder network respectively, and determine that the average value of the output quantities of the Transformer Encoder network corresponding to all words is the problem feature vector.

[0072] Step 108: Using the multi-modal fusion strategy, fuse the problem feature vector with the scene graph node feature, the scene graph main effect node feature, and the scene graph negative effect node feature respectively to obtain multiple joint features.

[0073] The joint features include: the joint feature corresponding to the scene graph node feature, the joint feature corresponding to the scene graph main effect node feature, and the joint feature corresponding to the scene graph negative effect node feature.

[0074] The joint feature corresponding to the scene graph node feature is:

[0075] 。

[0076] where is the joint feature corresponding to the scene graph node feature. is the multi-modal fusion strategy. is the scene graph node feature. is the problem feature vector.

[0077] The joint feature corresponding to the scene graph main effect node feature is:

[0078] .

[0079] where is the joint feature corresponding to the scene graph main effect node feature. is the scene graph main effect node feature.

[0080] The joint feature corresponding to the scene graph negative effect node feature is:

[0081] .

[0082] where is the joint feature corresponding to the scene graph negative effect node feature. is the scene graph negative effect node feature.

[0083] Step 109: Input the multiple joint features into a classifier to obtain the predicted probability distribution of the problem.

[0084] The predicted probability distribution is:

[0085] .

[0086] where .

[0087] .

[0088] .

[0089] Where P is the predicted probability distribution. is the output of the joint feature corresponding to the scene graph node feature input to the classifier. is the output of the joint feature corresponding to the main effect node feature of the scene graph input to the classifier. is the output of the joint feature corresponding to the negative effect node feature of the scene graph input to the classifier.

[0090] In another exemplary embodiment, as Figure 2 shown, a visual question answering method based on bipartite graph matching is provided, including:

[0091] Step 1: Keyword extraction: Use the pre-trained KeyBERT model to extract keywords in the question, and use the Glove word vector model to vectorize the keywords. Use the pre-trained KeyBERT model to construct a schematic diagram of the keyword graph as Figure 3 shown.

[0092] Step 1.1: Remove punctuation from the input question, and use the pre-trained KeyBERT model to extract k keywords according to the cosine similarity.

[0093] Step 1.2: Use the Glove word vector model to obtain the keyword vector K, expressed as:

[0094] .

[0095] where is the word vector of the k-th keyword, is the set of keyword vectors after being trained by the Glove word vector model.

[0096] Step 2: Keyword graph context fusion: Construct a fully connected keyword graph from the keywords extracted in Step 1, and encode the interaction information between the keywords through a graph attention network (GAT).

[0097] Step 2.1: Construct the keyword graph .

[0098] where represents the node set of the keyword graph, which is composed of the keyword vectors generated in Step 1; represents the edge set of the keyword graph. The keyword graph is an undirected fully connected graph, that is, for any node , there exists an edge .

[0099] Step 2.2: Use the graph attention network in different ways for The neighboring nodes of are weighted and aggregated to node The attention score when passing the message is which is expressed as follows:

[0100] .

[0101] where is a normalization to ensure that the sum of the attention scores from one node to its neighboring nodes is 1, and MLP represents a multi-layer perceptron. After calculating the attention scores, this application calculates the new representation of each node as the weighted average of its neighboring nodes, which is expressed by the formula as follows:

[0102] .

[0103] where, is the activation function, are the parameters to be learned. In practice, this application uses multiple attention heads.

[0104] Step 3: Scene graph generation: Use Faster R-CNN combined with the ResNet-101 network model to extract the regional features and spatial features of the image, add an attribute detector, the graph nodes represent the objects in the image, the edges of the graph are the relationships between node pairs, and the interaction relationships between node pairs within the scene are obtained through the relationship decoder. Use Faster R-CNN combined with the ResNet101 model to construct a schematic diagram of the scene graph as Figure 4 .

[0105] Step 4: Scene graph context fusion: Encode the interaction information between the scene graph nodes through GAT.

[0106] Step 4.1: The scene graph generated in Step 3 is represented as where represents the set of nodes in the scene, including the labels and attributes of the nodes, represents the set of relationships existing between the nodes, including semantic relationships and spatial relationships. Similar to Step 2.2, use the attention mechanism (GAT) to weight and aggregate the features of neighboring nodes, so as to achieve an adaptive allocation of the importance of different neighboring nodes and edges, which is expressed by the formula as follows:

[0107] .

[0108] where represents the state of the node numbered i at the (L - 1)th layer, represents the state of the neighboring nodes of node i, Represents the state of the relationship between node i and node j. Then, this application calculates the new representation of each node as the weighted average of its neighboring nodes, which is expressed by the formula as follows:

[0109] .

[0110] Where is the activation function, is the parameter to be learned. In practice, this application uses multiple attention heads.

[0111] Step 5: Bilinear bipartite graph matching: Consider the keyword graph updated in Step 2 and the scene graph updated in Step 4 as two subgraphs of a bipartite graph. After low-rank bilinear pooling and the Softmax function, obtain the bipartite graph weight matrix, and use the Hungarian algorithm to find the scene graph nodes that best match and least match the keywords respectively. The schematic diagram of using the Hungarian algorithm to achieve the alignment between the keyword graph and the heterogeneous graph of the scene graph is as Figure 5 .

[0112] Step 5.1: Represent the node features of the keyword graph updated in Step 2 as:

[0113] .

[0114] Represent the node features of the scene graph updated in Step 4 as:

[0115] .

[0116] Use bilinear pooling to map and to the same space, and then perform Softmax on the bilinear fusion matrix to obtain the bipartite graph weight matrix, which is expressed by the formula as follows:

[0117] .

[0118] Where represents normalizing the entire matrix; and represent two fully connected networks used to map and to a low-dimensional space, which is expressed as follows:

[0119] .

[0120] Where represents the input vector, represents the output vector, represents the linear layer, represents the activation function, represents the dropout layer.

[0121] Step 5.2: Bipartite graph weight matrix value represents the weights of keyword graph node i and scene graph node j for the entire VQA task. Use the Hungarian algorithm to find the K pairs of node pairs with the highest "contribution" to the entire task and the K pairs of node pairs with the highest "deviation" for the entire task. The scene graph nodes with the highest "contribution" are regarded as "main effect" nodes, denoted as , and the scene graph nodes with the highest "deviation" are regarded as "negative effect" nodes, denoted as .

[0122] Step 6: Keyword-guided attention: Pass the keyword features in Step 5 and the features of the scene graph nodes that match them most through the keyword-guided attention module to generate the "main effect" nodes of the scene graph; pass the features of the scene graph nodes that least match the keywords through another keyword-guided attention module to generate the "negative effect" nodes of the scene graph.

[0123] Use the matching keyword node features to guide the attention allocation of the "main effect" nodes and "negative effect" nodes of the scene graph, so as to extract the scene graph node information related to the keywords, and at the same time embed the keyword information into the feature space of the scene graph; it is expressed by the formula as follows:

[0124] .

[0125] .

[0126] .

[0127] .

[0128] where , represent learnable parameters, is a linear transformation implemented by a feed-forward module. Here, is used to stabilize the training, and d is the dimension of the node vector. The nodes are "main effect" nodes and "negative effect" nodes.

[0129] Step 7: Question embedding: Divide the question into independent words, use the Glove word vector model to vectorize the words, and use the Transformer Encoder to extract the sentence features to obtain the question feature vector Q.

[0130] Step 7.1: Divide the input question into separate words according to punctuation marks and spaces; the input question is converted into an array of words, expressed by the following formula:

[0131] .

[0132] Among them, N is the number of words contained in the sentence, are N individual words, and q is the word set.

[0133] Step 7.2: Use the Glove word vector model to obtain word vectors , which is expressed as:

[0134] = , ,..., .

[0135] Among them, is the word vector of the word , is the word vector set after training by the Glove word vector model;

[0136] Step 7.3: Use the Transformer Encoder network to extract sentence features, sum all word vectors on the final output vector and take the average to obtain the question feature vector Q, which is expressed by the formula as follows:

[0137] .

[0138] Among them, represents the Transformer Encoder operation, represents taking the average.

[0139] Step 8: Multimodal fusion and answer prediction: Send all node features of the scene graph generated by question 4, the "main effect" node features and "negative effect" node features of the scene graph generated in step 6, and the question features generated in step 7 into the multimodal fusion module to generate the final features, and then input them into the classifier to predict the answer. Step 8.1: Respectively fuse the question feature Q with the scene graph feature generated in step 4 and the "main effect" feature and "negative effect" feature of the scene graph generated in step 6 through the multimodal fusion strategy to obtain the joint feature representation:

[0140] .

[0141] .

[0142] .

[0143] Step 8.2: Send F to the fully connected layer of the classifier to obtain the corresponding unnormalized probability scores, which are expressed by the formula as follows:

[0144] 。

[0145] 。

[0146] 。

[0147] Step 8.3: To strengthen the "main effect" and reduce the "negative effect", the present application combines the obtained probability scores and performs to obtain the probability distribution of the final prediction, which is expressed as follows:

[0148] 。

[0149] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a visual question answering method based on bipartite graph matching.

[0150] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which when executed by a processor, implements the steps in the above method embodiments.

[0151] In an exemplary embodiment, a computer program product is provided, including a computer program, which when executed by a processor, implements the steps in the above method embodiments.

[0152] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0153] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0154] The databases involved in the various embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the various embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.

[0155] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0156] Specific examples are used in this article to elaborate on the principles and implementation manners of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A visual question answering method based on bipartite graph matching, characterized in that, Including: Obtain a question and an image; Construct a keyword graph based on the question, including: Delete punctuation marks in the text of the question to obtain a de-punctuated question text; Input the de-punctuated question text into a pre-trained KeyBERT model, and extract k keywords according to the cosine similarity; Use the Glove word vector model to determine the word vectors of each keyword; Determine the word vectors of the k keywords as keyword vectors; Construct a keyword graph based on the keyword vectors; the keyword graph is an undirected fully-connected graph; Use a graph attention network to encode the interaction information between nodes in the keyword graph to obtain keyword graph node features; Construct a scene graph based on the image, including: Use an attribute detector to identify the attributes of the objects in the image; Use a relationship decoder to determine the interaction relationships between object pairs in the image; Construct a scene graph with objects as nodes and the interaction relationships between object pairs as edges; Use a graph attention network to encode the interaction information between nodes in the scene graph to obtain scene graph node features; Based on the keyword graph node features and the scene graph node features, use the Hungarian algorithm to determine the scene graph main effect node features and the scene graph negative effect node features; Construct a question feature vector based on the question, including: According to the punctuation marks and spaces in the text of the question, divide the text of the question into a word array; Use the Glove word vector model to determine the word vectors of each word in the word array; Input the word vectors of each word in the word array into the Transformer Encoder network respectively, and determine the average value of the output quantities of the Transformer Encoder network corresponding to all words as the question feature vector; Use a multimodal fusion strategy to fuse the question feature vector with the scene graph node features, the scene graph main effect node features, and the scene graph negative effect node features respectively to obtain multiple joint features; Input the multiple joint features into a classifier to obtain the predicted probability distribution of the question.

2. The visual question answering method based on bipartite graph matching according to claim 1, wherein Use a graph attention network to encode the interaction information between nodes in the keyword graph to obtain keyword graph node features, including: Use a graph attention network to encode the interaction information between nodes in the keyword graph to obtain an updated keyword graph; Obtain all node features in the updated keyword graph as keyword graph node features; The graph attention network includes multiple attention heads; The i-th node after the L-th layer keyword graph update is: Among them, In the formula, is the i-th node after the update of the keyword graph of the L-th layer; σ is the activation function; N i is the number of graph nodes; is the attention score when the keyword graph node V j of the L-th layer transmits a message to the node V i ; is the first parameter to be learned; is the normalization function; MLP(*) is the multi-layer perceptron; is the i-th node after the normalization process of the keyword graph of the (L-1)-th layer; is the j-th node after the normalization process of the keyword graph of the (L-1)-th layer.

3. The visual question answering method based on bipartite graph matching according to claim 2, wherein, Use a graph attention network to encode the interaction information between nodes in the scene graph to obtain scene graph node features, including: Use a graph attention network to encode the interaction information between nodes in the scene graph to obtain an updated scene graph; Obtain all node features in the updated scene graph as scene graph node features; The graph attention network includes multiple attention heads; The i-th node after the L-th layer scene graph update is: Among them, Wherein, is the i-th node after the update of the L-th layer scene graph; is the node V of the L-th layer scene graph j when sending a message to node V i the attention score at this time; is the second parameter to be learned; is the L - The state of the node numbered the i-th in the 1-layer scenario graph; The state of the neighbor nodes of the i-th node in the (L-1)-layer scenario graph; The state of the relationship between the i-th node and the j-th node in the (L-1)-layer scenario graph.

4. The visual question answering method based on bipartite graph matching according to claim 1, characterized in that, Based on the keyword graph node features and the scene graph node features, use the Hungarian algorithm to determine the scene graph main effect node features and the scene graph negative effect node features, including: The keyword graph node features and the scene graph node features are mapped to the same space by using bilinear pooling to obtain a bipartite graph weight matrix; The bipartite graph weight matrix is as follows: Among them, F is the weight matrix of the bipartite graph; Softmax(*) represents the matrix normalization function; ° represents the Hadamard product; P T represents the first fully connected network; U T represents the second fully connected network; V T represents the third fully connected network. The principle of the fully connected network is: f' = Lin(Dropout(Relu(Lin(f)))), where f' is the output vector, f is the input vector, Lin(*) is the linear layer, Relu(*) is the activation function, and Dropout(*) is the Dropout layer; K Q is the keyword graph node feature; S F is the scene graph node feature; Based on the bipartite graph weight matrix, the Hungarian algorithm is used to determine the K node pairs with the highest contribution to the task and the K node pairs with the highest deviation from the task; the node pairs include a keyword graph node and a scene graph node; The K node pairs with the highest contribution to the task are input into the attention module guided by the first keyword to obtain the scene graph main effect node features; The K node pairs with the highest deviation from the task are input into the attention module guided by the second keyword to obtain the scene graph negative effect node features.

5. The visual question answering method based on bipartite graph matching according to claim 1, characterized in that, The joint features include: the joint features corresponding to the scene graph node features, the joint features corresponding to the scene graph main effect node features, and the joint features corresponding to the scene graph negative effect node features; The joint features corresponding to the scene graph node features are: Where, F F is the combined feature corresponding to the scene graph node feature; cat(*) is the multimodal fusion strategy; S F is the scene graph node feature; represents the Hadamard product; Q is the question feature vector; The joint features corresponding to the scene graph main effect node features are: where F P is the combined feature corresponding to the main effect node feature of the scene graph; S P is the main effect node feature of the scene graph; The joint features corresponding to the scene graph negative effect node features are: where F N is the combined feature corresponding to the negative effect node feature of the scene graph; S N is the negative effect node feature of the scene graph.

6. The visual question answering method based on bipartite graph matching according to claim 1, wherein The predicted probability distribution is: P = Softmax(L P + L F - L N ); where L F = Lin(F F ); L P = Lin(F P ); L N = Lin(F N ); Where P is the predicted probability distribution; L F is the output of the classifier when the joint feature F corresponding to the scene graph node feature is F input; L P is the output of the classifier when the joint feature F corresponding to the main effect node feature of the scene graph is P input; L N is the output of the classifier when the joint feature F corresponding to the negative effect node feature of the scene graph is N input.

7. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and capable of running on the processor, wherein the processor executes the computer program to implement the bipartite graph matching-based visual question answering method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Product knowledge aggregation method and device, computer equipment and storage medium

    CN112328857A

  • Scientific research literature recommendation system based on keyword and graph attention mechanism

    CN118170868A