Implementation method, device, equipment and storage medium for visual question answering

By extracting and fusing the object features, relational features and attribute features of the target picture, the problem of insufficient answer accuracy in the prior art is solved, and more efficient answers to visual questions are achieved.

CN114155422BActive Publication Date: 2025-05-27BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111402921.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-24
Publication Date
2025-05-27
Estimated Expiration
2041-11-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate object features, relational features and object attribute features in visual question answering tasks, resulting in insufficient answer accuracy.

Method used

By obtaining target pictures and questions, extracting object features, relational features and attribute features, and fusing them into comprehensive features, and making answer predictions based on problem features and comprehensive features.

Benefits of technology

It improves the accuracy of answers to visual questions, balances the attention of object information and inter-object relationship information, and makes the fusion information more complete and comprehensive.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155422B_ABST
    Figure CN114155422B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, device, and storage medium for implementing visual question answering, which relates to the field of computer technologies, and particularly to the fields of artificial intelligence and computer vision technologies. The specific implementation solution is as follows: after obtaining a specified target picture and a target question for the target picture, for the target question, extract its question features; for the target picture, respectively extract its object features and relationship features, fuse the object features, relationship features with the attribute features of the target object to obtain the comprehensive features of the target picture, and perform answer prediction based on the question features and the comprehensive features of the target picture to obtain the answer to the target question. Applying the embodiments of the present disclosure, the object features and the object relationship features are respectively extracted, which balances the attention to object information and object relationship information, and by combining object attribute information, the obtained fused information is more complete and comprehensive, so that the visual question answer obtained based on the fused information is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to artificial intelligence and computer vision technologies. Background Art

[0002] The Visual Question Answering (VQA) task refers to answering a series of natural language questions about the content of a given picture based on the provided picture information. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, device, and storage medium for implementing visual question answering to improve the accuracy of answers.

[0004] According to one aspect of the present disclosure, a method for implementing visual question answering is provided, including:

[0005] Obtaining a specified target picture and a target question for the target picture;

[0006] Converting the target question into a question feature;

[0007] Performing object feature extraction and relationship feature extraction on the target picture to obtain object features and relationship features respectively;

[0008] Fusing the object features, relationship features, and attribute features of each target object to obtain a comprehensive feature of the target picture;

[0009] Performing answer prediction based on the question feature and the comprehensive feature of the target picture to obtain an answer to the target question.

[0010] According to another aspect of the present disclosure, an apparatus for implementing visual question answering is provided, including:

[0011] A picture-question acquisition unit for obtaining a specified target picture and a target question for the target picture;

[0012] A question conversion unit for converting the target question into a question feature;

[0013] A feature extraction unit for performing object feature extraction and relationship feature extraction on the target picture to obtain object features and relationship features respectively;

[0014] A feature fusion unit for fusing the object features, relationship features, and attribute features of each target object to obtain a comprehensive feature of the target picture;

[0015] An answer prediction unit for performing answer prediction based on the question feature and the comprehensive feature of the target picture to obtain an answer to the target question.

[0016] According to another aspect of the present disclosure, there is provided an electronic device, including:

[0017] at least one processor; and

[0018] a memory communicatively connected to the at least one processor; wherein,

[0019] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the implementation method of visual question answering described in any one of the above.

[0020] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the implementation method of visual question answering described in any one of the above.

[0021] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements the implementation method of visual question answering described in any one of the above when executed by a processor.

[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0024] Figure 1 is a schematic diagram of a first embodiment of the implementation method of visual question answering provided by the present disclosure;

[0025] Figure 2 is a schematic diagram of a structure of the LSTM network used in the present disclosure;

[0026] Figure 3 is a schematic diagram of a second embodiment of the implementation method of visual question answering provided by the present disclosure;

[0027] Figure 4 is a schematic diagram of a first embodiment of the GGNN two-tower encoder used in the present disclosure;

[0028] Figure 5 is a schematic diagram of a second embodiment of the GGNN two-tower encoder used in the present disclosure;

[0029] Figure 6 is a schematic diagram of a structure of the GRU network used in the present disclosure;

[0030] Figure 7 It is a schematic diagram of the third embodiment of the implementation method for answering visual questions provided by the present disclosure;

[0031] Figure 8 It is a schematic diagram of obtaining the answer to a visual question in the present disclosure;

[0032] Figure 9 It is a schematic diagram of a specific example of the implementation method for answering visual questions provided by the present disclosure;

[0033] Figure 10 It is a schematic diagram of the first embodiment of the implementation device for answering visual questions provided by the present disclosure;

[0034] Figure 11 It is a block diagram of an electronic device for implementing the implementation method of answering visual questions in the embodiments of the present disclosure. Detailed implementation manners

[0035] The following makes an explanation of the exemplary embodiments of the present disclosure in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.

[0036] In order to improve the answer accuracy of answering visual questions, the present disclosure provides an implementation method, device, electronic device, and storage medium for answering visual questions. First, an implementation method for a visual question provided by the present disclosure will be described below.

[0037] See Figure 1 , Figure 1 It is a schematic diagram of the first embodiment of the implementation method for answering visual questions provided by the present disclosure. In the embodiments of the present disclosure, the above method may include the following steps:

[0038] Step S110, obtain a specified target picture and a target question for the target picture.

[0039] In the embodiments of the present disclosure, both the above target picture and the target question for the target picture can be arbitrarily specified by the user, and no specific limitation is made thereto in the present disclosure.

[0040] Step S120, convert the target question into a question feature.

[0041] In one embodiment of the present disclosure, a preset question encoding model can be used to convert a target question into question features. As a specific implementation, the target question can be input into the preset question encoding model, and the preset question encoding model can convert the text of the target question into a vector as the question features. In this way, the acquisition efficiency of the question features of the target question can be improved.

[0042] In one embodiment of the present disclosure, the above-mentioned question encoding model may include a word embedding layer and an LSTM network layer, etc. As a specific implementation, a GLOVE (Global Vectors for Word Representation) pre-trained word vector dataset can be used. This word vector dataset may include the corresponding relationship between words and vectors. After the above-mentioned target question is input into the word embedding layer, the word embedding layer can convert the text of the target question into word vectors based on the above-mentioned word vector dataset. Then, a position encoding matrix can be added to the obtained word vectors to obtain the target word vector sequence to be extracted. The above-mentioned position encoding matrix can represent the positions of the word vectors corresponding to the respective words in the target question in the entire target word vector sequence. After obtaining the above-mentioned target word vector sequence to be extracted, it can be input into the LSTM network, and the LSTM network can extract the corresponding question features. As described above, the above-mentioned question features can be expressed in vector form. For example, q = R dim , where dim can represent the output dimension of the question vector.

[0043] As Figure 2 shown, Figure 2 is a schematic structural diagram of the LSTM network used in the embodiment of the present disclosure.

[0044] The LSTM network (Long Short-Term Memory) is a long short-term memory RNN network (Recurrent Neural Network). It effectively retains the long-term information and short-term information in the word sequence through the memory gate and forget gate structures, and abstracts the question word vector sequence into a sentence vector.

[0045] Figure 2 In, ⊙ is the Hadamard Product, that is, multiplying the corresponding elements in the matrix. Therefore, it is required that the two matrices to be multiplied are of the same type. + represents matrix addition. z f , z i , z o is converted into a value between 0 and 1 through a sigmoid activation function after multiplying the concatenated vector by the weight matrix, as a kind of gating state. And z is converted into a value between -1 and 1 through a tanh activation function.

[0046] There are mainly three stages inside the LSTM:

[0047] 1. Forget stage. In this stage, the input passed in from the previous node is selectively forgotten. Simply put, it will "forget the unimportant and remember the important".

[0048] Specifically, it is through the calculated z f (where f represents forget) as the forget gate to control the previous state of c t-1 to determine what to keep and what to forget.

[0049] 2. Selective memory stage. In this stage, the input of this stage is selectively "remembered". It mainly selectively remembers the input x t (the target word vector sequence to be extracted). What is important is emphasized and recorded, and what is unimportant is remembered less. The current input content is represented by z calculated previously. And the selected gate signal is controlled by z i (where i represents information).

[0050] Adding the results obtained from the above two steps gives the c to be transmitted to the next state t . That is Figure 2 the formula 1 in

[0051] 3. Output stage. This stage determines what will be regarded as the output of the current state. It is mainly controlled by z o . And it also scales the c obtained in the previous stage o (through a tanh activation function for transformation, Figure 2 the formula 2 in

[0052] Similar to the ordinary RNN, the output y t (the problem feature) is often finally obtained through the transformation of h t ( Figure 2 the formula 3 in

[0053] As Figure 1 shown, in step S130, object feature extraction and relationship feature extraction are performed on the target picture to obtain object features and relationship features respectively.

[0054] In the embodiments of the present disclosure, a preset model can also be used to perform object feature extraction and relationship feature extraction on the target picture respectively.

[0055] Step S140, fuse the object features, relationship features, and attribute features of each target object to obtain the comprehensive feature of the target picture.

[0056] In the embodiments of the present disclosure, when obtaining the attribute features of the above target objects, it may be to obtain the preset target attribute texts of each target object, and convert each target attribute text into an information word vector respectively as the attribute feature of each target object.

[0057] In the embodiments of the present disclosure, the above preset target attribute texts may include texts corresponding to each attribute of the object, and the object attributes may include materials, colors, shapes, etc. The texts corresponding to the material attribute may be metals, plastics, cloth, etc. indicating the material; the texts corresponding to the color attribute may be red, purple, blue, etc. indicating the color; and the preset texts corresponding to the shape may be square, circular, diamond, etc. identifying the shape.

[0058] In the embodiments of the present disclosure, a pre-trained target recognition model may be used to recognize each attribute (such as color, shape) of each object in the target picture, so as to obtain the target attributes of each object, and further obtain the target attribute texts.

[0059] In the embodiments of the present disclosure, it may also be to perform word embedding on the above target attribute texts based on the word vector dataset pre-trained by the above GLOVE to obtain the respective attribute features of each target object.

[0060] By extracting and fusing the attribute features of each object in the target picture, the features of the target picture can be obtained more comprehensively, making the answers to visual question answering more accurate and supporting more types of questions.

[0061] For ease of description, in the present disclosure, L may be defined as the category of all attributes (such as materials, colors, shapes, etc.). For each object n i , a set of L + 1 attribute variables may be defined where may represent the word vector of the object's own name, and may represent the information word vector corresponding to the l-th attribute of the object.

[0062] Step S150, perform answer prediction based on the question feature and the comprehensive feature of the target picture to obtain the answer to the target question.

[0063] In the embodiments of the present disclosure, a corresponding model may also be used to perform answer prediction based on the question feature and the comprehensive feature of the target picture.

[0064] The implementation method of visual question answering provided by the embodiments of the present disclosure, after obtaining a specified target picture and a target question for the target picture, for the target question, extracts its question features; for the target picture, respectively extracts its object features and relationship features, fuses the object features, relationship features and the attribute features of the target object to obtain the comprehensive features of the target picture, and performs answer prediction based on the question features and the comprehensive features of the target picture to obtain the answer to the target question. Applying the embodiments of the present disclosure, separately extracting object information and object relationship information balances the attention to object information and object relationship information, and by combining object attribute information, the obtained fused information is more complete and comprehensive, so that the visual question answer obtained based on the fused information is more accurate.

[0065] In an embodiment of the present disclosure, when performing object feature extraction and relationship feature extraction on a target picture, the target picture can first be subjected to feature recognition, and based on the recognition result, a directed graph including the target object and the target association relationship is generated as a scene graph, and based on the scene graph, object features and relationship features are respectively extracted.

[0066] A scene graph is an interpretable and structured scene representation that summarizes the entities in the scene and the reasonable relationships between them. In the embodiments of the present disclosure, when generating a scene graph, an RPN network can be used to recognize the target picture to obtain proposal boxes of each object in the target picture, and an RNN network can be used for reasoning to construct a scene graph for the target picture based on the above proposal boxes. Of course, other scene graph construction methods can also be used in the present disclosure to obtain a scene graph.

[0067] Since the scene graph contains objects and the relationships between objects, therefore, feature extraction can be performed on the scene graph to separately obtain the object features and relationship features in the target picture.

[0068] And separately performing object feature extraction and relationship feature extraction based on the scene graph can reduce the influence of the background image in the target picture on the extraction result, making the extraction result more accurate.

[0069] As a specific implementation, as Figure 3 shown, Figure 1 step S130 in

[0070] Specifically, step S131 may include the following steps:

[0071] As described above, the scene graph contains the objects in the target image and the relationships between the objects. Therefore, the above scene graph can be converted into an object graph and a relationship graph, and object features and relationship features can be extracted from the object graph and the relationship graph respectively.

[0072] In the embodiments of the present disclosure, each node in the object graph can represent an object that actually exists in the target image, and each edge can represent the relationship between two objects. In the embodiments of the present disclosure, the object graph (object-significant graph) can be represented as G obj , and at the same time, N can be defined as the set of nodes, and E can be defined as the set of edges. For n i , n j ∈ N, e k ∈ E, <n i -e k -n j > represents a relationship pair, where the relationship edge e k is the relationship between objects n i , n j . In the embodiments of the present disclosure, the edges in the object graph are directed edges and are not symmetric. That is, if <n i -e k -n j > is a valid relationship pair, but <n j -e k -n i > does not necessarily exist. At the same time, there can be multiple relationship edges between n i , n j .

[0073] In the embodiments of the present disclosure, each node in the relationship graph can represent the relationship between objects in the target image, and each edge represents the object shared between two relationships. In the embodiments of the present disclosure, the relationship graph (relation-significant graph) can be defined as G rel , and the structure of the relationship graph is exactly the opposite of that of the object graph. Among them, for e i , e j ∈ E, nk ∈ N, <e i -n k -e j > represents that there is a shared object n i between relationship e j and e k . In the embodiments of the present disclosure, the relationship graph has symmetry.

[0074] Step S132, extracting object features based on the object graph and extracting relationship features based on the relationship graph.

[0075] In the embodiments of the present disclosure, by separately extracting object features and relationship features for the object graph and the relationship graph, the extraction of the two types of features is made more accurate, and the attention to object information and inter-object relationship information is effectively balanced, further improving the accuracy of visual question answering.

[0076] Correspondingly, as Figure 3 shown, Figure 1 step S140 in

[0077] can be refined into the following steps:

[0078] In the embodiments of the present disclosure, the original features of the above relationships may be relationship word vectors obtained by performing word embedding on the text of the relationships (such as wearing, eating, left, above, etc.) based on the above GLOVE pre-trained word vector dataset. The text of the above relationships can be obtained during the conversion of the scene graph.

[0079] Step S142, fusing the object fusion feature and the relationship fusion feature to obtain the comprehensive feature of the target picture.

[0080] In the embodiments of the present disclosure, by fusing object features, relationship features, object attribute features, and relationship original features, the fusion features can be made more accurate.

[0081] As described above, in the present disclosure, when extracting the above object features and relationship features, a model can be used for extraction. As a specific implementation manner of the embodiments of the present disclosure, a preset dual-tower GGNN encoder can be used to separately extract object features and relationship features based on the above scene graph.

[0082] As Figure 4 shown, the above dual-tower GGNN encoder may include a scene graph conversion module 410, a first GGNN network 420, a second GGNN network 430, a first fusion module 440, a second fusion module 450, and a third fusion module 460.

[0083] Specifically, as a specific implementation manner of the embodiments of the present disclosure, the above steps of separately extracting object features and relationship features based on the scene graph may include:

[0084] Inputting the scene graph into a preset dual-tower GGNN encoder;

[0085] The scene graph conversion module 410 converts the input scene graph into an object graph and a relationship graph, and inputs them into the first GGNN network 420 and the second GGNN network 430 respectively.

[0086] The first GGNN network 420 extracts features from the object graph to obtain object features and outputs them to the first fusion module 440; the second GGNN network 430 extracts features from the relationship graph to obtain relationship features and outputs them to the second fusion module 450.

[0087] In the embodiments of the present disclosure, since the above-mentioned first GGNN network is used to extract object features, the first GGNN network can also be referred to as an object encoder, and the above-mentioned second GGNN network extracts relationship features from the relationship graph. Therefore, the second GGNN network can also be referred to as a relationship encoder.

[0088] The step of fusing the object features, relationship features, and attribute features of each target object to obtain the comprehensive features of the target image may include:

[0089] The first fusion module 440 fuses the object features output by the first GGNN network 420 with the attribute features of each target object to obtain object fusion features and outputs them to the third fusion module 460; the second fusion module 450 fuses the relationship features output by the second GGNN network 430 with the original features of each relationship to obtain relationship fusion features and outputs them to the third fusion module 460;

[0090] The third fusion module 460 fuses the object fusion features with the relationship fusion features to obtain the comprehensive features of the target image.

[0091] In the embodiments of the present disclosure, by separately using an object encoder and a relationship encoder to extract object features and relationship features, the attention to object information and inter-object relationship information can be better balanced, so that the obtained comprehensive features contain more comprehensive and specific information.

[0092] In the embodiments of the present disclosure, the above-mentioned dual-tower GGNN is a Gated Graph Neural Networks. The GNN network is a neural network structure that directly acts on graph structures (such as the above-mentioned object graph and relationship graph). Its purpose can be understood as using labeled nodes to predict the labels of unlabeled nodes. When calculating for a node, the GNN network can learn from the nodes and edges around the node to obtain the influence of the nodes and edges around the node on the node, and update the hidden state of the node based on this influence. The hidden state of a node can represent the state of the node at a certain moment, and the hidden states of different time nodes can be different.

[0093] As Figure 5 shown, in an embodiment of the present disclosure, Figure 4The first GGNN network 420 in it may include: a first data conversion module 421, a first information transfer module 422, and a first hidden state graph generation module 423;

[0094] The first data conversion module 421 is configured to convert the input object graph into a first data combination; the first data combination may include: the own feature vectors of all object nodes in the object graph, the own feature vectors of all edges in the object graph, the adjacency matrix of the incoming edges in the object graph, and the adjacency matrix of the outgoing edges in the object graph.

[0095] As a specific implementation, the first data conversion module 421 may convert the input scene graph into a set of data combinations (A in , A out , N, E), where N represents the set of the own feature vectors of each node, E represents the set of the feature vectors of the directed relationship edges representing the object relationships, A in represents the adjacency matrix of the incoming edges, and A out represents the adjacency matrix of the outgoing edges.

[0096] In the embodiments of the present disclosure, may be used to mark the hidden state of the n i node at the t-th time step of the GGNN encoder. When t = 0, the word vector generated by the GLOVE word embedding of n i and padded with 0 may be used as the initial value of :

[0097]

[0098] The incoming edge information and the outgoing edge information may be retrieved in A in , A out .

[0099] The first information transfer module 422 is configured to, at any time step, for the object node to be updated in the object graph, calculate the first incoming edge information gain and the first outgoing edge information gain of the object node to be updated based on the adjacency matrix of the incoming edges and the adjacency matrix of the outgoing edges in the object graph; obtain the first reference feature learned by the object node to be updated from all its incoming edges, outgoing edges, and adjacent nodes based on the first incoming edge information gain and the first outgoing edge information gain; and obtain the updated hidden state of the object node to be updated based on the reference feature and the hidden state of the object node to be updated at the previous time step.

[0100] In the embodiments of the present disclosure, when the above-mentioned first information transmission module 422 calculates for the object node to be updated, it can learn the representations of the surrounding nodes and edges of the object node, thereby updating the hidden state of the node. During the calculation process, the connection between the node and the surrounding nodes and edges is enhanced, more complete information of the node can be obtained, and the obtained object features also contain more complete information.

[0101] As a specific implementation manner of the embodiments of the present disclosure, the above-mentioned first information transmission module 422 may include a first GRU network and a second GRU network.

[0102] GRU is a variant of the RNN network. It is a bidirectional network with an effect similar to that of LSTM and is easier to train. This GRU network can update the hidden state of the node. As Figure 6 shown, Figure 6 is a schematic structural diagram of the GRU network used in the embodiments of the present disclosure.

[0103] GRU uses the gating r for resetting and only uses one gating z to simultaneously perform forgetting and selective memory, while LSTM requires multiple gates, which reduces the training calculation cost. The specific formula is as follows:

[0104] Selective forgetting gate: (1 - z)⊙h (t-1)

[0105] Selective memory gate: z⊙h'

[0106] Updated hidden layer: h t =(1 - z)⊙h (t-1) +z⊙h'

[0107] In the above formula, z is a pre-calculated value with a value range between 0 and 1 and can be used as a gate. h (t-1) is the hidden state of the node at the previous time step, h t is the updated hidden state of the node, and h' is the original feature of the node. Figure 6 in is an intermediate quantity during the calculation process and can be obtained by splicing the reset h (t-1) with the input data.

[0108] In the embodiments of the present disclosure, calculating the first incident edge information gain and the first outgoing edge information gain of the object node to be updated based on the adjacency matrix of the incident edges and the adjacency matrix of the outgoing edges in the object graph includes:

[0109] Input the hidden state of the object node to be updated at the previous time step, the adjacency matrix of the incoming edges, and the adjacency matrix of the outgoing edges in the object graph into the first GRU network to obtain the first incoming edge information gain and the first outgoing edge information gain of the object node to be updated.

[0110] When the first GRU network calculates for a node to be updated, it can calculate the influence of the nodes and edges around the node to be updated on the node to be updated.

[0111] Taking the above relationship for <n i -e k -n j > as an example, the surrounding nodes of node n i include node n j , and the connected edge is e k . The word vector state e k of the edge and the hidden state h j of the adjacent node n j can be input into the above first GRU network together as a sequence of length 2.

[0112] Meanwhile, the hidden state h i of the node n i to be updated will also be input into the first GRU network as the initial hidden state of its network. The output of the first GRU network is both the information update of this set of relationships on the hidden state h i of the node to be updated, and at the same time, it is also the refinement of the key information related to e k , n j and n i . For all the adjacent nodes and associated edges of the object node to be updated, the sum of the outputs of the first GRU network is the total information gain of node n i from all its adjacent nodes and associated edges. The specific formula of the first GRU network is as follows:

[0113]

[0114]

[0115] Where EF i (A in ) is the incoming edge information gain of node n i , and EF i (A out ) is the information gain of the outgoing edge of node n i .

[0116] In the embodiments of the present disclosure, the first GRU network can calculate the influence (information gain) of the adjacent nodes and associated edges of the node to be updated, that is, the information of other nodes and edges flows to the node to be updated. Therefore, the first GRU network can also be referred to as an energy flow module (EF).

[0117] Correspondingly, obtaining the first reference feature learned by the object node to be updated from all its incoming edges, outgoing edges, and adjacent nodes based on the first incoming edge information gain and the first outgoing edge information gain; obtaining the updated hidden state of the object node to be updated based on the reference feature and the hidden state of the object node to be updated at the previous time step may include:

[0118] Input the hidden state, the first incoming edge information gain, and the first outgoing edge information gain of the object node to be updated at the previous time step into the second GRU network, so that the second GRU network can obtain the first reference feature learned by the object node to be updated from all its incoming edges, outgoing edges, and adjacent nodes based on the first incoming edge information gain and the first outgoing edge information gain; obtain the updated hidden state of the object node to be updated based on the reference feature and the hidden state of the object node to be updated at the previous time step.

[0119] Similar to the first GRU network above, the second GRU network is also a gated structure.

[0120] As described above, the energy flow module can calculate the incoming edge gain and the outgoing edge gain of the node to be updated respectively. Therefore, by splicing the two gains, the influence (gain) of all adjacent nodes and connected edges of the node to be updated on the node to be updated can be obtained, as shown in the following formula:

[0121]

[0122] Where represents the representation learned by node n i from all its incoming edges, outgoing edges, and adjacent nodes. And EF() is the information gain learned through the energy flow module.

[0123] After that, the second GRU network described above can be used to combine the information learned from the surrounding nodes and edges and the information of itself at the previous time step, and the combination result is used to update the hidden information of the node:

[0124]

[0125]

[0126]

[0127] Where U z ,Ur , where \(W\) represents a trainable weight matrix and \(b\) is a bias term. represents the update gate and reset gate at time step \(t\).

[0128]

[0129]

[0130] Among them, \(U\) 1 , \(U\) 2 are trainable parameters of the linear layer, and the operator \(\odot\) represents element-wise matrix multiplication. is the updated hidden state of node \(n\) i .

[0131] As can be seen from the above, the first GRU network can learn relevant information from surrounding nodes and edges, while the second GRU network can combine the information learned from surrounding nodes and edges and its own information at the previous time step, and update the hidden state of the node to be updated based on this, which is equivalent to passing the information of surrounding nodes and edges to the node to be updated. Therefore, the first GRU network and the second GRU network together can be called a transfer module, and this transfer module can include the above energy flow module.

[0132] It can be seen that using the GGNN network structure provided by the embodiments of the present disclosure can better obtain information from the surrounding nodes and edges of the node to be updated.

[0133] The first hidden state graph generation module 423 is used to generate a first hidden state graph corresponding to the object graph as object features based on the final hidden state of each object node in the object graph and the own feature vector of the object node output by the first information transfer module after iteration for a preset number of time steps.

[0134] The number of the above preset time steps can be set artificially.

[0135] As a specific implementation manner, the first hidden state graph generation module 423 may include: a first MLP network and a first fully connected network.

[0136] After iteration for a preset number of time steps, for each object node in the object graph, generating a first hidden state graph corresponding to the object graph based on the final hidden state of the object node output by the first information transfer module and the own feature vector of the object node includes:

[0137] For each object node in the object graph, input the final hidden state of the object node output by the first information transfer module and the own feature vector of the object node into the first MLP network, so that the first MLP network generates a first hidden state graph corresponding to the object graph and inputs it into the first fully connected network, so that the first fully connected network outputs the first hidden state graph.

[0138] In the embodiments of the present disclosure, the following formula can be used for node n i to calculate g i ∈G:

[0139]

[0140] where f is a multi-layer perception network (MLP), which receives n i the concatenated vector of two vectors and generates the final representation of n i i.e., the representation learned by the object from the surrounding objects and relationships. In this way, the ability to obtain relevant information from the surrounding nodes and edges of the object node is further enhanced.

[0141] Since in the embodiments of the present disclosure, the above GGNN encoder updates and calculates the state for a single node, therefore, a fully connected layer can be used to directly output the final representation of the object node to be updated.

[0142] The extraction of the relationship features of the relationship graph is exactly the same as the above object feature extraction process. Only a simple description will be given below and will not be elaborated.

[0143] As Figure 5 shown, Figure 4 the second GGNN network 430 in

[0144] includes: a second data conversion module 431, a second information transfer module 432, and a second hidden state graph generation module 433;

[0145] The second data conversion module 431 can be used to convert the input relationship graph into a second data combination; the second data combination can include: the own feature vectors of all relationship nodes in the relationship graph, the own feature vectors of all edges in the relationship graph, the adjacency matrix of the incoming edges, and the adjacency matrix of the outgoing edges in the relationship graph;The second information transfer module 432 is configured to, in any time step, for the relationship nodes to be updated in the relationship graph, calculate the second incident edge information gain and the second outgoing edge information gain of the relationship nodes to be updated based on the adjacency matrix of the incident edges and the adjacency matrix of the outgoing edges in the relationship graph; obtain the second reference features learned by the relationship nodes to be updated from all their incident edges, outgoing edges, and adjacent nodes based on the second incident edge information gain and the second outgoing edge information gain; and obtain the updated hidden state of the relationship nodes to be updated based on the reference features and the hidden state of the relationship nodes to be updated in the previous time step.

[0146] The second information transfer module 432 may include a third GRU network and a fourth GRU network.

[0147] For the relationship nodes to be updated in the relationship graph, calculating the second incident edge information gain and the second outgoing edge information gain of the relationship nodes to be updated based on the adjacency matrix of the incident edges and the adjacency matrix of the outgoing edges in the relationship graph includes:

[0148] Input the hidden state of the relationship nodes to be updated in the previous time step, the adjacency matrix of the incident edges, and the adjacency matrix of the outgoing edges in the relationship graph into the third GRU network to obtain the second incident edge information gain and the second outgoing edge information gain of the relationship nodes to be updated.

[0149] Obtaining the second reference features learned by the relationship nodes to be updated from all their incident edges, outgoing edges, and adjacent nodes based on the second incident edge information gain and the second outgoing edge information gain; and obtaining the updated hidden state of the relationship nodes to be updated based on the reference features and the hidden state of the relationship nodes to be updated in the previous time step includes:

[0150] Input the hidden state of the relationship nodes to be updated in the previous time step, the second incident edge information gain, and the second outgoing edge information gain into the fourth GRU network, so that the fourth GRU network obtains the second reference features learned by the relationship nodes to be updated from all their incident edges, outgoing edges, and adjacent nodes based on the second incident edge information gain and the second outgoing edge information gain; and obtains the updated hidden state of the relationship nodes to be updated based on the reference features and the hidden state of the relationship nodes to be updated in the previous time step.

[0151] The second hidden state graph generation module 433 is configured to, after the iteration of a preset number of time steps, for each relationship node in the relationship graph, generate a second hidden state graph corresponding to the relationship graph as the relationship feature based on the final hidden state of the relationship node output by the second information transfer module and the own feature vector of the relationship node.

[0152] The second hidden state graph generation module 433 may include: a second MLP network and a second fully connected network;

[0153] After the iteration of the preset number of time steps, for each relationship node in the relationship graph, based on the final hidden state of the relationship node output by the second information passing module and the own feature vector of the relationship node, generating a second hidden state graph corresponding to the relationship graph, including:

[0154] For each relationship node in the relationship graph, inputting the final hidden state of the relationship node output by the second information passing module and the own feature vector of the relationship node into the second MLP network, so that the second MLP network generates a second hidden state graph corresponding to the relationship graph and inputs it into the second fully connected network, so that the second fully connected network outputs the second hidden state graph.

[0155] The above second hidden state graph can be the representation learned by each relationship node from surrounding objects and relationships, and can be denoted as G E 。

[0156] As described above, after obtaining the above object features and relationship features, the first fusion module fuses the object features output by the first GGNN network with the attribute features of each target object, and outputs the object fusion features to the third fusion module; the second fusion module fuses the relationship features output by the second GGNN network with the original features of each relationship, and outputs the relationship fusion features to the third fusion module.

[0157] In the embodiments of the present disclosure, after receiving the representation vectors of nodes and edges from the dual-tower GGNN encoder, the attribute information of the object can be first fused into the feature matrix of the scene graph. For the object feature matrix G N of the object node, and the relationship feature matrix G E of the relationship, the fused feature matrix F N (object fusion feature), F E (relationship fusion feature) can be generated by the following formula:

[0158]

[0159] F = [F N , F E

[0160] Where is the fusion feature of the object node n i . is the representation obtained by the object node n i from the GGNN encoder. For object node n i is the word vector of the object attribute. For relation edge e j is the fused feature. For relation edge e j is the representation obtained from the GGNN encoder. e j is the original word vector representation of relation edge j. The final overall feature matrix F is F N concatenated with F E , that is, the third fusion module fuses the object fusion feature and the relation fusion feature to obtain the comprehensive feature of the target picture.

[0161] See Figure 7 , Figure 1 The steps in S150 in

[0162] Step S151, calculate and obtain the attention feature based on the problem feature and the comprehensive feature;

[0163] In the embodiments of the present disclosure, after obtaining the comprehensive feature matrix F, the problem feature q generated by the LSTM and the comprehensive feature matrix F can be input into a multi-head attention layer. This neural network layer can calculate the scoring situation of each feature. After weighted summation of all features, the attention feature r inferred from the scene graph and the problem can be obtained.

[0164] r = Attention(F, q)

[0165] Since F is obtained by concatenating the object fusion feature and the relation fusion feature, each of the above attention features can include the attention features of each object and the attention features of each edge.

[0166] Step S152, fuse the attention feature and the problem feature to obtain the feature to be predicted;

[0167] In this step, the above attention feature r and the problem feature can be concatenated to obtain the feature to be predicted.

[0168] Step S153, input the feature to be predicted into a preset prediction model, and obtain the target answer selected by the prediction model from multiple candidate answers preset for the target picture as the answer to the target question.

[0169] In the embodiments of the present disclosure, in the above prediction model, a two-layer MLP network can be used as a classifier. For the sake of convenience of description, this network can be defined as f(*), where * can represent the input of the network. The input of the MLP network is q and each attention feature r iThe splicing, i.e., the to-be-predicted feature mentioned above.

[0170]

[0171] In this formula, represents the final inference answer. Softmax normalizes the probabilities of each answer output by the above two-layer MLP network, and argmax represents selecting the answer with the highest probability from each answer as the final target answer. Each of the above answers can include objects, relationships, and object attributes in the target image, etc.

[0172] In this way, by performing visual question answering based on the attention feature and the question feature, the accuracy of visual question answering is further improved, and it can also support the answering of various types of visual questions.

[0173] As Figure 8 shown, after obtaining the feature map (object fusion feature and relationship fusion feature) of the target image and the question feature, the feature map of the target image and the question feature can be input into the above-mentioned third fusion module. The third fusion module can perform attention calculation on the feature map and the question feature to obtain the attention feature, and input the attention feature and the question feature into a two-layer MLP network. The MLP network calculates the probabilities of each answer for each attention feature and the question feature, and selects the answer with the highest probability from the list of answers composed of each answer as the final answer.

[0174] As Figure 9 shown, Figure 9 is a schematic diagram of the process of visual answering using the implementation method of visual question answering provided by the present disclosure. Specifically, it can include the following steps:

[0175] 1) Input the target image into the scene graph generation module, and the scene graph generation module can generate a scene graph including each object and the relationship between objects in the target image. Figure 9 In the scene graph shown, the objects can include a child, a sandwich, and sunglasses, and the relationships among the three can be: the child eats the sandwich, the child wears sunglasses, and the sunglasses are on top of the sandwich.

[0176] 2) Input the question "What is on the child's head" into the question encoder to obtain the question feature, and input the question feature into the fusion module and the answer prediction module.

[0177] 3) Input the scene graph into the dual - tower GGNN encoder. For the object graph by the energy - flow module in the object encoder, calculate the incoming - edge information gain and outgoing - edge information gain of each object node (the representation learned by the object node from other nodes and edges), so as to obtain each object feature, and fuse each object feature with each object - attribute feature to obtain the object - fused feature. For the relationship graph by the energy - flow module in the relationship encoder, calculate the incoming - edge information gain and outgoing - edge information gain of each relationship node (the representation learned by the relationship node from other nodes and edges), so as to obtain the relationship feature, and fuse the relationship feature with the original relationship feature to obtain the relationship - fused feature. Fuse the relationship - fused feature with the object - fused feature to obtain the comprehensive feature of the target picture.

[0178] 4) Input the comprehensive feature of the target picture into the fusion module. The fusion module calculates the attention for the comprehensive feature of the target picture and the question feature input in step 2) to obtain the attention feature.

[0179] 5) Input the attention feature into the answer prediction module. The answer prediction module predicts the answer based on the attention feature and the question feature input in step 2), so as to obtain the answer "sunglasses" to the above "what is on the child's head".

[0180] In the embodiments of the present disclosure, the specific training process of the above - mentioned dual - tower GGNN model can be end - to - end training. The datasets used can be the Visual Genome dataset and the GQA dataset. These two datasets contain rich scene graph annotations and high - quality question - answer pairs. The above - mentioned dual - tower GGNN model can use 50 - dimensional GLOVE word - vector encoding, and the loss function can use the Adam function.

[0181] It can be seen from the foregoing that, compared with the problems in the prior art where using the state - machine transition mechanism for visual question answering results in limited application scenarios, inability to effectively fuse the three representations of objects, object attributes, and object - object relationships, and a greater preference for object features while relatively neglecting object - object relationship features; and compared with the problems where using the inference model of the graph neural network results in the inability to effectively utilize the representation of object attributes and a greater preference for object features while relatively neglecting object - object relationship features, the method for implementing visual question answering provided by the present disclosure extracts object features and relationship features using an object encoder and a relationship encoder, effectively balancing the attention to object - object relationship information and object - itself information. At the same time, by using the energy - flow module, it enhances the ability of nodes to obtain relevant information from surrounding information, and at the same time, through explicit attribute splicing, effectively combines the three types of information of object - itself information, object - object relationship information, and object - attribute information to obtain a joint representation.

[0182] The method for implementing visual question answering provided by the present disclosure can be used in AI assistant systems, multi - modal customer - service dialogues, computer question - answering systems, and so on.

[0183] According to an embodiment of the present disclosure, the present disclosure further provides an implementation device for visual question answering, as Figure 10 shown, the device may include:

[0184] A picture question acquisition unit 1010, configured to acquire a specified target picture and a target question for the target picture;

[0185] A question conversion unit 1020, configured to convert the target question into a question feature;

[0186] A feature extraction unit 1030, configured to perform object feature extraction and relationship feature extraction on the target picture, and respectively obtain an object feature and a relationship feature;

[0187] A feature fusion unit 1040, configured to fuse the object feature, the relationship feature, and the attribute features of each target object to obtain a comprehensive feature of the target picture;

[0188] An answer prediction unit 1050, configured to perform answer prediction based on the question feature and the comprehensive feature of the target picture to obtain an answer to the target question.

[0189] After the implementation device for visual question answering provided by the embodiment of the present disclosure acquires a specified target picture and a target question for the target picture, for the target question, it extracts its question feature; for the target picture, it respectively extracts its object feature and relationship feature, fuses the object feature, the relationship feature, and the attribute features of the target object to obtain a comprehensive feature of the target picture, and performs answer prediction based on the question feature and the comprehensive feature of the target picture to obtain an answer to the target question. Applying the embodiment of the present disclosure, the object information and the object relationship information are respectively extracted, the attention to the object information and the object relationship information is balanced, and by combining the object attribute information, the obtained fused information is more complete and comprehensive, so that the visual question answer obtained based on the fused information is more accurate.

[0190] In an embodiment of the present disclosure, the feature extraction unit 1030 may further be configured to perform feature recognition on the target picture, and based on the recognition result, generate a directed graph including the target object and the target association relationship as a scene graph;

[0191] Based on the scene graph, the object feature and the relationship feature are respectively extracted.

[0192] In an embodiment of the present disclosure, the feature extraction unit 1030 may specifically be configured to convert the scene graph into an object graph and a relationship graph; wherein, the object graph is a directed graph with each target object as a node and each target association relationship as an edge; the relationship graph is a directed graph with each target association relationship as a node and each target object as an edge;

[0193] Extract object features based on the object graph and relationship features based on the relationship graph.

[0194] In one embodiment of the present disclosure, the feature fusion unit 1040 can be used to fuse the object features with the attribute features of each target object to obtain object fusion features; and fuse the relationship features with the original features of each relationship to obtain relationship fusion features;

[0195] Fuse the object fusion features with the relationship fusion features to obtain the comprehensive features of the target picture.

[0196] In one embodiment of the present disclosure, the feature extraction unit 1030 is used for:

[0197] Input the scene graph into a preset two-tower GGNN encoder; the two-tower GGNN encoder includes: a scene graph conversion module, a first GGNN network, a second GGNN network, a first fusion module, a second fusion module, and a third fusion module;

[0198] The scene graph conversion module converts the input scene graph into an object graph and a relationship graph, and inputs them into the first GGNN network and the second GGNN network respectively;

[0199] The first GGNN network extracts features from the object graph to obtain object features and outputs them to the first fusion module; the second GGNN network extracts features from the relationship graph to obtain relationship features and outputs them to the second fusion module;

[0200] The feature fusion unit 1040 is used for:

[0201] The first fusion module fuses the object features output by the first GGNN network with the attribute features of each target object to obtain object fusion features and outputs them to the third fusion module; the second fusion module fuses the relationship features output by the second GGNN network with the original features of each relationship to obtain relationship fusion features and outputs them to the third fusion module;

[0202] The third fusion module fuses the object fusion features with the relationship fusion features to obtain the comprehensive features of the target picture.

[0203] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0204] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0205] Figure 11 FIG. shows a schematic block diagram of an example electronic device 1100 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0206] As Figure 11 shown, the device 1100 includes a computing unit 1101 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the device 1100 can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.

[0207] A plurality of components in the device 1100 are connected to the I / O interface 1105, including: an input unit 1106, such as a keyboard, a mouse, etc.; an output unit 1107, such as various types of displays, speakers, etc.; a storage unit 1108, such as a magnetic disk, an optical disk, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows the device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0208] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 executes the various methods and processes described above, such as the implementation method of visual question answering. For example, in some embodiments, the implementation method of visual question answering can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the implementation method of visual question answering described above can be executed. Alternatively, in other embodiments, the computing unit 1101 can be configured to execute the implementation method of visual question answering in any other suitable way (e.g., by means of firmware).

[0209] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0210] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0211] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0212] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0213] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0214] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0215] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0216] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for implementing visual question answering, including: obtaining a specified target image and a target question for the target image; converting the target question into question features; performing object feature extraction and relationship feature extraction on the target image to obtain object features and relationship features respectively; fusing the object features, relationship features, and attribute features of each target object to obtain comprehensive features of the target image; performing answer prediction based on the question features and the comprehensive features of the target image to obtain an answer to the target question; The step of performing object feature extraction and relationship feature extraction on the target image to obtain object features and relationship features respectively includes: performing feature recognition on the target image, and based on the recognition result, generating a directed graph including target objects and target association relationships as a scene graph; extracting object features and relationship features respectively based on the scene graph; The step of extracting object features and relationship features respectively based on the scene graph includes: inputting the scene graph into a preset dual-tower GGNN encoder; the dual-tower GGNN encoder includes: a scene graph conversion module, a first GGNN network, a second GGNN network, a first fusion module, a second fusion module, and a third fusion module; converting the input scene graph into an object graph and a relationship graph by the scene graph conversion module, and inputting them into the first GGNN network and the second GGNN network respectively; extracting object features from the object graph by the first GGNN network, outputting them to the first fusion module; extracting relationship features from the relationship graph by the second GGNN network, outputting them to the second fusion module; The step of fusing the object features, relationship features, and attribute features of each target object to obtain comprehensive features of the target image includes: fusing the object features output by the first GGNN network with the attribute features of each target object by the first fusion module to obtain object fusion features and output them to the third fusion module; fusing the relationship features output by the second GGNN network with the original features of each relationship by the second fusion module to obtain relationship fusion features and output them to the third fusion module; fusing the object fusion features and the relationship fusion features by the third fusion module to obtain comprehensive features of the target image.

2. The method according to claim 1, wherein, The step of extracting object features and relationship features respectively based on the scene graph includes: converting the scene graph into an object graph and a relationship graph; wherein, the object graph is a directed graph with each target object as a node and each target association relationship as an edge; the relationship graph is a directed graph with each target association relationship as a node and each target object as an edge; extracting object features based on the object graph and extracting relationship features based on the relationship graph.

3. The method according to claim 1, wherein, The step of fusing the object features, relationship features, and attribute features of each target object to obtain comprehensive features of the target image includes: Fuse the object features with the attribute features of each target object to obtain object fusion features; and fuse the relationship features with the original features of each relationship to obtain relationship fusion features; Fuse the object fusion features with the relationship fusion features to obtain the comprehensive features of the target image.

4. The method according to claim 1, wherein, the first GGNN network includes: a first data conversion module, a first information transfer module, and a first hidden state graph generation module; the first data conversion module is used to convert the input object graph into a first data combination; the first data combination includes: the own feature vectors of all object nodes in the object graph, the own feature vectors of all edges in the object graph, the adjacency matrix of the incoming edges in the object graph, and the adjacency matrix of the outgoing edges in the object graph; the first information transfer module is used to, at any time step, for the object node to be updated in the object graph, calculate the first incoming edge information gain and the first outgoing edge information gain of the object node to be updated based on the adjacency matrix of the incoming edges and the adjacency matrix of the outgoing edges in the object graph; based on the first incoming edge information gain and the first outgoing edge information gain, obtain the first reference feature learned by the object node to be updated from all its incoming edges, outgoing edges, and adjacent nodes; obtain the updated hidden state of the object node to be updated based on the reference feature and the hidden state of the object node to be updated at the previous time step; the first hidden state graph generation module is used to, after iteration of a preset number of time steps, for each object node in the object graph, generate a first hidden state graph corresponding to the object graph based on the final hidden state of the object node output by the first information transfer module and the own feature vector of the object node, as the object feature; the second GGNN network includes: a second data conversion module, a second information transfer module, and a second hidden state graph generation module; the second data conversion module is used to convert the input relationship graph into a second data combination; the second data combination includes: the own feature vectors of all relationship nodes in the relationship graph, the own feature vectors of all edges in the relationship graph, the adjacency matrix of the incoming edges in the relationship graph, and the adjacency matrix of the outgoing edges in the relationship graph; the second information transfer module is used to, at any time step, for the relationship node to be updated in the relationship graph, calculate the second incoming edge information gain and the second outgoing edge information gain of the relationship node to be updated based on the adjacency matrix of the incoming edges and the adjacency matrix of the outgoing edges in the relationship graph; based on the second incoming edge information gain and the second outgoing edge information gain, obtain the second reference feature learned by the relationship node to be updated from all its incoming edges, outgoing edges, and adjacent nodes; obtain the updated hidden state of the relationship node to be updated based on the reference feature and the hidden state of the relationship node to be updated at the previous time step; The second hidden state graph generation module is configured to, after iteration for a preset number of time steps, for each relationship node in the relationship graph, generate a second hidden state graph corresponding to the relationship graph based on the final hidden state of the relationship node output by the second information transfer module and the own feature vector of the relationship node, as relationship features.

5. The method according to claim 4, wherein, the first information transfer module includes a first GRU network and a second GRU network; for the to-be-updated object node in the object graph, calculating a first incident edge information gain and a first outgoing edge information gain of the to-be-updated object node based on the adjacency matrix of the incident edges and the adjacency matrix of the outgoing edges in the object graph includes: inputting the hidden state of the to-be-updated object node at the previous time step, the adjacency matrix of the incident edges and the adjacency matrix of the outgoing edges in the object graph into the first GRU network to obtain the first incident edge information gain and the first outgoing edge information gain of the to-be-updated object node; obtaining a first reference feature learned by the to-be-updated object node from all its incident edges, outgoing edges and adjacent nodes based on the first incident edge information gain and the first outgoing edge information gain; obtaining the updated hidden state of the to-be-updated object node based on the reference feature and the hidden state of the to-be-updated object node at the previous time step includes: inputting the hidden state of the to-be-updated object node at the previous time step, the first incident edge information gain and the first outgoing edge information gain into the second GRU network, so that the second GRU network obtains the first reference feature learned by the to-be-updated object node from all its incident edges, outgoing edges and adjacent nodes based on the first incident edge information gain and the first outgoing edge information gain; obtaining the updated hidden state of the to-be-updated object node based on the first reference feature and the hidden state of the to-be-updated object node at the previous time step; the second information transfer module includes a third GRU network and a fourth GRU network; for the to-be-updated relationship node in the relationship graph, calculating a second incident edge information gain and a second outgoing edge information gain of the to-be-updated relationship node based on the adjacency matrix of the incident edges and the adjacency matrix of the outgoing edges in the relationship graph includes: inputting the hidden state of the to-be-updated relationship node at the previous time step, the adjacency matrix of the incident edges and the adjacency matrix of the outgoing edges in the relationship graph into the third GRU network to obtain the second incident edge information gain and the second outgoing edge information gain of the to-be-updated relationship node; obtaining a second reference feature learned by the to-be-updated relationship node from all its incident edges, outgoing edges and adjacent nodes based on the second incident edge information gain and the second outgoing edge information gain; obtaining the updated hidden state of the to-be-updated relationship node based on the reference feature and the hidden state of the to-be-updated relationship node at the previous time step includes: Input the hidden state, the second incident edge information gain, and the second outgoing edge information gain of the relationship node to be updated at the previous time step into the fourth GRU network, so that the fourth GRU network can obtain the second reference feature learned by the relationship node to be updated from all its incident edges, outgoing edges, and adjacent nodes based on the second incident edge information gain and the second outgoing edge information gain; and obtain the updated hidden state of the relationship node to be updated based on the second reference feature and the hidden state of the relationship node to be updated at the previous time step.

6. The method according to claim 4, wherein, the first hidden state graph generation module includes: a first MLP network and a first fully connected network; After iterating for a preset number of time steps, for each object node in the object graph, generating a first hidden state graph corresponding to the object graph based on the final hidden state of the object node output by the first information transfer module and the own feature vector of the object node, includes: For each object node in the object graph, input the final hidden state of the object node output by the first information transfer module and the own feature vector of the object node into the first MLP network, so that the first MLP network generates the first hidden state graph corresponding to the object graph and inputs it into the first fully connected network, so that the first fully connected network outputs the first hidden state graph; the second hidden state graph generation module includes: a second MLP network and a second fully connected network; After iterating for a preset number of time steps, for each relationship node in the relationship graph, generating a second hidden state graph corresponding to the relationship graph based on the final hidden state of the relationship node output by the second information transfer module and the own feature vector of the relationship node, includes: For each relationship node in the relationship graph, input the final hidden state of the relationship node output by the second information transfer module and the own feature vector of the relationship node into the second MLP network, so that the second MLP network generates the second hidden state graph corresponding to the relationship graph and inputs it into the second fully connected network, so that the second fully connected network outputs the second hidden state graph.

7. The method according to claim 1, wherein, the step of converting the target problem into a problem feature includes: Input the target problem into a preset problem encoding model, and convert the text of the target problem into a vector as the problem feature.

8. The method according to claim 1, wherein, the attribute features of each target object are obtained through the following steps: Obtain the preset target attribute texts of each target object; Convert each target attribute text into an information word vector as the attribute feature of each target object.

9. The method according to claim 1, wherein, the step of predicting an answer to the target problem based on the problem feature and the comprehensive feature of the target picture to obtain the answer to the target problem includes: Calculate and obtain an attention feature based on the problem feature and the comprehensive feature; Fuse the attention feature and the problem feature to obtain a feature to be predicted; Input the to-be-predicted feature into a preset prediction model to obtain a target answer selected by the prediction model from multiple candidate answers preset for the target picture as the answer to the target question.

10. An implementation device for visual question answering, comprising: a picture question acquisition unit configured to acquire a specified target picture and a target question for the target picture; a question conversion unit configured to convert the target question into a question feature; a feature extraction unit configured to perform object feature extraction and relationship feature extraction on the target picture to obtain an object feature and a relationship feature respectively; a feature fusion unit configured to fuse the object feature, the relationship feature, and the attribute features of each target object to obtain a comprehensive feature of the target picture; an answer prediction unit configured to perform answer prediction based on the question feature and the comprehensive feature of the target picture to obtain the answer to the target question; the feature extraction unit is configured to perform feature recognition on the target picture, and based on the recognition result, generate a directed graph including target objects and target association relationships as a scene graph; Based on the scene graph, extract the object feature and the relationship feature respectively; The feature extraction unit is configured to: input the scene graph into a preset two-tower GGNN encoder; the two-tower GGNN encoder includes: a scene graph conversion module, a first GGNN network, a second GGNN network, a first fusion module, a second fusion module, and a third fusion module; the scene graph conversion module converts the input scene graph into an object graph and a relationship graph, and inputs them into the first GGNN network and the second GGNN network respectively; the first GGNN network performs feature extraction on the object graph to obtain an object feature and outputs it to the first fusion module; the second GGNN network performs feature extraction on the relationship graph to obtain a relationship feature and outputs it to the second fusion module; The feature fusion unit is configured to: the first fusion module fuses the object feature output by the first GGNN network with the attribute features of each target object to obtain an object fusion feature and outputs it to the third fusion module; the second fusion module fuses the relationship feature output by the second GGNN network with the original features of each relationship to obtain a relationship fusion feature and outputs it to the third fusion module; the third fusion module fuses the object fusion feature and the relationship fusion feature to obtain a comprehensive feature of the target picture.

11. The device according to claim 10, wherein the feature extraction unit is configured to convert the scene graph into an object graph and a relationship graph; wherein, the object graph is a directed graph with each target object as a node and each target association relationship as an edge; the relationship graph is a directed graph with each target association relationship as a node and each target object as an edge; extract the object feature based on the object graph and extract the relationship feature based on the relationship graph.

12. The device according to claim 10, wherein the feature fusion unit is configured to fuse the object feature with the attribute features of each target object to obtain an object fusion feature; and fuse the relationship feature with the original features of each relationship to obtain a relationship fusion feature; Fuse the object fusion feature and the relationship fusion feature to obtain the comprehensive feature of the target image.

13. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-9.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, the computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.

15. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Image question-answering method and device, computer equipment and medium

    CN111782840A

  • Scene graph generation method and device, computer readable medium and electronic equipment

    CN113568983A