A visual question answering method and system based on a visual question answering model with division of labor decision-making

By introducing a division of labor decision module into the image question and answer model, the processing image and text features are separated and processed, and fusion is carried out when answers are generated, the problem of cross-modal fusion in the existing technology is solved, and the model's understanding and reasoning ability is improved.

CN114283292BActive Publication Date: 2025-05-06XIAMEN JUDING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111483361.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-07
Publication Date
2025-05-06
Estimated Expiration
2041-12-07

AI Technical Summary

Technical Problem

When existing image question-and-answer models merge across modalities, it is difficult to effectively combine the high-level semantic information of the image with the semantic information of the problem text, resulting in increased inference difficulty.

Method used

A visual question-and-answer model based on division of labor decisions is proposed. Image and text features are obtained through feature acquisition module, and separation processing and iterative updates are performed in the division of labor decision module, and answers are generated in the answer output module.

Benefits of technology

By separating and processing visual spatial information and fusion at the end, the inference difficulty brought about by cross-modal fusion is reduced, and the understanding and reasoning ability of the question-and-answer model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114283292B_ABST
    Figure CN114283292B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of image question answering, and specifically relates to a visual question answering method and system of a visual question answering model based on division of labor decision-making, the method comprising: obtaining a visual image and a question to be answered, inputting the visual image and the question to be answered into an LRBNet model, and obtaining a question answering result; the LRBNet model comprises a visual understanding module, a text understanding module and an exchange module; the visual understanding module is used to obtain a visual feature map, the text understanding module is used to obtain a text feature map, the exchange module is used to perform data interaction between the visual feature map and the text feature map, and update the node according to the interaction data; the visual space feature map and the text semantic information are associated and updated to obtain a final question answering result; the present invention processes the text semantic information and the visual space information separately, and only fuses the processed results at the end, thereby reducing the reasoning difficulty of other VQA models that is increased due to cross-modal fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image question answering, and specifically relates to a visual question answering method and system based on a visual question answering model with division of labor decision-making. Background Art

[0002] Deep learning algorithms have achieved great success in both vision-related and language-related tasks, and visual question answering is a task that tests the cross-modal understanding of vision and language. In the most common form of visual question answering, the computer presents an image and a text question about the image, and the visual question answering algorithm needs to give an answer to the question based on the image.

[0003] At present, most image question answering (VQA) models use neural networks such as recurrent neural networks (RNNs) and long short-term memory networks (LSTM) to learn the encoding representation of questions. In order to encode images, early VQA models used neural networks pre-trained on ImageNet such as Resnet or VGG to extract visual information from images. In order to obtain the features of different regions in the image and reduce the impact of invalid regions on answer prediction, Teney et al. proposed a BUTD model that uses Faster R-CNN to detect objects in the image, obtain the features of related regions, and calculate the attention weight of each region according to the question encoding to predict the answer. Lu et al. proposed a hierarchical collaborative attention network, which not only realizes question-guided visual attention, but also realizes visual-guided question attention. Pan Lu et al. constructed the Relation-VQA dataset to directly mine VQA-specific relations and provide additional semantic information for the model. The above VQA system can be roughly divided into four modules: question encoding module, image encoding module, cross-modal fusion module and question prediction module. The question encoding module usually uses RNN, LSTM and other models to embed questions into vectors; the image encoding module first uses the Faster R-CNN model to extract image features, and then adds or concatenates the question encoding and image features for joint encoding and relational modeling to learn the relationship between text and image, and obtain joint features. The cross-modal fusion module fuses the question encoding and joint features, and finally inputs them into the question prediction module for answer prediction.

[0004] The above existing technologies only focus on the cross-modal fusion of images and texts. Although they involve cross-modal conversion, they do not jointly encode the high-level semantic information of the image and the semantic information of the question text, which makes the model more difficult to reason due to cross-modal fusion. Summary of the invention

[0005] In order to solve the problems existing in the above prior art, the present invention proposes a visual question answering system based on a division of labor decision-making visual question answering model, the system comprising: a feature acquisition module, a division of labor decision-making module and an answer output module;

[0006] The feature acquisition module is used to acquire the visual features of the image and the text features of the question, and input them into the division of labor decision module;

[0007] The division of labor decision module includes a preprocessing module, a visual understanding module, a text understanding module, an exchange module and an answer prediction module;

[0008] The preprocessing module is used to convert the question text into visual features, and extract local visual features and local text information of the image, input the visual features converted from the question text and the local visual features of the image into the visual understanding module, and input the local text information into the text understanding module;

[0009] The visual understanding module is used to process the output from the preprocessing module, obtain a visual feature map after screening, graph construction and spatial relationship modeling, and input it into the exchange module;

[0010] The text understanding module is used to process text information. After screening, counting and semantic relationship modeling, the obtained text feature map is input into the exchange module, and the one-hot vector of the counting result is input into the question prediction module; the text information includes the question text and the local text information of the image from the data preprocessing module;

[0011] The exchange module is used to perform data exchange between the visual understanding module and the text understanding module, receive the visual feature map from the visual understanding module and the text feature map from the text understanding module, perform one or more rounds of iterative updates on the visual feature map and the text feature map through data exchange, and feed back the visual feature map and the text feature map of the last round of iterative updates to the visual understanding module and the text understanding module respectively;

[0012] The question prediction module is used to obtain the updated text feature map, updated visual feature map and one-hot vector in the text understanding module and the visual understanding module, and obtain the answer to the question based on the obtained features;

[0013] The answer output module is used to output the answer to the question obtained by the question prediction module.

[0014] Preferably, the process of converting the question text into visual features by the data preprocessing module includes: using the text-image network DM-GAN to transform the image-related questions in the training set to obtain the image of the question, and using the ResNet50 network to extract features from the converted image to obtain the visual feature Q2I feature related to the question.

[0015] Preferably, the process of processing visual information by the visual understanding module includes: using the bounding box clipping module Bounding Box Clipping and the matrix creation module Adjacency Matrix Creating to screen and construct graphs of local image features Imagefeatures and Q2I features to obtain an adjacency matrix and a visual feature graph; splicing the Imagefeatures and the Q2I feature and inputting them together with the adjacency matrix into the spatial relationship learning module SpatialRelation Learning to perform spatial relationship modeling; using the residual connection module Add&Norm to add and normalize the visual features after relationship modeling to the features before modeling to obtain visual spatial features.

[0016] Preferably, the process of using the text understanding module to process the text information in the training set includes: using LSTM to encode the text information Image captions and question text Question of the image; using the bounding box clipping module Bounding Box Clipping and the adjacency matrix construction module Create adjacency matrix to screen and construct the encoded Image captions and Question to obtain the adjacency matrix and the text feature map; sending the screening results to the Count module for counting to obtain C; splicing the encoded Image captions and Question and inputting them together with the adjacency matrix into the semantic relationship learning module Semantic Relation Learning for semantic relationship modeling; using the Add&Norm module to add and normalize the features after relationship modeling and the features before modeling to obtain text semantic features.

[0017] Preferably, the process of updating the visual feature map and the text feature map by the exchange module includes: respectively obtaining the feature value sets of each node in the visual feature map and the text feature map, using the two feature value sets to calculate the attention coefficient matrix between the two feature maps, using the attention coefficient matrix to perform weighted averaging on each node of the two feature maps, and updating each node using feature linear modulation.

[0018] Preferably, the process of processing the features by the question prediction module includes: using the attention mechanism to calculate the attention coefficient of the question text feature and the text semantic feature, and performing weighted average of the attention coefficient and the text semantic feature to obtain the text semantic embedding cap emb , cap embThe result is sent to the multi-layer perceptron MLP to obtain the probability p2 predicted by the text understanding module. The attention mechanism is used to calculate the attention coefficient of the visual feature Q2I feature and the visual space feature of the problem transformation, and the attention coefficient and the visual space feature are weighted averaged to obtain the visual space embedding V. emb , V emb Send it to the multi-layer perceptron MLP to get the probability p3 predicted by the visual understanding module; emb ,V emb And C are concatenated and sent to the multi-layer perceptron MLP to obtain the probability p1 of joint embedding prediction.

[0019] A visual question answering method based on a visual question answering model with division of labor decision-making, comprising: obtaining a visual image and a question to be answered, inputting the visual image and the question to be answered into the image visual question answering model with division of labor decision-making, and obtaining a question answering result; the image visual question answering model with division of labor decision-making comprises a visual understanding module, a text understanding module and an exchange module, and the visual understanding module, the text understanding module and the exchange module cooperate in answering questions;

[0020] The process of training the image visual question answering model based on division of labor decision-making includes:

[0021] S1: Obtain an original question-answering visual image set, preprocess the data in the original question-answering visual image set, and divide the preprocessed question-answering visual image set into a training set and a test set;

[0022] S2: Input the data in the training set into the LRBNet model for training;

[0023] S3: Convert the question text in the training set into visual features and extract local visual features and local text information of the image;

[0024] S4: Use the visual understanding module to process the visual information and obtain visual spatial features;

[0025] S5: Use the text understanding module to process the text information and obtain the text semantic features;

[0026] S6: Use the exchange module to iteratively update the visual feature map and text feature map;

[0027] S7: The answer prediction module is used to process the visual spatial features and text semantic features to predict the answer of the visual image;

[0028] S8: Calculate the model's loss function based on the predicted visual image answer;

[0029] S9: Input the data in the test set into the model, continuously adjust the model parameters, and complete the model training when the loss function is minimized.

[0030] Preferably, the process of converting the question text into visual features includes: using the text-image network DM-GAN to convert the image-related questions in the training set to obtain the image of the question, and using the ResNet50 network to extract features from the converted image to obtain visual features related to the question.

[0031] Preferably, the process of updating the visual feature map and the text feature map using the exchange module includes: the exchange module uses attention-based feature linear modulation to iteratively update the visual feature map and the text feature map to connect them.

[0032] Preferably, the loss function expression of the model is:

[0033]

[0034] L=αL1+βL2+(1-α-β)L3

[0035] in, N represents the number of samples, y j represents the true label, ρ(x) is the sigmoid function, and p ·j The answer is y j , p1 is the joint probability, p2 is the probability predicted by the text understanding module, p3 is the probability predicted by the visual understanding module, α, β represent the weights of losses L1 and L2, L1, L2, L3 are the joint embedding multi-label classification loss, text embedding multi-label classification loss and visual embedding multi-label classification loss respectively.

[0036] Beneficial effects of the present invention:

[0037] The LRBNet model constructed by the present invention is provided with Right module, Left module and Change Module. The Left and Right modules process text semantic information and visual spatial information respectively, and communicate through the Change Module, which takes into account both the understanding ability and the reasoning ability. At the same time, the model processes the text semantic information and the visual spatial information separately, and only fuses the processing results at the end, reducing the reasoning difficulty increased by cross-modal fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a structural diagram of the LRBNet model of the present invention. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] A visual question answering method based on a visual question answering model with division of labor decision-making, the method comprising: acquiring a visual image and a question to be answered, inputting the visual image and the question to be answered into the image visual question answering model with division of labor decision-making, and obtaining a question answering result; the image visual question answering model with division of labor decision-making comprises a visual understanding module, a text understanding module and an exchange module, and the visual understanding module, the text understanding module and the exchange module cooperate in answering questions in a division of labor decision-making.

[0041] like Figure 1 As shown in FIG. 1 , the process of training the image visual question answering model based on division of labor decision-making includes:

[0042] S1: Obtain an original question-answering visual image set, preprocess the data in the original question-answering visual image set, and divide the preprocessed question-answering visual image set into a training set and a test set;

[0043] S2: Input the data in the training set into the LRBNet model for training;

[0044] S3: Convert the question text in the training set into visual features and extract local visual features and local text information of the image;

[0045] S4: Use the visual understanding module to process the visual information and obtain visual spatial features;

[0046] S5: Use the text understanding module to process the text information and obtain the text semantic features;

[0047] S6: Use the exchange module to iteratively update the visual feature map and text feature map;

[0048] S7: The answer prediction module is used to process the visual spatial features and text semantic features to predict the answer of the visual image;

[0049] S8: Calculate the model's loss function based on the predicted visual image answer;

[0050] S9: Input the data in the test set into the model, continuously adjust the model parameters, and complete the model training when the loss function is minimized.

[0051] The process of preprocessing the data in the original question-answer visual image set includes: filtering, completing and cleaning the images in the original question-answer visual image set to obtain complete and high-definition visual images; and cleaning the text question-answer data in the original question-answer visual image set to obtain text question-answer data with complete information.

[0052] The data preprocessing module is used to preprocess data. First, the image is passed through the Densecap network for object detection to obtain the image features of M objects in the image and a text description of the object Image captions; the questions related to the image are passed through the DM-GAN network, which converts the question text into an image related to the object mentioned in the question. The image is then passed through ResNet50 to extract the image feature Q2I feature.

[0053] The Visual Understand Module (Right) is used to process image information. The Image features and Q2I features are filtered out by the Bounding Box Clipping module to obtain N Image features that are closely related to the Q2I feature. These N Image features are constructed into a graph structure by the Create adjacency matrix module to obtain an adjacency matrix. After that, the Image features and Q2I features are connected and sent to Spatialrelation Learning for relationship modeling. Then, they are sent to the Add&Norm module to obtain the final output V.

[0054] The textual understanding module (Textual Understand Module, Left) is used to process text information. This module has the same structure as the visual understanding module (Visual Understand Module (Right). First, the image captions and question text obtained in the data preprocessing module are encoded by LSTM, and the encoded features are also passed through the Bounding Box Clipping, Create adjacency matrix, Semantic relation Learning and Add&Norm modules to finally get the output T. The only difference is that in the Textual Understand Module (Left), there is an additional counting module branch after Bounding Box Clipping to specifically deal with counting problems.

[0055] The Change Module is used to handle the interaction between the Left Module and the Right Module. The Left Module and the Right Module complete text understanding and image understanding respectively, but the two modules still need to communicate and work together to complete reasoning. The Change Module plays the role of the corpus callosum connecting the left and right brains, providing a communication channel for the Left Module and the Right Module.

[0056] The answer prediction module (Answer Predictor) is used to predict the answer. The output V of the model's Right module, the output T of the Left module and the output C of the counting module are connected and sent to the Answer Predictor to predict the answer.

[0057] Specifically, the process of using the data preprocessing module to process the visual images in the training set includes: using the dense caption generation network Densecap to perform target detection on the visual image to obtain the visual features Imagefeatures of M objects and the text information Image captions of each image; using the text-image network DM-GAN to transform the image-related questions in the training set to obtain the image of the question, and using the ResNet50 network to extract features from the transformed image to obtain the visual feature Q2I feature of the image generated by the DM-GAN network and extracted by the residual network.

[0058] The Visual Understand Module (Right) uses Densecap to extract a set of Object Region Proposals (target region proposals), including the features of Regions. And the corresponding text description. V contains the Bounding box feature (x, y, w, h) represent the coordinates, width and height of the bounding boxes respectively. The question text is passed through the DM-GAN network and ResNet to obtain image features related to the question.

[0059] The Bounding Box Clipping module receives V and Q and calculates v i The similarity with Q is mapped to [0, 1] using the sigmoid function, and the similarity is i Multiply to achieve the purpose of screening V, the formula is as follows:

[0060] v i ′=a i v i , i=1,2,...,n

[0061] a i =sigmoid([v i w v ] T Q q ]), i = 1, 2, ..., n

[0062] Among them, v i ′ represents the visual features after screening, a i Indicates v i Similarity with Q, v i represents the visual features before screening, n represents n objects, sigmoid represents the nonlinear silver snake function, w v Indicates v i , T represents the transpose, Q represents the visual features extracted by the residual network from the image generated by the DM-GAN network, and w q Represents the parameter of parameter Q. in represents the set of real numbers, d v Indicates v i The dimension, d out Indicates v i The dimension of ′, d Q represents the dimension of Q.

[0063] The Create adjacency matrix module receives the filtered V and Q and constructs the adjacency matrix:

[0064] e i =F([v i ′||Q]), i=1, 2,...,n

[0065] in is a nonlinear function, then the adjacency matrix can be expressed as

[0066] A=sigmoid(EE T )

[0067] Among them, A represents the adjacency matrix, sigmoid represents the activation function, and E represents the i The matrix composed of.

[0068] The spatial relation learning module receives the concatenated features of V and Q and performs relation modeling. It can make any graph neural network:

[0069] V out =Gr([V||Q])+V

[0070] Among them, Gr represents the spatial relationship learning module.

[0071] The Textual Understand Module (Left) receives the question text and the Image captions from the Right module, and obtains the descriptive text features through LSTM encoding. and question text features

[0072] The Bounding Box Clipping module receives q and Cap and calculates cap i The similarity with q is mapped to [0, 1] using the sigmoid function, and the similarity with v i Multiply to achieve the purpose of screening Cap, the formula is as follows:

[0073] cap i ′=a i cap i , i=1,2,...,n

[0074] cap i =sigmoid([cap i w cap ] T [qw q ]), i = 1, 2, ..., n

[0075] in cap i ′ represents the text features after screening, a i Indicates cap i Similarity with q, cap i represents the text feature, w cap Indicates the parameters of capi, w q Represents the parameter of q.

[0076] The Create adjacency matrix module receives the filtered Cap and q and constructs the adjacency matrix:

[0077] e i =F([cap i ||q]), i=1, 2, ..., n

[0078] in is a nonlinear function, then the adjacency matrix can be expressed as:

[0079] A=sigmoid(EE T )

[0080] The Semantic relation Learning module receives the concatenated features of Cap and q and performs relation modeling. It can make any graph neural network:

[0081] Cap out =Gl([Cap||q])+Cap

[0082] Among them, Cap out Represents the text features after modeling, Gl represents the semantic relationship learning module, and Cap represents the text features before modeling.

[0083] The Change Module is used to connect the visual space feature map and text semantic information. In order to avoid cross-modal fusion, the communication between modules adopts the feature linear modulation (FiLM) method. Take the linear modulation of the Right module to the Left module as an example. out As input, the attention mechanism is used to obtain out,i Most relevant objects:

[0084]

[0085]

[0086] Among them, α ij Indicates cap out,i and v out,j Similarity, softmax represents a nonlinear function, W Q Indicates cap out,i Parameters, cap out,i represents the i-th text feature, W K Indicates v out,j The parameter, v out,j represents the jth image visual feature, d out Represents the dimension of the output, and n represents n objects.

[0087] Using the learnable network g, with v′ out,i As input, the linear modulation parameters β, γ are obtained; their expressions are:

[0088] β i,c , γ i,c =g(v′ out,i ; θ), i = 1, 2, ..., n

[0089] Cap is linearly modulated using β and γ, and the modulation formula is:

[0090] cap' i =σ(γ i,c ⊙W i capout,i +β i,c ), i = 1, 2, ..., n

[0091] Among them, θ represents the learnable network g parameter, n represents n objects, σ ​​represents the nonlinear function softmax, γ i,c represents the linear modulation parameter, ⊙ represents element-by-element multiplication, W i Indicates cap out,i The parameter, β i,c Indicates linear modulation parameters.

[0092] The Count module adopts the Counting module of Yan Zhang et al. Receive the similarity a of Bounding Box Clipping and obtain the similarity matrix A=aa T , and get the output Among them, Counting module represents the counting module.

[0093] Answer Predictor receives the output Cap' of the above modules out , V' out and C, using the attention mechanism to calculate the text semantic embedding and visual space embedding, which is expressed as:

[0094]

[0095]

[0096] in:

[0097]

[0098]

[0099] Among them, α i and β i represents the attention coefficient, W1 represents Cap' out,i Parameters, W1 represents Cap' out,i , W2 represents the parameter of q, and W3 represents V' out,i , W4 represents the parameter of Q, d out Cap' out,i and V' out,i Dimension.

[0100] The multi-layer perceptron is used to predict the answer, and the formula is as follows

[0101] p1=σ(MLP1([cap emb ||V emb ||C]))

[0102] p2=σ(MLP2(cap emb ))

[0103] p3=σ(MLP3(V emb ]))

[0104] Among them, p i , i = 1, 2, 3 represents the probability of the predicted answer, MLP i represents a multilayer perceptron, σ represents the softmax function, p1 is the joint probability, p2 is the probability predicted by the Left module, and p3 is the probability predicted by the Right module.

[0105] The loss function expression of the model is:

[0106]

[0107] L=αL1+βL2+γL3

[0108]

[0109] Where N represents the number of samples, y j represents the true label, ρ(x) is the sigmoid function, and p ·j The answer is y j , α, β, γ represent the weights of losses L1, L2, and L3, where L1, L2, L3 are the joint embedding multi-label classification loss, text embedding multi-label classification loss, and visual embedding multi-label classification loss, respectively.

[0110] A visual question answering system based on a visual question answering model of division of labor decision-making, the system comprising: the system comprising: a feature acquisition module, a division of labor decision-making module and an answer output module;

[0111] The feature acquisition module is used to acquire the visual features of the image and the text features of the question, and input them into the division of labor decision module;

[0112] The division of labor decision module includes a preprocessing module, a visual understanding module, a text understanding module, an exchange module and an answer prediction module;

[0113] The preprocessing module is used to convert the question text into visual features, and extract local visual features and local text information of the image, input the visual features converted from the question text and the local visual features of the image into the visual understanding module, and input the local text information into the text understanding module;

[0114] The visual understanding module is used to process the output from the preprocessing module, obtain a visual feature map after screening, graph construction and spatial relationship modeling, and input it into the exchange module;

[0115] The text understanding module is used to process text information. After screening, counting and semantic relationship modeling, the obtained text feature map is input into the exchange module, and the one-hot vector of the counting result is input into the question prediction module; the text information includes the question text and the local text information of the image from the data preprocessing module;

[0116] The exchange module is used to perform data exchange between the visual understanding module and the text understanding module, receive the visual feature map from the visual understanding module and the text feature map from the text understanding module, perform one or more rounds of iterative updates on the visual feature map and the text feature map through data exchange, and feed back the visual feature map and the text feature map of the last round of iterative updates to the visual understanding module and the text understanding module respectively;

[0117] The question prediction module is used to obtain the updated text feature map, updated visual feature map and one-hot vector in the text understanding module and the visual understanding module, and obtain the answer to the question based on the obtained features;

[0118] The answer output module is used to output the answer to the question obtained by the question prediction module.

[0119] The specific implementation of the system of the present invention is the same as the method embodiment of the present invention.

[0120] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation modes of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A visual question answering system based on a visual question answering model with division of labor decision-making, characterized in that: The system includes: a feature acquisition module, a division of labor decision module and an answer output module; The feature acquisition module is used to acquire the visual features of the image and the text features of the question, and input them into the division of labor decision module; The division of labor decision module includes a preprocessing module, a visual understanding module, a text understanding module, an exchange module and an answer prediction module; The preprocessing module is used to convert the question text into visual features, and extract local visual features and local text information of the image, input the visual features converted from the question text and the local visual features of the image into the visual understanding module, and input the local text information into the text understanding module; the process of converting the question text into visual features by the data preprocessing module includes: using the text-image network DM-GAN to convert the questions related to the image in the training set to obtain the image of the question, and using the ResNet50 network to extract features from the converted image to obtain the visual feature Q2Ifeature related to the question; The visual understanding module is used to process the output from the preprocessing module, and obtains a visual feature map after screening, graph construction and spatial relationship modeling, and inputs it into the exchange module; specifically, it includes: using the bounding box clipping module Bounding Box Clipping and the matrix creation module Adjacency Matrix Creating to screen and graph the local features of the image Image features and Q2I feature to obtain the adjacency matrix and the visual feature map; splicing the Image features and the Q2I feature and inputting them together with the adjacency matrix into the spatial relationship learning module Spatial Relation Learning to perform spatial relationship modeling; using the residual connection module Add&Norm to add and normalize the visual features after the relationship modeling and the features before the modeling to obtain the visual spatial features; The text understanding module is used to process text information, and after screening, counting and semantic relationship modeling, the obtained text feature map is input into the exchange module, and the one-hot vector of the counting result is input into the question prediction module; the text information includes the question text and the local text information of the image from the data preprocessing module; specifically, it includes: using LSTM to encode the text information Image captions and the question text Question of the image; using the bounding box clipping module Bounding Box Clipping and the adjacency matrix construction module Create adjacency matrix to screen and construct the encoded Image captions and Question to obtain the adjacency matrix and the text feature map; sending the screening result to the Count module for counting to obtain C; splicing the encoded Image captions and Question and inputting them together with the adjacency matrix into the semantic relationship learning module Semantic Relation Learning for semantic relationship modeling; using the Add&Norm module to add and normalize the features after relationship modeling with the features before modeling to obtain text semantic features; The exchange module is used to perform data exchange between the visual understanding module and the text understanding module, receive the visual feature map from the visual understanding module and the text feature map from the text understanding module, perform one or more rounds of iterative updates on the visual feature map and the text feature map through data exchange, and feed back the visual feature map and the text feature map of the last round of iterative updates to the visual understanding module and the text understanding module respectively; the iterative update specifically includes: obtaining the feature value set of each node in the visual feature map and the text feature map respectively, calculating the attention coefficient matrix between the two feature maps using the two feature value sets, performing weighted average with each node of the two feature maps using the attention coefficient matrix, and updating each node using feature linear modulation; The question prediction module is used to obtain the updated text feature map, updated visual feature map and one-hot vector in the text understanding module and the visual understanding module, and obtain the answer to the question based on the obtained features; specifically, it includes: using the attention mechanism to calculate the attention coefficient of the question text feature and the text semantic feature, and performing weighted average of the attention coefficient and the text semantic feature to obtain the text semantic embedding cap emb , cap emb The result is sent to the multi-layer perceptron MLP to obtain the probability p2 predicted by the text understanding module. The attention mechanism is used to calculate the attention coefficient of the visual feature Q2I feature and the visual space feature of the problem transformation, and the attention coefficient and the visual space feature are weighted averaged to obtain the visual space embedding V. emb , V emb Send it to the multi-layer perceptron MLP to get the probability p3 predicted by the visual understanding module; emb ,V emb And C is concatenated and sent to the multi-layer perceptron MLP to obtain the probability p1 of joint embedding prediction; The answer output module is used to output the answer to the question obtained by the question prediction module.

2. A visual question answering method based on a visual question answering model with division of labor decision, the method adopts a visual question answering system based on a visual question answering model with division of labor decision according to claim 1, characterized in that: include: Obtaining visual images and questions to be answered, inputting the visual images and questions to be answered into an image visual question answering model based on division of labor decision, and obtaining question answering results; The image visual question answering model based on division of labor decision-making includes a visual understanding module, a text understanding module and an exchange module. The visual understanding module, the text understanding module and the exchange module work together to answer questions. The process of training the image visual question answering model based on division of labor decision-making includes: S1: Obtain an original question-answering visual image set, preprocess the data in the original question-answering visual image set, and divide the preprocessed question-answering visual image set into a training set and a test set; S2: Input the data in the training set into the LRBNet model for training; S3: Convert the question text in the training set into visual features and extract local visual features and local text information of the image; S4: Use the visual understanding module to process the visual information and obtain visual spatial features; S5: Use the text understanding module to process the text information and obtain the text semantic features; S6: Use the exchange module to iteratively update the visual feature map and text feature map; S7: The answer prediction module is used to process the visual spatial features and text semantic features to predict the answer of the visual image; S8: Calculate the loss function of the model based on the predicted visual image answer; S9: Input the data in the test set into the model, continuously adjust the model parameters, and complete the model training when the loss function is minimized.

3. The visual question answering method based on the visual question answering model of division of labor decision according to claim 2, characterized in that: The process of converting question text into visual features includes: using the text-image network DM-GAN to transform the image-related questions in the training set to obtain the image of the question, and using the ResNet50 network to extract features from the converted image to obtain visual features related to the question.

4. The visual question answering method based on the visual question answering model of division of labor decision according to claim 2, characterized in that: The process of updating the visual feature map and the text feature map by using the exchange module includes: the exchange module uses attention-based feature linear modulation to iteratively update the visual feature map and the text feature map to connect them.

5. The visual question answering method based on the visual question answering model of division of labor decision according to claim 2, characterized in that: The loss function expression of the model is: L=αL1+βL2+(1-α-β)L3 in, N represents the number of samples, y j represents the true label, ρ(x) is the sigmoid function, and p· j The answer is y j , p1 is the joint probability, p2 is the probability predicted by the text understanding module, p3 is the probability predicted by the visual understanding module, α, β represent the weights of losses L1 and L2, L1, L2, L3 are the joint embedding multi-label classification loss, text embedding multi-label classification loss and visual embedding multi-label classification loss respectively.

Citation Information

Patent Citations

  • Visual question and answer method based on multi-modal fusion and structural control

    CN113010656A

  • Visual question answering method and device based on question semantic mapping

    CN113420833A