A Visual Question Answering Method Based on a Module Routing Network Model

Through the module routing network model, the integration of vision and text modes is solved, and the problem of fusion difficulties in the existing technology is realized, and the answers to complex questions under unsupervised conditions is improved, which is the generalization and universality of the model.

CN114138946BActive Publication Date: 2025-07-08FUDAN UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010811525.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-13
Publication Date
2025-07-08
Estimated Expiration
2040-08-13

AI Technical Summary

Technical Problem

The prior art is difficult to integrate visual and text modalities at multiple semantic levels and requires expert knowledge and supervised information, resulting in limited generalization and universality.

Method used

The module routing network model is adopted to extract problem features through the text network, and the module layer in the visual network is activated using the routing path. The final features are generated by combining the visual network and the routing network, and the answer is generated by a predetermined answer. The training loss function includes visual question-and-answer loss and load balancing loss.

Benefits of technology

It realizes the integration of visual and text modalities at multiple semantic levels, and can answer complex questions without expert knowledge and supervised information, improving generalization and universality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114138946B_ABST
    Figure CN114138946B_ABST
Patent Text Reader

Abstract

The visual question answering method based on the module routing network provided by the present invention is used to solve the problem of processing natural language question text and input question photos according to the module routing network model and generating question answers. It is characterized in that the module routing network model has a text network, a routing network and a visual network. The method includes the following steps: Step 1, input the natural language question text into the text network to extract question features; Step 2, activate the corresponding module in the visual network based on the routing path generated at least based on the question features to become an activated module, and input the question photo into the visual network. The activated module extracts image features from the question photo to form corresponding final features; Step 3, input the final features into an answerer to generate question answers. The method of the present invention fuses the text and visual modalities at multiple levels, and does not require expert knowledge and supervision information when answering complex questions, and can be widely applied to situations that require the combination of multiple modalities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a visual question answering method based on a module routing network model, belonging to the field of artificial intelligence and used for solving visual question answering tasks. Background Art

[0002] Historically, computer vision and natural language processing have been developing independently as separate research directions. With the revival of neural networks, new research tasks have been emerging in these two fields. Among them, some tasks connecting the two fields have been proposed, and the visual question answering task involved in the present invention is one of them [1].

[0003] Visual question answering means that given an image and a question pair, the model needs to answer the question based on the content of the image. Compared with classical computer vision tasks such as image recognition, detection, and segmentation, the visual question answering method based on a module routing network model provided by the present invention can fuse the visual and text modalities more fully. Therefore, the module routing network model provided by the present invention has a certain "understanding" of images and can answer questions better.

[0004] Visual question answering faces two core problems. One is how to better fuse the visual and text modalities, and the other is how to enable the model to have a certain visual reasoning ability to answer more complex questions.

[0005] To solve the first problem, since most existing works [2, 3] are based on such a mode: first, use a convolutional neural network and a recurrent neural network to extract the features of the image and the question respectively, and then perform feature-level fusion on these two features. However, since the objects of fusion, that is, the features extracted separately, are already at a high semantic level (because a widely recognized view is that the features extracted by the neural network are at a higher semantic level as it progresses), the fusion only occurs at a high semantic level and fails to fuse at multiple semantic levels.

[0006] To solve the second problem, the current mainstream method is based on neural module networks [4 - 8]. This method believes that questions are compositional, so the answer to a question can be decomposed into a series of sub-question answers. Therefore, this method first needs to decompose the question, then design exclusive modules for each sub-question, and finally use a neural network to learn a layout method for organizing these modules from the question, then organize the modules according to the layout method to form a model, and finally use the model to process the input image. However, both decomposing the question and designing the modules require expert knowledge, and most of the time this series of methods require additional expensive supervision information. These two drawbacks have a certain impact on its generalization and universality.

[0007] In summary, the image features and problem features extracted by convolutional neural networks and recurrent neural networks cannot be fused at multiple semantic levels, and the inherent complex processing method of neural module networks leads to the impact on their generalization and universality.

[0008] [1] ANTOLS, AGRAWALA, L U J, et al. Vqa: Visual question answering[C] / / ICCV. 2015.

[0009] [2] BEN-YOUNES H, CADENE R, CORD M, et al. Mutan: Multimodal tucker fusion for visual question answering[C] / / ICCV. 2017.

[0010] [3] YANG Z, HE X, GAO J, et al. Stacked attention networks for image question answering[C] / / CVPR. 2016:21-29.

[0011] [4] ANDREAS J, ROHRBACH M, DARRELL T, et al. Neural module networks[C] / / CVPR. 2016.

[0012] [5] JOHNSON J, HARIHARAN B, VAN DER MAATEN L, et al. Inferring and executing programs for visual reasoning[C] / / ICCV. 2017.

[0013] [6] HU R, ANDREAS J, ROHRBACH M, et al. Learning to reason: End-to-end module networks for visual question answering[C] / / ICCV. 2017.

[0014] [7]MASCHARKA D,TRAN P,SOKLASKI R,et al.Transparency by design:Closing the gap between performance and interpretability in visual reasoning[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2018:4942-4950.

[0015] [8]MAO J,GAN C,KOHLI P,et al.The neuro-symbolic concept learner:Interpreting scenes,words,and sentences from natural supervision[J].arXiv preprint arXiv:1904.12584,2019. Summary of the Invention

[0016] The present invention is made to solve the above problems, and provides a module routing network-based visual question answering method that can fuse two modalities of vision and text at multiple semantic levels and can reason about complex problems based on the module routing network model. The following technical solutions are specifically adopted:

[0017] The module routing network-based visual question answering method provided by the present invention is used to process natural language question texts and related input question photos according to the module routing network model and generate question answers. It is characterized in that the module routing network model has a text network, a routing network, and a visual network including L module layers, and each module layer includes multiple modules. The method includes the following steps: Step 1, input the natural language question text into the text network to extract question features; Step 2, activate the corresponding modules in the visual network based on the routing path generated at least based on the question features to form activated modules, and input the question photo into the visual network. The activated modules extract image features from the question photo to form corresponding final features; Step 3, input the final features into a predetermined answer generator to generate question answers.

[0018] The visual question answering method based on the module routing network provided by the present invention may further have the following technical features. The sub-steps included in step 2 are as follows: Step 2-11, input the question features into the routing network to generate routing paths corresponding to all module layers; Step 2-12, activate the corresponding modules in all module layers of the visual network as activated modules according to the routing paths; Step 2-13, input the question picture into the visual network and sequentially extract the final features through the activated modules in each module.

[0019] The visual question answering method based on the module routing network provided by the present invention may further have the following technical features. That is, the sub-steps included in step 2' are as follows: Step 2-21, input the question features into the routing network to generate a routing path corresponding to the first module layer, and take the first module layer as the current module layer; Step 2-22, activate the corresponding module in the current module layer as the activated module according to the routing path; Step 2-23, input the question picture into the visual network, and the activated module in the current module layer extracts image features from the question photo as the current image features; Step 2-24, input the image features and the question features into the routing network to generate a routing path corresponding to the next module layer, and take the next module layer as the new current module layer; Step 2-25, input the current image features into the current module layer, and the activated module in the current module layer extracts image features as the new current image features; Step 2-26, repeat steps 2-24 to 2-25 until the final features are obtained by the activated module of the last layer.

[0020] The visual question answering method based on the module routing network provided by the present invention may further have the following technical features. That is, the visual network includes a first attention module and a second attention module. The first attention module is used to paste the image features and the question features together, and establish the connection between two positions or objects in the image features through the spatial self-attention mechanism to obtain a fused feature map. The second attention module is used to weight-average the image features by the spatial attention mechanism to obtain a normalized weight map, and combine the fused feature map and the weight map to obtain the final features.

[0021] The visual question answering method based on the module routing network provided by the present invention may further have the following technical features. That is, when training the module routing network model and constructing the training loss function, the training loss function includes a visual question answering loss function and a load balancing loss function. The visual question answering loss function is used to adjust the parameters in the text network, the routing network, the visual network, and the answerer. The load balancing loss function is used to adjust the parameters in the text network and the routing network.

[0022] The visual question answering method based on the module routing network provided by the present invention may also have the following technical features. That is, the number of modules in each module layer is fixed at M, the routing path is a binary matrix P, the routing path follows a Bernoulli distribution of L×M dimensions, and the parameters of the Bernoulli distribution are generated by the routing network.

[0023] The visual question answering method based on the module routing network provided by the present invention may also have the following technical features. That is, the number of modules in each module layer is different.

[0024] The visual question answering method based on the module routing network provided by the present invention may also have the following technical features. That is, when the question picture reaches each module layer of the visual network, the features output after the activation module processes and the residual input are summarized as the image features output by each module layer.

[0025] The visual question answering method based on the module routing network provided by the present invention may also have the following technical features. That is, each module in each module layer is a neural network with different granularities, and at least one module in each module layer is consistent with the CNN model.

[0026] Function and effect of the invention

[0027] According to a visual question answering method based on the module routing network proposed by the present invention, since the question features are obtained by processing the natural language question text through the text network, and then the corresponding image features are obtained by the visual network processing the picture and the routing path, and then the features corresponding to the text network and the visual network in the routing network are processed through the routing path generated based on the question features, so that the final features are obtained and the final answer is extracted from the predetermined answers of the answerer through the predetermined training model. Therefore, the visual question answering method based on the module routing network provided by the present invention can fuse the text and visual modalities at multiple levels, enabling the text and picture features to be fully analyzed and processed, so as to answer complex questions without the need for expert knowledge and supervision information. Therefore, this visual question answering method based on the module routing network can be more widely applied to other situations that require the combination of multiple modalities. Brief description of the drawings

[0028] Figure 1 It is a framework diagram of the module routing network in the embodiment of the present invention;

[0029] Figure 2 It is a flowchart of the visual question answering method based on the module routing network model in the first embodiment of the present invention;

[0030] Figure 3It is a schematic diagram of the input form of the routing network in the first embodiment of the present invention;

[0031] Figure 4 It is a flowchart of the visual question answering method based on the modular routing network model in the second embodiment of the present invention; and

[0032] Figure 5 It is a schematic diagram of the input form of the routing network in the second embodiment of the present invention. Detailed implementation manners

[0033] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the following specifically describes a visual question answering method based on the modular routing network model of the present invention in conjunction with embodiments and drawings.

[0034] <Embodiment 1>

[0035] Figure 1 It is a framework diagram of the modular routing network in the embodiment of the present invention.

[0036] In this embodiment, the visual question answering method based on the modular routing network model is implemented by a pre-trained modular routing network model. As Figure 1 shown, the modular routing network model has a text network, a routing network, and a visual network. Among them, the text network is used to process natural language questions to obtain question features; the routing network is used to generate discrete routing paths; the visual network selectively activates certain modules through the routing paths generated by the routing network to extract image features, and the visual network consists of a series of general neural modules; the question features are generated by receiving the text network. In this embodiment, the visual network includes L module layers and each module layer includes M modules. The routing path is a binary matrix P and follows a Bernoulli distribution of L×M dimensions, and the parameters of the Bernoulli distribution are generated by the routing network.

[0037] Figure 2 It is a flowchart of the visual question answering method based on the modular routing network model in the first embodiment of the present invention.

[0038] As Figure 2 shown, the process of the visual question answering method based on the modular routing network model includes steps 1 to 3, specifically including the following steps 1 to 3.

[0039] Step 1, input the natural language question text into the text network to extract question features.

[0040] Among them, the natural language question text is a question described in natural language and stored in text form. After the natural language question text is input into the text network, the text network performs feature extraction to obtain the question features corresponding to the questions in the natural language question text.

[0041] In this embodiment, as Figure 1 shown, the text network employs a classical bidirectional GRU network, which takes the natural language question text as input and outputs question features. In this bidirectional GRU network, the dimension of the word vector is set to 200, the hidden feature dimension of the GRU is set to 256, the word embedding layer is initialized uniformly, and the GRU layer is initialized orthogonally. In addition, the structure of the text network is not limited to this, and a unidirectional GRU can also be used. The dimension of the hidden features can also be adjusted, and there is not much difference between 256 and 512.

[0042] Step 2: Activate the corresponding modules in the visual network based on the routing path generated by the routing network at least based on the question features to become activation modules, and input the question photo into the visual network. The activation modules extract image features from the question photo to form corresponding final features.

[0043] Among them, the routing network consists of fully connected layers, and for the routing network taking the question features as input, the part with trainable parameters is a fully connected layer.

[0044] In this embodiment, Step 2 includes the following sub-steps 2-11 to 2-13.

[0045] Step 2-11: Input the question features into the routing network to generate routing paths corresponding to all module layers.

[0046] In Step 2 of this embodiment, taking the routing network with only the question features as input as an example, the pseudo-code corresponding to the routing network when calculating the routing path is as follows:

[0047] / * Sample from the Logistics distribution * /

[0048] 1: First, sample a random tensor u of the same dimension as the routing path from a uniform distribution from 0 to 1, u ∼ Uniform([0,1] L×M ), where u ∈ R L×M

[0049] 2: Then calculate t ← log(u) - log(1 - u) through the random tensor u. At this time, t is equivalent to reparameterized sampling from the Logistics distribution

[0050] / * Calculate the routing path * /

[0051] 3: First, obtain the real-valued routing path through the fully connected layer

[0052] 4: Then use the Gumbel-Sigmoid trick to calculate the random tensor

[0053] 5: Obtain through thresholding Discrete version of

[0054] / / detach means only passing values ​​without accumulating gradients

[0055] 6: Finally, use the direct connection technique to calculate the routing path

[0056] Through the processing corresponding to the above pseudo code, the routing path can be obtained.

[0057] Routing path P∈{0,1} L×M Each value in controls the activation of any module in the visual network. It can be generally understood that each module in the visual network is bound to a switch, and the routing path summarizes the status of all switches (open or closed). If the switch is turned on, the module is activated and executed, and if the switch is turned off, the module is not executed. In this embodiment, 1 is used in the routing path to indicate that the switch is turned on, and 0 is used to indicate that the switch is turned off.

[0058] The routing path activation module processes the input image or image features and aggregates the output and residual input to the next module layer. Specifically, when the input reaches the lth layer, the corresponding module with all values ​​1 in the lth row of the routing path P processes the input.

[0059] The routing path has two properties: one is discreteness and the other is randomness. In the specific implementation, the Gumbel-Max method is first used to perform reparameterized sampling from the Bernoulli distribution. However, since the formula finally generated by this method contains a non-differentiable unit cross-order function, the method of using a sigmoid function with temperature to approximate this unit cross-order function is the Gumbel-Sigmoid method. When the temperature is greater than 0, when the value generated by the Gumbel-Sigmoid method is still not a discrete value, the direct connection method is used to threshold it so that the gradient can be back-propagated normally.

[0060] After the routing path activates the corresponding modules in the visual network to form multiple activation modules, the activation module will be used to process the input and aggregate the output and residual input to the next module layer. Specifically, when the input reaches the lth layer, the corresponding module with all values ​​1 in the lth row of the routing path P will process the input.

[0061] Step 2-12, activate the corresponding modules in all module layers of the visual network according to the routing path as activation modules.

[0062] Among them, the visual network consists of a series of modules. Each module is a small neural network that takes an image as input and extracts visual image features through routing path control. The modules have different granularities. By adjusting the granularity of the modules, the modules can be constructed into different visual networks in the form of filters in the convolutional layer or branches in ResNext. The modules form a module layer. Each module layer has at least one module consistent with the convolutional neural network (CNN) model and has residual connections. The L-layer modular layer forms a framework. When constructing the visual network, the visual network follows the modular routing framework.

[0063] Step 2-13: Input the problem picture into the visual network and sequentially extract the final features through the activation modules in each module.

[0064] Figure 3 It is a schematic diagram of the input form of the routing network in the first embodiment of the present invention.

[0065] The Figure 3 shows a solution when only the problem features are used as input and new image features are generated through the routing network according to the routing path generated based on the problem features. The input form of this solution is similar to the SENet method.

[0066] Step 3: Input the final features into a predetermined answerer to generate the question answer.

[0067] When the problem picture reaches each module layer of the visual network, the features and residuals output after the activation module processing are aggregated as the image features output by each module layer.

[0068] In this embodiment, after the corresponding modules in the visual network are activated by the routing path to form multiple activation modules, the activation modules are used to process the input, and the output and residual input are aggregated and sent to the next module layer, and finally input into a predetermined answerer to generate the question answer.

[0069] The answerer is a classifier, and the form of the answer is to select a correct answer from a plurality of preset answers according to the input final image features.

[0070] The above model is established in a predetermined training model. Before establishing the training model, it is necessary to train the module routing network model and construct a loss function for training.

[0071] When training the module routing network model and constructing the loss function for training, the loss function for training includes a visual question answering loss function and a load balancing loss function. The visual question answering loss function is used to adjust the parameters in the text network, routing network, visual network, and answerer. The load balancing loss is used to modify the parameters in the text network and the routing network. The load of a certain module is the number of samples that activate the module in a training batch. For the load of any module layer, the more balanced the loads of all the modules in the module layer, the smaller the load balancing loss function, where the load balancing loss function is implemented through the coefficient of variation.

[0072] Function and Effect of Embodiment 1

[0073] A visual question answering method based on a module routing network proposed by the present invention is to process the natural language question text through a text network to obtain question features, then the visual network processes the picture and the routing path to obtain corresponding image features, and then the routing network processes the features corresponding to the text network and the visual network through the routing path generated based on the question features, so as to obtain the final features and extract the final answer from the predetermined answers of the answerer through the predetermined training model. Therefore, the visual question answering method based on a module routing network provided by the present invention can fuse the text and visual modalities at multiple levels, enabling the text and picture features to be fully analyzed and processed, so as to answer complex questions without the need for expert knowledge and supervision information. This visual question answering method based on a module routing network can be more widely applied to other situations that require the combination of multiple modalities.

[0074] In addition, in the embodiment, the routing path obtained by activating the modules in all module layers of the visual network while activating the modules of the visual network through the routing path can simplify the working process of the visual network and improve the working efficiency while ensuring the accuracy of information.

[0075] In addition, in the embodiment, first train the training visual network model and construct the loss function for training, then when establishing the visual network, integrate the question text into the training visual network, construct the training model of the training visual network for the question text and the question image, and then the trained visual network can fully analyze the problem to be solved and obtain a more practical answer. This can not only avoid the loss of visual question answering, but also avoid model collapse additionally, and balance the load and make up for the loss of visual question answering by adjusting the load.

[0076] The adjustable granularity of the visual network can change the characteristics of the input routing network, enabling the routing network to receive and fuse question features and image features simultaneously to obtain an answer that better matches the question image.

[0077] In addition, the method of the present invention belongs to the implicit and general version of the neural module network method, and can obtain the final answer without expert knowledge and additional supervision information. Among them, the supervision information of the neural module network is that the dataset needs to provide sub-questions obtained by splitting the question and the order of answering these sub-questions in sequence in addition to images, questions, and answers.

[0078] Compared with classical computer vision tasks such as image recognition, detection, and segmentation, the visual question answering method based on the module routing network model provided in this embodiment can fuse the visual and text modalities more fully. Therefore, the module routing network model provided in this embodiment has a certain "understanding" of images and can answer questions better.

[0079] <Embodiment 2>

[0080] For the sake of easy expression, in the second embodiment, the same symbols are given to the same structures as those in the first embodiment, and the corresponding descriptions are omitted.

[0081] Figure 4 It is a flowchart of the visual question answering method based on the module routing network model in the second embodiment of the present invention.

[0082] As Figure 4 shown, in the second embodiment, steps 1 and 3 are the same as steps 1 and 3 in the first embodiment, and will not be described in detail here.

[0083] Compared with the first embodiment, in the second embodiment, not only the question features are input into the routing network, but also the question features and image features are merged to generate a routing path. Step 2' includes the following sub-steps from step 2-21 to step 2-26.

[0084] Step 2-21: Input the question features into the routing network to generate a routing path corresponding to the first module layer, and take the first module layer as the current module layer.

[0085] Step 2-22: Activate the corresponding modules in the current module layer according to the routing path as the activated modules.

[0086] Step 2-23: Input the question picture into the visual network, and the activated modules in the first module layer extract image features from the question photo as the current image features.

[0087] Step 2-24: Input the image features and question features into the routing network to generate a routing path corresponding to the next module layer, and take the next module layer as the new current module layer.

[0088] Step 2-25: Input the current image features into the current module layer, and the activated modules in the current module layer extract image features as the new current image features.

[0089] Step 2-26: Repeat steps 2-24 to 2-25 until the final feature is obtained from the activation module of the last module layer.

[0090] Figure 5 It is a schematic diagram of the input form of the routing network in the second embodiment of the present invention.

[0091] As Figure 5 shown, Figure 5 It shows the control of the routing network over the l-th layer of the visual network, and gives the routing path obtained by processing the image feature and the text feature when using both the image feature and the text feature as inputs at the same time. The next-layer routing path and the image feature are generated through the routing network. The above steps are repeated until the final feature is obtained after the last module is activated, and the final answer is obtained through the answerer in this way.

[0092] Function and effect of the second embodiment

[0093] In the embodiment, processing the image feature and the problem feature of each module layer through the routing network can more accurately combine the image feature and the problem feature to obtain a more practical final feature, and the answer obtained by inputting the final feature into the answerer can also solve the problem more effectively.

[0094] The routing network aggregates the features and residual inputs output after the activation module processes them as the image feature output by each module layer. The image feature of each layer can consider the residual input on the basis of the image feature of the previous module layer, and can aggregate the effective content that may be lost in the previous step with the processed image feature, so as to avoid obtaining an answer that does not conform to the problem photo and the problem text.

[0095] <Embodiment Three>

[0096] On the basis of the first embodiment, this third embodiment adds an attention module to the visual network, which is used to paste the image feature and the problem feature, and establish the connection between two positions or objects in the image feature through the spatial self-attention mechanism to obtain the final feature.

[0097] The attention module in the visual network includes a first attention module and a second attention module.

[0098] The first attention module is used to paste the image feature and the problem feature, and then use the spatial self-attention mechanism to model the connection between two positions or objects in the feature map to obtain a final feature with stronger representation ability. The spatial self-attention mechanism can have different implementation methods, such as the encoder in Transformer.

[0099] The second attention module is used to obtain a normalized weight map by weighted averaging the image features through a spatial attention mechanism and then passing it through softmax, and combines the fused feature map and the weight map to obtain the final features. The spatial attention mechanism can also have different implementation methods, such as reducing the number of channels of the final features to 1 through a 1×1 convolution.

[0100] The first attention module uses a Transformer encoder stacked with 3 layers; the second attention module reduces the number of channels of the fused feature map to 1 through a 1×1 convolution, obtains a normalized weight map through softmax, and then obtains the final features by combining the weight map and the fused feature map, calculates the loss function, and obtains the trained model.

[0101] Next, an effectiveness test is carried out on the models in Embodiment 1 and Embodiment 3.

[0102] The model in Embodiment 1 is built on ResNext, and when the filter of the module is a convolutional layer called FRN, the module performs routing control at the second module layer of the ResNext residual block.

[0103] The model in Embodiment 3 is built on ResNext, and when the branch in ResNext is called BRN, the module performs routing control on the second grouped convolution in the ResNext basic block and pays attention to adjusting the number of branches. The experimental results of BRN with 8, 16, and 32 branches show that the model performance increases in turn. Replace the BatchNorm in the third normalization layer of the ResNext basic block with GroupNorm.

[0104] When training with the ADAM optimizer, the hyperparameters betas of ADAM are set to (0.9, 0.999). After training with ADAM, the model performance is slightly improved by fine-tuning the epochs with SGD. When using the warmup strategy, the learning rate slowly increases to 3e-4 in the first few epochs. When conducting experiments on CLEVR and CLEVR-Humans respectively, the CLEVR dataset contains 100,000 images and 1 million questions, and CLEVR-Humans contains approximately 18,000 training questions, of which approximately 7,000 questions are used to verify the model.

[0105] The experiment on the trained model is carried out on the CLEVR and CLEVR-Humans datasets. The answer form is to select a correct answer from the pre-given answers, so the answerer is a classifier.

[0106] In this embodiment, in step 2, when calculating the routing path to specifically implement the algorithm pseudocode, the hyperparameter temperature is constantly set to 1.0 (of course, it can also be done in the way of simulated annealing), and the bias of the fully connected layer is initially set to Then the initialization makes the probability of each module executed in the initial stage of training about 0.7.

[0107] In the model without the attention module in Embodiment 1, when the branches of BRN are 8, 16, and 32, the overall accuracies on the CLEVR dataset are 86.8%, 94.7%, and 97.9% respectively, and the accuracy of BRN with a branch of 32 on the CLEVR-Humans dataset is 77.9%; the overall accuracy of FRN on the CLEVR dataset is 98.2%, and the accuracy on the CLEVR-Humans dataset is 79.9%.

[0108] In the model with the attention module added in this Embodiment 3, the overall accuracy of BRN with a branch number of 32 on the CLEVR dataset is 98.6% respectively, and the accuracy on the CLEVR-Humans dataset is 79.3%; the overall accuracy of FRN on the CLEVR dataset is 98.9% respectively, and the overall accuracy on the CLEVR-Humans dataset is 81.8%.

[0109] Combined with the above test results, it can also be known that removing the attention module and only using the image features processed by the visual network to answer questions can also obtain answers that match the images. However, compared with the visual network using the attention module, the training module using the attention module can better fuse the image features and question features, and can give more practical answers and higher accuracies when more complex problems need to be solved.

[0110] The above embodiments are only used to illustrate the specific implementation manners of the present invention, and the present invention is not limited to the description scope of the above embodiments.

[0111] In the above embodiments, the visual network has the same number of modules in all module layers. However, in actual implementation, the number of modules in each module layer is set to different numbers according to actual needs, and these different module parts are activated through the routing path, so as to construct the corresponding routing network according to the actual problem and make corresponding analyses according to the characteristics of the actual problem, which can avoid the unnecessary time cost caused by applying the actual complex problem to a fixed template.

[0112] In the above embodiments, the load balancing loss function is implemented through the coefficient of variation. However, in actual implementation, the expression method of the load balancing loss function is not limited to being implemented through the coefficient of variation. When actually constructing the training function, it can be implemented in other ways, such as simply making the standard deviation smaller.

[0113] In the above embodiments, only the implementation manner in the visual question answering task is given. However, according to the present invention, the model can be easily applied to multi-modal fusion tasks, especially the fusion of two-modal information, and can be more widely applied.

[0114] In the above embodiments, the answer of the answerer is in the form of selecting a correct answer from the pre-set answers. However, in actual implementation, the answerer can directly generate corresponding answers according to the input question text and question image, without the need to additionally set the answers to be selected, saving the time for analyzing and understanding the question text and question image.

Claims

1. A visual question answering method based on a modular routing network model, which is used to process natural language question texts and related input question photos according to the modular routing network model and generate question answers, characterized in that, The module routing network model has a text network, a routing network, and a visual network including L module layers, each module layer including multiple modules, and the method includes the following steps: Step 1, input the natural language question text into the text network to extract question features; Step 2, activate the corresponding modules in the visual network as activated modules according to the routing paths generated by the routing network at least based on the question features, and input the question photo into the visual network, and the activated modules extract image features from the question photo to form corresponding final features; Step 3, input the final features into a predetermined answerer to generate the question answer, wherein, Step 2 includes the following sub-steps: Step 2-21, input the question features into the routing network to generate the routing paths corresponding to the first module layer, and take the first module layer as the current module layer; Step 2-22, activate the corresponding modules in the current module layer as activated modules according to the routing paths; Step 2-23, input the question photo into the visual network, and the activated modules in the current module layer extract the image features from the question photo as the current image features; Step 2-24, input the image features and the question features into the routing network to generate the routing paths corresponding to the next module layer, and take the next module layer as the new current module layer; Step 2-25, input the current image features into the current module layer, and the activated modules in the current module layer extract the image features as the new current image features; Step 2-26, repeat Steps 2-24 to 2-25 until the final features are obtained by the activated modules of the last layer.

2. The visual question answering method based on the module routing network model according to claim 1, It is characterized in that: wherein, Step 2 includes the following sub-steps: Step 2-11, input the question features into the routing network to generate the routing paths corresponding to all the module layers; Step 2-12, activate the corresponding modules in all the module layers of the visual network as activated modules according to the routing paths; Step 2-13, input the question photo into the visual network, and the final features are extracted by passing through the activated modules in each module in turn.

3. The visual question answering method based on the module routing network model according to claim 1, characterized in that: Among them, The visual network includes the modules as the first attention module and the second attention module, The first attention module is used to paste the image features and the question features together, and establish the connection between two positions or objects in the image features through the spatial self-attention mechanism to obtain a fused feature map, The second attention module is used to weighted average the image features by the spatial attention mechanism to obtain a normalized weight map, and combine the fused feature map and the weight map to obtain the final features.

4. The visual question answering method based on the module routing network model according to claim 1, characterized in that: Among them, When training the module routing network model and constructing the loss function for training, the loss function for training includes a visual question answering loss function and a load balancing loss function. The visual question answering loss function is used to adjust the parameters in the text network, the routing network, the visual network, and the answerer. The load balancing loss function is used to adjust the parameters in the text network and the routing network.

5. The visual question answering method based on the module routing network model according to claim 1, wherein: Among them, The number of the modules in each module layer is fixed at M. The routing path is a binary matrix P. The routing path follows a Bernoulli distribution of L×M dimensions, and the parameters of the Bernoulli distribution are generated by the routing network.

6. The visual question answering method based on the module routing network model according to claim 1, wherein: Among them, The number of the modules in each module layer is different.

7. The visual question answering method based on the module routing network model according to claim 1, wherein: Among them, When the question photo reaches each module layer of the visual network, the features and residual inputs output after being processed by the activation module are summarized as the image features output by each module layer.

8. The visual question answering method based on the module routing network model according to claim 1, wherein: Each of the modules in each module layer is a neural network with different granularities. Among them, each module layer has at least one module whose neural network is consistent with the CNN model.

Citation Information

Patent Citations

  • Automatic Question-Answering System

    KR1020180088192A