Grounded visual question-answering method based on dynamic dual-level visual information fusion

By adopting a dynamic two-level visual information fusion method in the visual question and answer system, combined with the language-guided area-level and pixel-level feature modules, the simultaneous output of text answers and grounded answer masks is realized, solving the problem of insufficient answer accuracy and interpretability in traditional systems, and improving the accuracy and credibility of the question and answer system.

WO2025091218A1PCT designated stage expired Publication Date: 2025-05-08DALIAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/128345
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-31
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Traditional visual question and answer systems usually only output the final text answer, which lacks verification of visual evidence, resulting in insufficient accuracy and interpretability of the answers, making it difficult to meet the need to provide credible answers in practical applications.

Method used

The grounded visual question-and-answer method based on dynamic two-level visual information fusion is adopted. Through the design language-guided area-level feature module and pixel-level feature module, combined with the cross-modal multi-scale fusion module, the simultaneous output of text answers and grounded answer masks is realized, and model training is carried out through multiple loss functions to improve the accuracy and interpretability of answers.

Benefits of technology

It realizes a grounded visual question-and-answer system that can achieve better segmentation effects without relying on large-scale pre-trained models, improves the understanding and positioning ability of complex scenarios, and enhances the accuracy and credibility of the question-and-answer system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023128345_08052025_PF_FP_ABST
    Figure CN2023128345_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present invention is a grounded visual question-answering method based on dynamic dual-level visual information fusion. A dual-level multi-scale network is used for constructing a grounded visual question-answering system, that is, performing division into a language-guided pixel-level feature and a language-guided region-level feature, the two scale branches being combined so as to perform final text answer and grounded answer prediction. In addition, a question-guided dynamic region-level feature positioning network is provided, so that the visual information is positioned by means of question guidance and masks of different dimensions are adaptively allocated to grounded answers, thereby improving the accuracy of positioning and segmentation of small targets. Furthermore, a cross-modal aggregation module is further designed to fuse the features of the two levels, so that feature fusion of the pixel-level and region-level features can be enhanced, thereby improving the segmentation effect for the edge of grounded answer masks. In the present invention, by using the grounded visual question-answering system established on the basis of a language-guided adaptative dual-level feature fusion network, answer grounding masks can be generated while questions are answered, thereby effectively improving the accuracy of the whole model.
Need to check novelty before this filing date? Find Prior Art

Description

Grounded visual question answering method based on dynamic two-level visual information fusion Technical Field

[0001] The present invention belongs to the technical field of computer vision and natural language processing, and specifically relates to a grounded visual question answering method based on dynamic two-level visual information fusion. Background Art

[0002] In recent years, VQA (Visual Question Answering) technology has developed rapidly, with a growing number of practical applications, such as answering questions from visually impaired patients, helping radiologists diagnose life-threatening diseases early, and addressing human-computer interaction. As these systems mature, simply generating good answers is no longer sufficient; it is also crucial for the answers to be well-reasoned for various research and applications. By considering the model's reasoning mechanism, it is possible to provide explainable support for the answers to a certain extent. An ideal VQA system for such purposes should not only generate accurate answers but also provide a mechanism for verifying them.

[0003] However, traditional VQA typically only outputs the final textual answer, lacking visual verification. Recent work has attempted to address this issue. For example, the MAC-CAPS method (Capsule-based Weakly Supervised Grounded Visual Question Answering) proposes providing a visual attention map alongside the textual answer to better evaluate the accuracy of the system's answer localization. Similar methods include LXMERT (Transformer-based Cross-Modal Encoder) and DCAMN (Dual Capsule Attention Mask Network with Mutual Learning for Visual Question Answering), which also simultaneously generate textual answers and output the corresponding grounded answer region in the image. However, these methods typically output attention maps or boxes related to the question to illustrate the grounded answer region. Providing a well-founded image grounded answer mask when answering visual questions allows direct verification of the convincingness of the obtained answer, making VQA systems more reliable. Furthermore, obtaining an image grounded answer mask can expand applications. For example, when answering questions from visually impaired users, relevant content can be segmented from the background, blurring the background to protect privacy, or magnifying the relevant visual area, allowing users with low vision to find the information they are looking for more quickly.

[0004] Therefore, the answer grounding task was proposed. Unlike conventional VQA tasks, it starts from the actual application of visually impaired people and aims to output a mask map of the visual area corresponding to the answer while answering the text answer. For this task, DAVI (Answer Grounding Based on Dual Vision-Language Interaction) combines two large pre-trained models, BLIP (Guided Language Image Pre-training for Unified Visual Language Understanding and Generation) and VIT (Multimodal Framework Based on Vision and Language Research). It contains two encoders and two decoders, combining the text image segmentation task model and the vision-to-language generation task model. However, it is actually equivalent to dividing the two interrelated tasks of generating text answers and outputting grounding masks into two independent tasks. The newly published DDTN (Grounded Visual Answer Based on Dual Decoder Transformer Network) does not use a large-scale pre-training model, but the segmentation effect is much worse than DAVT.

[0005] Summary of the Invention

[0006] In response to the above-mentioned problems existing in the prior art, the present invention proposes a grounded visual question answering method based on dynamic two-level visual information fusion, which can achieve relatively good segmentation effect without relying on a large-scale pre-trained model, and realize the output of two answer modalities under the conditions of one encoder and decoder, which can better realize the interaction between the two modalities.

[0007] To achieve the above object, the technical solution of the present invention is:

[0008] The grounded visual question answering method based on dynamic two-level visual information fusion includes the following steps:

[0009] Step 1: As shown in Figure 1, the present invention adopts a question-guided regional-level dynamic multi-scale method to locate and segment the ground answer, and designs a language-guided regional-level feature module QGDR, which consists of a cross-attention module and a spatial attention module, and finally obtains the regional-level mask prediction feature F with small to large resolution. i ∈F t ,F s ,F m ,F l ; where F t ,F s ,F m ,F l There are four types of regional feature hierarchical structures, from F t to F l The spatial resolution increases by a factor of two layer by layer;

[0010] Step 2: To reduce computational overhead while maintaining performance, a dynamic method is used to adaptively assign a mask of appropriate resolution to each localized object, while also imposing a budget constraint on resource consumption. The QGDR output has four different on-off states, corresponding to four different mask resolutions: [14×14, 28×28, 56×56, 112×112].

[0011] Step 3: In order to better fuse the features of the two levels, a cross-modal multi-scale fusion module FPA is designed to combine the features F output by the language-guided pixel-level feature module PWAM and the language-guided region-level feature module QGDR. i and P i Perform polymerization;

[0012] Step 4: Construct an information flow between each level of the language-guided pixel-level feature module (PWAM) and the language-guided region-level feature module (QGDR), perform hierarchical decoding, and finally obtain the grounded answer by the image segmentation decoder and the text answer by the text decoder. The grounded visual question answering model composed of two-level feature branches is trained using mask loss, edge loss, budget constraint, and text loss.

[0013] Step 5: Load the model in step 4, input the required images and their corresponding questions into the trained grounded visual question answering model, and obtain the corresponding grounded answers and text answers.

[0014] Based on the above scheme, this method utilizes multi-scale information fusion to better understand and process visual information at different scales, helping to improve comprehension and localization of complex scenes, thereby enhancing question answering accuracy. Adaptive resolution mask allocation dynamically allocates a mask of appropriate resolution size based on the needs of each localized object, improving resource utilization while maintaining high resolution in critical areas. A cross-modal multi-scale fusion module is introduced to aggregate language-guided pixel-level and region-level features at multiple scales, effectively integrating textual and image information, improving question understanding and answer generation. A hierarchical decoding approach is employed to decode information from pixel-level features to region-level features and finally to the answer, helping to better capture image details and relate them to the question, thereby improving question answering accuracy. Multiple loss functions, including mask loss, edge loss, budget constraint, and text loss, are used to comprehensively consider different objectives, thereby optimizing model training and improving performance. This method can be applied to grounded visual question answering, providing an efficient and accurate method for machine understanding of images and answering questions, with potential applications in various fields such as autonomous driving, medical imaging analysis, and image retrieval.

[0015] Furthermore, step 1 specifically includes:

[0016] Step 1.1: First extract the ROI aligned regional feature Z from swin-transformer i Perform average pooling to obtain Combined with the question feature K extracted from BERT i .Will and K i The input is then fed into the cross-modal attention. This step can be considered as injecting the attention of the words in the question into different visual channels to guide visual localization and promote the complementary enhancement of multimodal information. Where T represents the transposition operation, after two linear transformations, the specific formula is as follows:

[0017] where Q i represents the attention weight; d i express and the length of the vector; Represents the vector generated by linear transformation of problem features;

[0018] Step 1.2: Get Q i Perform global pooling operation to obtain information weight It is fed into the attention module SE-block to weight different channels of visual information for screening. Then it is classified using several convolutional and fully connected layers to obtain region-level mask prediction features F of different sizes. i The specific formula is as follows:

[0019] Where, Represents Flattehen operation, F ex represents the operation in the SE-block module, w represents the weight; F ex The specific operation formula is as follows:

[0020] Where δ represents the sigmoid function, ρ represents the ReLU function, and Represents the weight matrix dimension.

[0021] Furthermore, the step 2 specifically includes:

[0022] The QGDR module is actually a lightweight classifier that aims to select the best mask resolution from the k candidate targets of different scales to accurately locate and segment the ground truth at the minimum resource cost. iDivided into four categories of regional feature hierarchy F t ,F s ,F m ,F l , from F t to F l The spatial resolution increases by two times layer by layer. And the probability vector ε is output by performing a softmax operation. k =[ε 1 ,…,ε k ]. Each element of the probability vector represents the probability that the corresponding candidate resolution is selected. The soft output ε of QGDR k should be converted into a one-hot prediction, denoted as H = [h1,…,h k This process can be completed through discrete sampling, and then Gumbel-Softmax is used to back propagate the gradient to update QGDR. The specific formula is as follows:

[0023] Where τ is a parameter; when τ is close to 0, Gumbel-softmax is close to one-hot. i represents Gumbel distribution; ε k′ represents k' discrete probability vectors.

[0024] Furthermore, the step 3 specifically includes:

[0025] Step 3.1: The two modal information of image and question are processed by the language-guided pixel-level feature module (PWAM) and the language-guided region-level feature module (QGDR) to obtain the transmembrane fusion feature. and F i ∈R C×H×W Next, we need to perform multi-scale aggregation on the outputs of these two modules. Due to the upsampling and ROI pooling operations of these two modules, F i and P i There is spatial misalignment between them. In order to enhance the segmentation performance of the boundary area, this paper designs a cross-modal multi-scale fusion module FPA that adaptively aggregates multi-scale features. As shown in Figure 1, FPA contains a deformable convolution and a dynamic convolution. First, F i After deconvolution (Deconv) upsampling, then F i With P i Connect them in series and pass the concatenated features through a 3×3 conv to obtain the offset mapping, represented by ΔO. Finally, use the learned offset o to convert F i Align P i , adjust the output F of QGDR through deformable convolution deform conv1 i position so that it better matches the output of PWAMi Alignment, the specific formula is as follows: i =Φ[conv(ρ(F i )||P i )] (5)

[0026] Where ρ represents the Deconv operation, Φ represents the Deform conv1 operation, and || is the connection operation.

[0027] Step 3.2: After the variable convolution operation O i With P i The sum is then passed through a 1×1 convolution to achieve an output channel of C. Finally, the conditional convolution CondConv is used, which is similar to the attention mechanism and focuses more on the prominent parts of the object. The cross-modal multi-scale fusion module FPA is inserted into different stages of the swin-transformer decoding and plays a key role in improving the ground answer mask prediction. The specific formula is as follows: i =ψ(conv 1×1 (O i +P i )) (6)

[0028] Where Y i represents the regional features; ψ represents the CondConv operation.

[0029] Furthermore, the step 4 is specifically as follows:

[0030] QGDR is guided by language to dynamically locate the ground answer in the image and provide ground answer masks of different resolutions for different aggregation stages. While ensuring accuracy, it reduces the computational resource cost and thus adopts three loss functions to train the dynamic multi-scale module.

[0031] Step 4.1: First, the mask loss (mask loss), given a VQA instance, first use QGDR to predict its mask switching state H = [h1,…,h k ] and is passed to different stages of the decoding end through the fusion of the FPA module to obtain a set of K mask prediction maps The mask loss function is defined as follows:

[0032] Where N represents N different instances, represents the ground answer mask of the k-th prediction, Indicates the corresponding true ground answer mask, h i Indicates whether to select the kth mask resolution as the output resolution. Expressed as binary cross entropy loss.

[0033] Step 4.2: The second is edge loss. The mask generated by QGDR is dynamically selected. It is usually believed that the size of the mask loss is used as a measure of mask quality. However, in fact, the mask loss generated on different masks is very close, and it is difficult to distinguish the mask quality. In contrast, the edge loss generated by masks of different resolutions is quite different, which can better reflect the quality of the mask. Therefore, the present invention uses edge loss to measure the mask quality. Given the output of QGDR F = [f1,···,f k ] and edge maps of different resolutions, using The marginal loss is defined as follows:

[0034] Among them, represents the ground truth answer edge, which is obtained by first The soft edge map is obtained by applying the Laplacian operator on it, and then converted into a binary edge map by thresholding.

[0035] Step 4.3: The QGDR module is optimized by the edge loss in step 4.2, but there is a problem that the model training tends to converge to a suboptimal solution, that is, all instances are segmented with the maximum resolution mask, because the mask contains more detailed information, so the prediction loss is minimized. In fact, experiments have shown that not all samples require the largest mask for segmentation. In order to avoid the above problems, improve model efficiency and reduce the amount of computation, the present invention adopts budget constraint training QGDR. Specifically, let C represent the corresponding computational cost of the selected mask resolution. Indicates that the expected deviation (E(C)) calculated for the current batch data exceeds the target deviation (measured in C t denoted), add a penalty to the model.

[0036] Step 4.4: The overall objective function of the ground answer branch is as follows, where λ1 and λ2 are trade-off hyperparameters:

[0037] Finally, the question features and visual features are combined through element-wise product and classified through the Softmax function. The network is trained using the text answers and the binary cross entropy loss function of PWAM.

[0038] Furthermore, the step 5 specifically includes:

[0039] Load the model best trained in step 4, input the image and its corresponding question into the model, and output the answer and corresponding evaluation indicators.

[0040] Beneficial effects of the present invention: The present invention proposes a grounded visual question-answering method based on dynamic two-level visual information fusion, which constructs a multi-level direct flow from pixel-level features to region-level features, thereby promoting the complementary information aggregation of multi-level features. Specifically, the present invention proposes a problem-oriented dynamic region-level module, which can effectively locate region-level objects according to the question, and dynamically select masks of different resolutions, thereby realizing language-guided object-level features for multi-scale feature fusion. In addition, the present invention proposes a cross-modal multi-scale fusion module, which is guided by the language in the image, and adaptively aggregates pixel-level information and region-level content, thereby realizing the interaction and fusion of multimodal information from different levels, achieving high-quality information interaction, and effectively improving the accuracy of the entire model. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a framework diagram of a grounded visual question answering network based on dynamic two-level visual information fusion. DETAILED DESCRIPTION

[0042] The embodiments of the present invention are implemented on the premise of the technical solution of the present invention, and detailed implementation methods and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0043] The present invention provides a grounded visual question-answering method based on dynamic two-level visual information fusion. A grounded visual question-answering system is constructed through a two-level multi-scale network, which is divided into language-guided pixel-level features and region-level features. The two scale branches are combined to predict the final text answer and the grounded answer. A question-guided dynamic region-level feature localization network is proposed. Through question-guided visual information localization, masks of different sizes are adaptively assigned to the grounded answer, thereby improving the accuracy of positioning and segmentation of small targets. In addition, a cross-modal aggregation module is designed to fuse the two levels of features, which can enhance the feature fusion between pixel-level and region-level features to improve the segmentation effect of the edge of the grounded answer mask. The grounded visual question-answering system constructed by the language-guided adaptive two-level feature fusion network can generate the answer ground mask while answering the question, effectively improving the accuracy of the entire model.

[0044] Example 1

[0045] This embodiment uses the Windows system as the development environment, Pycharm as the development platform, and Python as the development language. It adopts the grounded visual question-answering method based on dynamic two-level visual information fusion of the present invention to complete the grounded answer prediction for pictures taken by visually impaired people and related questions.

[0046] In this embodiment, the visual question answering method based on dynamic two-level visual information fusion includes the following steps:

[0047] Step 1: Load the pre-trained weights of the Swin-Transformer and BERT encoder in the DDVT network into the grounded visual question answering network as shown in Figure 1;

[0048] Step 2: Input the ‘image-question-ground answer’ pairs in the training set into the grounded visual question answering network in step 1 for training;

[0049] Step 3: Take the required image and the corresponding question as input, load the network model trained and saved in step 2, and obtain the corresponding grounded answer and the corresponding evaluation index. The present invention uses the intersection-over-union ratio, that is, the overlapping area between the model predicted segmentation and the label divided by the joint area between the predicted segmentation and the label, as the evaluation index. Its calculation method can be expressed by formula (16), where S i and S u Represent the predicted segmentation answers and the true label answers, respectively.

[0050] According to the above steps, the present invention compares the LXMTRT model, Mac-Caps model, UNIFIED model, DDVT model, and MCAN model. As can be seen from Table 1, the accuracy of the proposed method on two common test sets is generally better than that of other methods.

[0051] Table 1. Comparison of model performance on the closed part of the VizWizGroundVQA validation set and the VQS test set.

[0052] The foregoing descriptions of specific exemplary embodiments of the present invention are for purposes of illustration and description. These descriptions are not intended to limit the invention to the precise forms disclosed, and it is apparent that many modifications and variations are possible in light of the foregoing teachings. The exemplary embodiments have been selected and described for the purpose of explaining the specific principles of the invention and their practical application, thereby enabling those skilled in the art to make and utilize a variety of exemplary embodiments of the invention and various options and variations. The scope of the invention is intended to be defined by the claims and their equivalents.

Claims

1. A grounded visual question answering method based on dynamic two-level visual information fusion, characterized in that: The method comprises the following steps: Step 1: Use the question-guided region-level dynamic multi-scale method to locate and segment the ground answer, design a language-guided region-level feature module QGDR, which consists of a cross-attention module and a spatial attention module, and finally obtain the region-level mask prediction feature F with small to large resolutions. i ∈F t ,F s ,F m ,F l , where F t ,F s ,F m ,F l There are four types of regional feature hierarchical structures, from F t to F l The spatial resolution increases by two times layer by layer; Step 2: A dynamic method is used to adaptively assign a mask of appropriate resolution size to each positioning object, and a budget limit is imposed on resource consumption; QGDR outputs four different switch states, corresponding to four different mask resolutions, namely [14×14, 28×28, 56×56, 112×112]; Step 3: Design a cross-modal multi-scale fusion module FPA to perform multi-scale aggregation on the features output by the language-guided pixel-level feature module PWAM and the language-guided region-level feature module QGDR; Step 4: construct an information flow between each level of the language-guided pixel-level feature module PWAM and the language-guided region-level feature module QGDR, perform hierarchical decoding, and finally obtain the grounded answer by the image segmentation decoder and the text answer by the text decoder; use mask loss, edge loss, budget constraint and text loss to jointly train the grounded visual question answering model composed of two-level feature branches; Step 5: Load the model in step 4, input the required image and its corresponding question into the trained grounded visual question answering model, and obtain the corresponding grounded answer and text answer.

2. The grounded visual question answering method based on dynamic two-level visual information fusion according to claim 1 is characterized in that: The problem in step 1 guides the regional level dynamic multi-scale method, which specifically includes: Step 1.1: First extract the ROI-aligned regional feature Z from the swin-transformer i Perform average pooling to obtain Combined with the question feature K extracted from BERT i ,Will and K i Input to cross-modal In attention, T represents the transposition operation, which undergoes two linear transformations. The specific formula is as follows: Where Q i represents the attention weight; d i express and The length of the vector; Represents the vector generated by linear transformation of the problem features; Step 1.2: Get Q i Perform global pooling operation to obtain information weight is fed into the attention module SE-block to weight different channels of visual information for screening; it is then classified using several convolutional and fully connected layers to obtain region-level mask prediction features F of different sizes. i , the specific formula is as follows: In the formula, represents the Flattehen operation, F ex represents the operation in the SE-block module, and w represents the weight; F ex The specific operation formula is as follows: Where δ represents the sigmoid function, ρ represents the ReLU function, and Represents the weight matrix dimension.

3. The grounded visual question answering method based on dynamic two-level visual information fusion according to claim 1 is characterized in that: Step 2 adopts a dynamic method to adaptively assign a mask of appropriate resolution size to each positioning object, specifically including: QGDR is a lightweight classifier that selects the best mask resolution from k candidate targets of different scales. i Divided into four categories of regional feature hierarchy F t ,F s ,F m ,F l , from F t to F l The spatial resolution doubles layer by layer, and the probability vector ε is output by performing a softmax operation k =[ε 1 ,…,ε k ]; Each element of the probability vector represents the probability that the corresponding candidate resolution is selected; The soft output ε of QGDR k Transformed into a one-hot prediction, denoted as H = [h1,…,h k ], this process is completed by discrete sampling, Then Gumbel-Softmax is used to back propagate the gradient to update QGDR. The specific formula is as follows: Where τ is a parameter; when τ is close to 0, Gumbel-softmax is close to one-hot; g i represents Gumbel distribution; ε k′ represents k' discrete probability vectors.

4. The grounded visual question answering method based on dynamic two-level visual information fusion according to claim 1 is characterized in that: The step 3 specifically includes: Step 3.1: The two modal information of image and question are processed by the language-guided pixel-level feature module PWAM and the language-guided region-level feature module QGDR to obtain the transmembrane fusion feature and F i ∈R C×H×W Next, the outputs of these two modules are multi-scale aggregated; a cross-modal multi-scale fusion module FPA is designed to adaptively aggregate multi-scale features. FPA contains a deformable convolution and a dynamic convolution; first, F i After deconvolution Deconv is performed for upsampling, F i With P i The concatenated features are passed through a 3×3 conv to obtain the offset mapping, denoted by ΔO; finally, F is mapped using the learned offset o. i Align P i , adjust the output F of QGDR through deformable convolution deform conv1 i position so that it matches the output P of PWAM i Alignment, the specific formula is as follows: The i =Φ[conv(ρ(F i )||P i )] (5) Where ρ represents the Deconv operation, Φ represents the deform conv1 operation, and || is the connection operation; Step 3.2: After the variable convolution operation O i With P i Add, and then pass 1×1 convolution to realize the output channel C; finally, through the conditional convolution CondConv, the cross-modal multi-scale fusion module FPA is inserted into the different stages of swin-transformer decoding. The specific formula is as follows: AND i =ψ(conv 1×1 (EITHER i +P i )) (6) Where Y i represents regional features; ψ represents the CondConv operation.

5. The grounded visual question answering method based on dynamic two-level visual information fusion according to claim 4, characterized in that: The mask loss in step 4 is as follows: given a VQA instance, first use QGDR to predict its mask switching states of different resolutions H = [h1,…,h k ], and is passed to different stages of the decoding end through the fusion of the FPA module to obtain a set of K mask prediction maps The mask loss function is defined as follows: Where N represents N different instances, represents the k-th predicted ground answer mask, represents the corresponding true ground answer mask, h i Indicates whether to select the kth mask resolution as the output resolution. Expressed as binary cross entropy loss.

6. The grounded visual question answering method based on dynamic two-level visual information fusion according to claim 5 is characterized in that: The edge loss in step 4 is as follows: the edge loss is used to measure the mask quality. Given the output of QGDR F = [f1, ···, f k ] and edge maps of different resolutions, using Indicates that the marginal loss is defined as follows: in, represents the ground truth answer edge, which is obtained by first filtering the ground truth answer mask The soft edge map is obtained by applying the Laplacian operator on it, and then converted into a binary edge map by thresholding.

7. The grounded visual question answering method based on dynamic two-level visual information fusion according to claim 6 is characterized in that: The budget constraint and text loss described in step 4 are as follows: QGDR is trained with a budget constraint. Specifically, let C denote the corresponding computational cost of the selected mask resolution, and let the expected deviation E(C) calculated for the current batch of data exceed the target deviation C t When , add a penalty to the model: The overall objective function for the grounded answer branch is as follows: where λ1 and λ2 are trade-off hyperparameters: Finally, the question features and visual features are combined through element-wise product, classified through the Softmax function, and trained using the text answer and the binary cross entropy loss function of PWAM.

Citation Information

Patent Citations

  • Visual question and answer oriented method of context awareness based on multi-modal interaction

    CN114970517A

  • Visual question and answer method based on multistage mesh interaction model

    CN115546596A

  • Answer positioning method and device based on weak supervision double-flow visual language interaction

    CN116010578A

  • Medical visual question answering

    US20220130499A1

  • Cross-modal processing for vision and language

    WO2022187063A1