A method and device for answer localization based on weakly supervised dual-stream visual language interaction

Through the visual language dual-stream interaction method, multimodal feature alignment and weak supervision learning are solved, and the problem of multimodal interaction and data labeling difficulties in visual question-and-answer positioning is realized, efficient visual question-and-answer positioning and answer positioning is improved, and the accuracy of visual question-and-answer is reduced.

CN116010578BActive Publication Date: 2025-08-26TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310067972.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-06
Publication Date
2025-08-26
Estimated Expiration
2043-02-06

AI Technical Summary

Technical Problem

The existing visual question-answer positioning methods have shortcomings in visual-language multimodal feature interaction and spatial relationship representation, and the lack of data labeling leads to difficulty in training, high data labeling costs, and difficult to effectively train under different weak supervision annotations.

Method used

Multimodal feature alignment is adopted for vision-based language encoder and language-based vision decoder, weak-supervised learning is introduced to generate pseudo-labels, trained through a unified weak-supervised network framework, use visual question-and-answer to locate the existing resources of the community, and generate intensive pseudo-labels to reduce the annotation cost.

Benefits of technology

It realizes efficient training under different weak supervision annotations, improves the accuracy and complexity of visual Q&A positioning, simplifies the training process, and is suitable for helping people with disabilities identify image information and protect privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116010578B_ABST
    Figure CN116010578B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for answer localization based on weakly supervised dual-stream visual-language interaction. The method comprises: linearly mapping visual cascade features and natural language features using a language-based visual decoder, aligning text features with the visual cascade features through multimodal fusion, and generating a final answer localization map; finally, fusing the visual features of interest with question features using an answer decoder to generate a joint embedding, and predicting the correct answer from a set of answers using a classifier; introducing weakly supervised learning into dual-stream visual question-answering localization, allowing the model to learn self-generated pseudo-labels to supplement missing real data; and using the correct answer and answer localization map to identify information in images for people with disabilities. The device includes a processor and a memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimodal technology, and in particular to a method and device for answer localization based on weakly supervised dual-stream visual language interaction. Background Art

[0002] Visual Question Answering (VQA) aims to answer questions about images and provide natural language answers, such as answering questions about images for the visually impaired. In order to improve the effectiveness of visual question answering, recent studies have begun to evaluate the intersection of the answer area on the image, because the image area used to answer the question is also important for visual question answering. Existing work usually obtains the answer location through the attention map of the visual question answering module and evaluates whether the model correctly focuses on the regional objects related to the answer. This task is usually called Visual Question Answering grounding, which spans vision-language multimodal multi-task. With this technology, people with disabilities can use mobile phones to identify information in images and highlight key information while protecting privacy.

[0003] Visual question answering localization, as an extension of visual question answering based on visual evidence, incorporates spatial information in addition to multimodal information. Visual question answering involves the model providing an answer given an image and a corresponding natural language question. However, visual question answering localization requires not only accurate answers but also the corresponding image evidence region. This requires the model to locate the relevant image region and answer the visual question. Compared to visual question answering, visual question answering localization faces numerous problems and challenges. For example, how should visual-linguistic multimodal features be represented? How should visual-linguistic multimodality interact? How should features be fed into specific downstream tasks? How should multiple tasks be balanced?

[0004] To achieve good localization accuracy in visual question answering, most methods in this field rely on input feature maps from object detection models, pre-processed with preset relevant object classes. Mac-Caps first proposed a visual capsule module with a query-based capsule feature selection mechanism. This enables the model to focus on relevant areas based on the textual clues of the visual information in the question and predict the background truth bounding box of the object associated with the correct answer. In addition, Khan AU utilizes a capsule network, grouping each visual marker in the visual encoder. Using the activation of the language self-attention layer as a text guide, a selection module is used to select capsules and mask them before forwarding them to the next layer.

[0005] However, to date, most methods use capsule networks to achieve weakly supervised answer localization based on the attention or gradient maps of visual question answering methods. Although such methods are effective, they completely ignore spatial relationships. At the same time, the problem of lack of data labels is also serious. Only a few real-world datasets provide localization labels, which makes this task challenging. Summary of the Invention

[0006] This paper provides a method and device for answer localization based on weakly supervised dual-stream visual-language interaction. This invention is a new end-to-end unified framework that combines visual question answering and answer localization capabilities. It can help people with disabilities identify information in images while emphasizing key points and protecting privacy. Details are described below:

[0007] A method for answer localization based on weakly supervised two-stream visual-language interaction, the method comprising:

[0008] The visual features and natural language features are linearly mapped separately through a vision-based language encoder, and the mapped visual features and natural language features are multimodally fused to align the visual features with the text features.

[0009] The language-based visual decoder linearly maps the visual cascade features and natural language features, and aligns the text features with the visual cascade features through multimodal fusion to generate the final answer location map.

[0010] The answer decoder finally fuses the visual features of interest with the question features to generate a joint embedding, and the classifier predicts the correct answer from the answer set;

[0011] By introducing weakly supervised learning in two-stream language visual question answering positioning, the model learns pseudo-labels generated by itself to supplement the lack of real data;

[0012] Based on the correct answers and the answer location map, it is used for people with disabilities to identify information in images.

[0013] A device for answer localization based on weakly supervised dual-stream visual language interaction, the device comprising: a processor and a memory, wherein program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute method steps.

[0014] The beneficial effects of the technical solution provided by the present invention are:

[0015] 1. This paper introduces a vision-based language encoder and a language-based visual decoder to enable the model to achieve multimodal interaction with different focuses, so as to facilitate different downstream tasks. This paper also proposes a unified weakly supervised network that can fully utilize the existing resources of the visual question answering and positioning community, train under different weakly supervised annotations, and achieve good performance, thus solving the current problems of inconsistent visual question answering and positioning data and high data annotation costs.

[0016] 2. This paper proposes a unified weakly supervised visual question answering localization model. Assuming that there is a natural constraint between answer localization and visual question answering, and that pixels in the same semantic region have highly similar cues, the model learns sparse weak labels related to answer localization and visual question answering, generates dense pseudo-labels for training, and reduces the annotation cost of the dataset. It can also take into account the existing answer grounding annotation (e.g., box-level, pixel-level, and question-answer-level) in the VQA community.

[0017] 3. The interaction algorithm of the present invention is simple, but significantly outperforms existing fully supervised and weakly supervised methods in terms of complexity and accuracy. This dual-stream visual-linguistic interaction method can be extended to other label-efficient tasks in the future without the need for complex training cycles and post-processing techniques.

[0018] 4. The present invention has a wide range of practical application scenarios, for example: helping people with disabilities identify information in images, answer natural language questions, and emphasize important areas in images while protecting other private information; this technology helps people with disabilities read pictures, recognize street scenes and objects, etc., improving the convenience of life and bringing substantial improvements to the work and life of visually impaired people. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is the architecture diagram of the visual language dual-stream interaction algorithm;

[0020] Figure 2 This is a framework diagram for introducing the weakly supervised image question answering and positioning algorithm.

[0021] Among them, VE is the visual encoder, VLE is the vision-based language encoder, LVD is the language-based visual decoder, and LM is the language model (answer decoder). DETAILED DESCRIPTION

[0022] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention are described in further detail below.

[0023] Example 1

[0024] Image question answering (Q&A) localization faces numerous problems and challenges. For example, how should visual-linguistic multimodal features be represented, how should visual-linguistic multimodal interactions be implemented, how should features be fed into specific downstream tasks, and how should multiple tasks be balanced? This embodiment of the present invention proposes a method for weakly supervised visual question answering localization based on dual-stream visual-linguistic interaction. This method utilizes dual-stream multimodal interaction and online pseudo-label supervision to achieve efficient image question answering and answer localization.

[0025] First, in order to overcome the difficulty of multi-modal interaction under multi-task, the embodiment of the present invention develops a method based on dual-stream interaction of visual language, see Figure 1 and Figure 2 It uses a two-stream cross-attention mechanism for feature alignment. Specifically, visual features are aligned and fused with text features through a cross-attention mechanism of a vision-based language encoder, and finally image question answering is performed. Text features are aligned with visual features through a language-based visual encoder, and finally answer localization is performed, achieving performance far superior to traditional methods.

[0026] At the same time, the embodiment of the present invention designs a unified weak supervision framework, which realizes the extension of weak supervision to image question answering positioning, and can also simultaneously utilize different labels of the image question answering positioning community through a single model.

[0027] Example 2

[0028] An ideal visual question answering and localization system must effectively model the complex interactions between language and vision to acquire useful knowledge and use it during reasoning to answer new questions and localize answers. Different tasks require different features. For example, localization tasks may prioritize visual features and spatial information, while image question answering tasks may prioritize language features, supplemented by visual features.

[0029] To address this issue, the present invention proposes a framework. Given an input image-question pair consisting of an image I and a question Q, the present invention aims to locate the relevant question-answer object under VQA-level supervision, point-level supervision, scribble-level supervision, or box-level supervision. The mathematical definition of weak supervision is shown in Table 1.

[0030] Table 1 Mathematical definitions of different weak supervision types

[0031]

[0032] The overall process is based on four main steps:

[0033] (1) Visual Encoder: Here, a Transformer-based visual encoder (ViT) is used to extract cascade features of the image.

[0034] (2) Vision-based language encoder: processes image-question pairs and linearly maps visual features to K and V and natural language features to Q. Finally, multimodal fusion is performed to align these visual features with text features to facilitate subsequent answer generation.

[0035] (3) Language-based visual decoder: Use the language-based visual decoder to decode the visual cascade feature F v The linear mapping is Q, the language features are linearly mapped to K and V, and finally the text features are aligned with the visual features through multimodal fusion to generate the answer location map.

[0036] (4) Answer decoder: Using a language generation model L similar to the BERT structure d ,The focused visual features are finally fused with the question features to generate a refined joint embedding, and then a classifier is used to predict the correct answer from the answer set.

[0037] 1. Visual Encoder

[0038] The visual encoder first uses convolutional kernels to divide the image into fixed-size patches, linearly embeds each patch, adds learnable positional embedding information, and inputs the resulting vector sequence into a standard Transformer encoder. To obtain global information, the encoder also adds an additional learnable "classification token" to the sequence using standard methods. Following the original Transformer architecture, its simplicity and ease of use are one of its advantages. It leverages the versatile and efficient Transformer architecture, making the visual encoder easy to apply.

[0039] 2. Vision-based Language Encoder and Answer Decoder

[0040] Vision-based Language Encoder (VL) e (Visual-based Linguistic Encoder). The visual cascade feature F extracted by the visual encoder v Through a separate linear mapping layer, it is mapped to the visual feature K vl 、V vl The language features are first extracted by self-attention and then processed by linear mapping Q vl Finally, the cross attention (CA) layer is interwoven between the self-attention layer and the feed-forward network of BERT (Bidirectional Encoder Representation of Transformer). The CA layer uses the language features as the visual reference of the query, and after 12 layers, it generates the feature F vl . V Le These visual language interactions in I d The answer decoding in provides multimodal evidence. The embodiment of the present invention provides a maximization objective to train the network, which is defined as:

[0041]

[0042] Among them, the data set is C, L Answer (C) is the answer decoder; L d is a vision-based language encoder; x is a token (embedded sequence), and y is the answer label.

[0043] 3. Language-based visual decoder

[0044] In order to predict the answer grounding mask, the language-based visual decoder LV d (Linguistic-based Visual Decoder) focuses visual features on evidence-related areas for answer localization. In this embodiment, visual features are mapped to image features Q through a linear mapping layer. lv , and the language feature F vl is mapped to K lv and V vl

[0045] , and then use the cross attention mechanism to align the visual features. Here, the embodiment of the present invention uses the fused cascade features to generate the answer location, and the network structure is the standard U-Net. The U-Net network minimizes the partial cross entropy loss L CE To train, we calculate the loss for n labeled pixels i, which is defined as:

[0046]

[0047] in, is the pixel category predicted by the model, is the true pixel category, and n is the total number of pixels.

[0048] 4. Refining

[0049] The embodiment of the present invention inputs the training image into the network to obtain a prediction map. Pseudo labels are generated by selecting the largest category of the prediction map. In order to ensure the quality of the pseudo labels, the present invention uses a refinement part to refine the pseudo labels and ensure local consistency. The refinement part is based on the pixel-adaptive convolution (PAC) technology, which assumes that adjacent regions have the same appearance and should be identified as the same category. The idea is to use its neighbors to identify the same category. The convex combination of labels iteratively updates the pixel label m:,i,j at t th In iterations:

[0050]

[0051] Among them, m represents the mask, :,i,j, : represents all values ​​of the dimension, and (i,j) represents the coordinate value.

[0052] To calculate α, embodiments of the present invention use a kernel function on the pixel-level intensity I:

[0053]

[0054] In this embodiment of the present invention, i,j Defined as the standard deviation of image intensity, it is calculated locally by the convolution kernel. The embodiment of the present invention applies softmax to obtain the final affinity distance α of each neighbor (l,n) of (i,j) i,j,l,n .

[0055] For this module, the present embodiment does not perform backpropagation, so it is always in "evaluation" mode:

[0056]

[0057] Among them, L Pseudo is the cross entropy loss of pseudo labels.

[0058] To summarize, in each iteration of the model, the network is trained using the overall loss L, which is defined as:

[0059] L=L Answer +L CE +L Pseudo #(6)

[0060] 5. Application of the Model

[0061] This model can effectively solve the problem of visual question answering positioning. First, 1) the model inputs the image into the visual encoder to extract the visual features F v , then in the vision-based language encoder, the natural language question related to the image is extracted and mapped to Q, visual features F v are mapped to K and V respectively, and aligned and fused through the cross-attention mechanism. 2) The vision-based language encoder can understand the questions corresponding to the visual features and generate language-oriented evidence, which is used for the answer decoder to predict the correct answer to the natural language question. 3) On the other hand, the language-based visual encoder converts the visual cascade feature F v The linear mapping is Q, and the language features are mapped to K and V. Finally, they are aligned and fused through the cross-attention mechanism to generate the answer location map.

[0062] Example 3

[0063] The VizWiz-VQA-Grounding dataset used in this embodiment of the present invention is derived from the VizWiz VQA dataset. This dataset, which utilizes a dataset captured by blind individuals, provides visual localization labels for questions and answers for evaluating visual question answering localization methods. This dataset consists of 32,842 image-question pairs, each with 10 crowdsourced answers. In addition to publicly available VQA triplets from the training and validation splits, there are also triplets from the test split. These images and questions are shared by visually impaired individuals seeking visual assistance in their daily lives. We then removed all unanswerable questions, questions with multiple sub-questions embedded, questions that could not be grounded due to ambiguity and covered multiple regions in the image, questions for which a majority of people disagreed, or questions for which a single answer was not agreed upon. This process left a total of 9,998 VQAs. This embodiment of the present invention used 6,494 triplets as the training set and 1,131 triplets as the validation set.

[0064] Based on extensive experiments, the AdamW optimizer was used with a weight decay of 0.05. The learning rate was preheated to 2e-4 and linearly decayed at a rate of 0.85. During pretraining, the present embodiment captured random image scaling at a resolution of 224 x 224. The model of the present embodiment was implemented in PyTorch and trained on an 8-GPU node. The image encoder was initialized from ViT, pretrained on ImageNet, and the text converter was initialized from BERT. The model was pretrained for 20 epochs.

[0065] The network of this embodiment of the present invention is trained using the aforementioned training set. The network input consists of a question and an image. After extracting features from the input image using a visual encoder, it is encoded using a vision-based language encoder, and the language decoder generates the answer. To predict the evidence mask, a visual decoder using language as a reference is used to generate visual localization. Furthermore, a teacher-student network is used to generate pseudo labels, enabling weakly supervised learning. The detailed network structure is shown in Table 1:

[0066] Table 1 Network structure

[0067]

[0068]

[0069] After obtaining the trained model, it was tested on the test set. The experimental results are shown in Table 2:

[0070] Table 2 Comparative experiment

[0071]

[0072] The weakly supervised experiments are shown in Table 3 below.

[0073] Table 3 Experiments with different weak supervision information

[0074]

[0075]

[0076] As can be seen from Table 3, the method proposed in the present invention has a significant improvement in the aforementioned indicators compared to previous methods, indicating that the method proposed in the present invention is universal and can achieve a strong effect.

[0077] Example 4

[0078] A device for answer localization based on weakly supervised dual-stream visual-language interaction, comprising: a processor and a memory, wherein program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to cause the device to perform the following method steps:

[0079] The visual features and natural language features are linearly mapped separately through a vision-based language encoder, and the mapped visual features and natural language features are multimodally fused to align the visual features with the text features.

[0080] The language-based visual decoder linearly maps the visual cascade features and natural language features, and aligns the text features with the visual cascade features through multimodal fusion to generate the final answer location map.

[0081] The answer decoder finally fuses the visual features of interest with the question features to generate a joint embedding, and the classifier predicts the correct answer from the answer set;

[0082] By introducing weakly supervised learning in two-stream language visual question answering positioning, the model learns pseudo-labels generated by itself to supplement the lack of real data;

[0083] Based on the correct answers and the answer location map, it is used for people with disabilities to identify information in images.

[0084] Among them, the vision-based language encoder is:

[0085] The visual features extracted by the visual encoder are converted into visual features K through two independent linear mapping layers vl and V vl , the natural language embedding extracts internal relevant information through the self-attention layer and is mapped to Q through the linear mapping layer vl ;

[0086] Q vl , K vl and V vlInput to the cross attention layer, the cross attention layer uses the language features as a reference for the visual query, and generates the feature F after passing through the feedforward network layer vl ;

[0087] The network is trained to maximize the objective, defined as:

[0088]

[0089] Among them, the data set is C, L Answer (C) is the answer decoder, x is the embedding sequence, and y is the answer label.

[0090] Among them, the language-based visual decoder is:

[0091] Linearly map the cascade features extracted by the visual encoder to Q lv , linearly mapping the language features extracted by the vision-based language encoder to K lv and V lv ,Finally, the cross attention layer is used to fuse multimodal features and generate cascade features;

[0092] The fused cascade features are used to generate answer positioning through the U-Net network, which minimizes the partial cross entropy loss L CE To train, calculate the loss of labeled n pixels i, defined as:

[0093]

[0094] in, is the pixel category predicted by the model, is the true pixel category, and n is the total number of pixels.

[0095] Among them, by introducing weak supervision learning in two-stream language visual question answering positioning, the model learns the pseudo labels generated by itself:

[0096] Use Neighbors The label convex combination iteratively updates the pixel label m:,i,j and refines the pseudo label. th In iterations:

[0097]

[0098] Among them, m represents the mask, :, i, j, : represents all values ​​of the dimension, and (i, j) represents the coordinate value;

[0099] Apply a kernel function on the pixel-level intensity I:

[0100]

[0101] Among them, σi,j Defined as the standard deviation of image intensity, softmax is applied to obtain the final affinity distance α for each neighbor (l,n) of (i,j) i,j,l,n ;

[0102]

[0103] Among them, L Pseudo is the cross entropy loss of pseudo labels;

[0104] The overall loss L is trained and is defined as:

[0105] L=L Answer +L CE +L Pseudo .

[0106] It should be noted here that the device description in the above embodiment corresponds to the method description in the embodiment, and the embodiment of the present invention will not be described in detail here.

[0107] The execution subjects of the above-mentioned processor and memory can be computers, single-chip microcomputers, microcontrollers and other devices with computing functions. In specific implementation, the embodiment of the present invention does not limit the execution subject and it can be selected according to the needs of actual application.

[0108] Data signals are transmitted between the memory and the processor via a bus, which will not be described in detail in the embodiment of the present invention.

[0109] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable storage medium, which includes a stored program, and when the program is running, controls the device where the storage medium is located to execute the method steps in the above embodiment.

[0110] The computer-readable storage medium includes but is not limited to a flash memory, a hard disk, a solid-state drive, and the like.

[0111] It should be noted here that the description of the readable storage medium in the above embodiment corresponds to the description of the method in the embodiment, and the embodiment of the present invention will not be described in detail here.

[0112] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part.

[0113] The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in or transmitted via computer-readable storage media. Computer-readable storage media can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that integrates one or more available media. Available media can include magnetic media or semiconductor media, etc.

[0114] Those skilled in the art will understand that the accompanying drawings are only a schematic diagram of a preferred embodiment, and the serial numbers of the embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0115] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for answer localization based on weakly supervised two-stream visual language interaction, characterized by: The method comprises: The visual features and natural language features are linearly mapped separately through a vision-based language encoder, and the mapped visual features and natural language features are multimodally fused to align the visual features with the text features. The language-based visual decoder linearly maps the visual cascade features and natural language features, and aligns the text features with the visual cascade features through multimodal fusion to generate the final answer location map. The answer decoder finally fuses the visual features of interest with the question features to generate a joint embedding, and the classifier predicts the correct answer from the answer set; By introducing weakly supervised learning in two-stream language visual question answering positioning, the model learns pseudo-labels generated by itself to supplement the lack of real data; Based on the correct answers and the answer location map, it is used for people with disabilities to identify information in images.

2. The answer localization method based on weakly supervised dual-stream visual-language interaction according to claim 1 is characterized in that: The vision-based language encoder is: The visual features extracted by the visual encoder are converted into visual features K through two independent linear mapping layers vl and V vl , the natural language embedding extracts internal relevant information through the self-attention layer and is mapped to Q through the linear mapping layer vl ; Q vl , K vl and V vl Input to the cross attention layer, the cross attention layer uses the language features as a reference for the visual query, and generates the feature F after passing through the feedforward network layer vl ; The network is trained to maximize the objective, defined as: Among them, the data set is C, L Answer (C) is the answer decoder, x is the embedding sequence, and y is the answer label.

3. The answer localization method based on weakly supervised dual-stream visual-language interaction according to claim 1 is characterized in that: The language-based visual decoder is: Linearly map the cascade features extracted by the visual encoder to Q lv , linearly mapping the language features extracted by the vision-based language encoder to K lv and V lv ,Finally, the cross attention layer is used to fuse multimodal features and generate cascade features; The fused cascade features are used to generate answer positioning through the U-Net network, which minimizes the partial cross entropy loss L CE To train, calculate the loss of labeled n pixels i, defined as: in, is the pixel category predicted by the model, is the true pixel category, and n is the total number of pixels.

4. The answer localization method based on weakly supervised dual-stream visual-language interaction according to claim 1 is characterized in that: By introducing weakly supervised learning in two-stream language visual question answering positioning, the model can learn the pseudo labels generated by itself: Use Neighbors The label convex combination iteratively updates the pixel label m:,i,j and refines the pseudo label. th In iterations: Among them, m represents the mask, :, i, j, : represents all values ​​of the dimension, and (i, j) represents the coordinate value; Apply a kernel function on the pixel-level intensity I: Among them, σ i,j Defined as the standard deviation of image intensity, softmax is applied to obtain the final affinity distance α for each neighbor (l,n) of (i,j) i,j,l,n ; Among them, L Pseudo is the cross entropy loss of pseudo labels; The overall loss L is trained and is defined as: L=L Answer +L CE +L Pseudo 。 5. An answer localization device based on weakly supervised dual-stream visual language interaction, characterized in that: The device includes: a processor and a memory, wherein program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the method steps according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Multi-modal information fusion method under visual question and answer task based on knowledge

    CN113240046A

  • Multi-clue reasoning with memory augmentation for knowledge-based visual question answering

    WO2022165858A1