Visual question and answer method based on multi-modal bidirectional attention guidance

An attention and multi-modal technology, applied in neural learning methods, character and pattern recognition, biological neural network models, etc., can solve problems such as loss of local information, disadvantages, etc.

Pending Publication Date: 2021-12-24
SICHUAN UNIV
View PDF4 Cites 2 Cited by
  • Summary
  • Abstract
  • Description
  • Claims
  • Application Information

AI Technical Summary

Problems solved by technology

Although this method is simple and direct, it loses important local information, which is not conducive to answering questions about local areas.

Method used

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
View more

Image

Smart Image Click on the blue labels to locate them in the text.
Viewing Examples
Smart Image
  • Visual question and answer method based on multi-modal bidirectional attention guidance
  • Visual question and answer method based on multi-modal bidirectional attention guidance
  • Visual question and answer method based on multi-modal bidirectional attention guidance

Examples

Experimental program
Comparison scheme
Effect test

Embodiment Construction

[0034] The present invention will be further described below in conjunction with accompanying drawing:

[0035] figure 1 It is the principle diagram of the image-guided problem attention module proposed by the present invention. This module is composed of 4 layers of guided attention units connected by stacking. It is mainly guided by image features and pays more attention to words containing effective information in the question. The module input is the weighted image attention feature output through the 6-layer SGA structure and the question self-attention feature through the 6-layer self-attention unit.

[0036] In order to verify the rationality of the image guidance problem proposed by the present invention and pay attention to the value of the module cascading layer number as 4, different values ​​have been experimentally verified, and the results are as shown in Table 1:

[0037] Table I

[0038]

[0039] It can be seen from Table 1 that when the number of GA unit...

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More

PUM

No PUM Login to View More

Abstract

The invention designs a visual question and answer method based on multi-modal bidirectional attention guidance, and relates to the two fields of computer vision and natural language processing. The full understanding that interactivity between different modes of vision and text and autocorrelation between the same modes are keys for overcoming the difficulty of a visual question and answer task. Effective information in images and problems is highlighted by reasonably utilizing an attention mechanism so as to be beneficial to improving the performance of the model, an attention guiding module for guiding attention of the problems by the images is designed based on an attention guiding mechanism, bidirectional attention guiding is constructed in combination with collaborative attention, the accuracy of answering questions or not and the overall accuracy are improved to a certain extent, and the counting capability of the model is improved by combining a Counter module. The method has certain significance in practical application aspects of helping visually impaired people and children to learn books and the like.

Description

technical field [0001] The invention relates to two fields of computer vision and natural language processing, in particular to obtaining weighted attention features of different modalities by using a self-attention mechanism and a guided attention mechanism, and in particular to increasing the guidance of images to problems based on collaborative attention. Background technique [0002] The visual question answering task aims to give an image and a question related to the image, and answer the correct answer to the question. The task involves the learning of both visual and textual modalities, bridging the gap between computer vision and natural language processing. The early visual question answering model mainly extracts the global features of images and questions, and then generates a predicted answer after simple feature fusion and classification. Although this method is simple and direct, it loses important local information, which is not conducive to answering questi...

Claims

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More

Application Information

Patent Timeline
no application Login to View More
IPC IPC(8): G06K9/62G06N3/04G06N3/08
CPCG06N3/049G06N3/08G06N3/045G06F18/254G06F18/253
Inventor何小海鲜荣吴晓红卿粼波吴小强滕奇志任超
OwnerSICHUAN UNIV