Visual question-answering method and system of visual question-answering model based on labor division decision

A question-answering system and visual technology, applied in semantic analysis, character and pattern recognition, instruments, etc., can solve problems such as model-enhanced reasoning, joint coding of semantic information, difficulty, etc., and achieve the effect of reducing the difficulty of reasoning

Pending Publication Date: 2022-04-05
CHONGQING UNIV OF POSTS & TELECOMM
View PDF0 Cites 0 Cited by
  • Summary
  • Abstract
  • Description
  • Claims
  • Application Information

AI Technical Summary

Problems solved by technology

[0004] The above existing technologies only focus on the cross-modal fusion of images and texts. Although they involve cross-modal conversion, they do not jointly encode the high-level semantic information of the image and the semantic information of the question text, which makes the model suffer from cross-modal fusion. which increases the difficulty of reasoning

Method used

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
View more

Image

Smart Image Click on the blue labels to locate them in the text.
Viewing Examples
Smart Image
  • Visual question-answering method and system of visual question-answering model based on labor division decision
  • Visual question-answering method and system of visual question-answering model based on labor division decision
  • Visual question-answering method and system of visual question-answering model based on labor division decision

Examples

Experimental program
Comparison scheme
Effect test

Embodiment Construction

[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some, not all, embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by persons of ordinary skill in the art without making creative efforts belong to the protection scope of the present invention.

[0040] A visual question answering method based on a division of labor decision-making visual question answering model, the method comprising: acquiring a visual image and a question to be answered, inputting the visual image and the question to be answered into an image visual question answering model based on a division of labor decision-making, and obtaining a question answering result ; The image visual question answering model based on the division of labo...

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More

PUM

No PUM Login to View More

Abstract

The invention belongs to the field of image question answering, and particularly relates to a visual question answering method and system of a visual question answering model based on division of labor decision, and the method comprises the steps: obtaining a visual image and a to-be-answered question, inputting the visual image and the to-be-answered question into an LRBNet model, and obtaining a question answering result; the LRBNet model comprises a visual understanding module, a text understanding module and an exchange module; the visual understanding module is used for obtaining a visual feature map, the text understanding module is used for obtaining a text feature map, and the exchange module is used for carrying out data interaction on the visual feature map and the text feature map and updating nodes according to interaction data; associating and updating the visual spatial feature map and the text semantic information to obtain a final question and answer result; according to the method, the text semantic information and the visual space information are separately processed, and only the processed results are finally fused, so that the reasoning difficulty of other VQA models due to cross-modal fusion is reduced.

Description

technical field [0001] The invention belongs to the field of image question answering, and in particular relates to a visual question answering method and system based on a division of labor decision-making visual question answering model. Background technique [0002] Deep learning algorithms have achieved great success in both vision-related and language-related tasks, and visual question answering is a task that tests cross-modal understanding of vision and language. In the most common form of visual question answering, a computer is presented with an image and a textual question about the image, and a visual question answering algorithm needs to answer the question based on the image. [0003] Most current image question answering models (VQA) use neural networks such as recurrent neural networks (RNNS) and long short-term memory networks (LSTM) to learn encoded representations of questions. In order to encode the image, the early VQA model used Resnet or VGG to extract...

Claims

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More

Application Information

Patent Timeline
no application Login to View More
Patent Type & AuthorityApplications(China)
IPC IPC(8): G06V10/44G06V10/74G06V10/80G06V10/774G06K9/62G06V30/148G06F40/30
Inventor丰江帆刘睿国龙仁华易成杰
OwnerCHONGQING UNIV OF POSTS & TELECOMM