Attention Memory for Visual Dialogue Object Location
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing neural network-based object identification methods in visual dialogue systems generate inefficient questions as they rely solely on the questioner's information and answerer's responses, lacking incorporation of critical object location information for accurate identification.
Innovation Solution
An attention memory method and system that generates questions by utilizing the location information of objects stored in attention memory, updating this information through dialogue to improve question efficiency and accuracy in identifying objects on an image.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If the questioner generates questions based only on current question and answer information, then the dialogue process is simple, but the question efficiency and object identification accuracy deteriorate
Solution Approach 1:
The system performs preliminary action by pre-extracting and storing location information of objects in the image before the dialogue begins. This preliminary processing of visual data allows the question generator to access rich spatial information during the dialogue, improving question efficiency without requiring complex real-time processing during the conversation itself.
Solution Approach 2:
The patent introduces an intermediary mechanism (attention memory module) that bridges the gap between visual input and question generation. This intermediary stores and manages object location information, allowing the question generator to efficiently query relevant spatial data without directly processing the entire image in real-time, thus improving efficiency while managing complexity.
2Measurement precision
If the questioner does not incorporate object location information, then the question generation process is simple, but the object identification accuracy deteriorates
Solution Approach 1:
The system extracts and stores location information of objects in advance before the dialogue begins. This preliminary extraction of spatial data ensures that accurate object identification can be achieved during the dialogue without requiring complex real-time information processing, as the necessary location data is already prepared and accessible.
Solution Approach 2:
The patent extracts critical location information from the image and separates it from the main dialogue processing flow. By taking out and storing this spatial information in attention memory, the system can access it efficiently during question generation without burdening the main dialogue process with complex information processing, thus maintaining accuracy while reducing processing complexity.
Data Source
AI summary
There are proposed an attention memory method and system for locating an object through visual dialogue. The attention memory system for identifying an object on an image includes: a control unit which generates a question for identifying a preset object on the image, derives an answer to the generated question, and identifies the preset object on the image based on the question and the answer; and memory which stores the image. The control unit generates the question by incorporating information about objects included in the image into the question, and updates the information about the objects based on the answer.

