Attention Memory for Visual Dialogue Object Location

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing neural network-based object identification methods in visual dialogue systems generate inefficient questions as they rely solely on the questioner's information and answerer's responses, lacking incorporation of critical object location information for accurate identification.

Innovation Solution

An attention memory method and system that generates questions by utilizing the location information of objects stored in attention memory, updating this information through dialogue to improve question efficiency and accuracy in identifying objects on an image.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If the questioner generates questions based only on current question and answer information, then the dialogue process is simple, but the question efficiency and object identification accuracy deteriorate

Engineering Contradiction:
Improvequestion efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system performs preliminary action by pre-extracting and storing location information of objects in the image before the dialogue begins. This preliminary processing of visual data allows the question generator to access rich spatial information during the dialogue, improving question efficiency without requiring complex real-time processing during the conversation itself.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary mechanism (attention memory module) that bridges the gap between visual input and question generation. This intermediary stores and manages object location information, allowing the question generator to efficiently query relevant spatial data without directly processing the entire image in real-time, thus improving efficiency while managing complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the questioner does not incorporate object location information, then the question generation process is simple, but the object identification accuracy deteriorates

Engineering Contradiction:
Improveobject identification accuracyVSAvoidinformation processing complexity
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system extracts and stores location information of objects in advance before the dialogue begins. This preliminary extraction of spatial data ensures that accurate object identification can be achieved during the dialogue without requiring complex real-time information processing, as the necessary location data is already prepared and accessible.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts critical location information from the image and separates it from the main dialogue processing flow. By taking out and storing this spatial information in attention memory, the system can access it efficiently during question generation without burdening the main dialogue process with complex information processing, thus maintaining accuracy while reducing processing complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11188774B2Attentive memory method and system for locating object through visual dialogue
Publication Date: 2021.11.30 SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
  • US11188774B2 patent drawing
  • US11188774B2 patent drawing

AI summary

There are proposed an attention memory method and system for locating an object through visual dialogue. The attention memory system for identifying an object on an image includes: a control unit which generates a question for identifying a preset object on the image, derives an answer to the generated question, and identifies the preset object on the image based on the question and the answer; and memory which stores the image. The control unit generates the question by incorporating information about objects included in the image into the question, and updates the information about the objects based on the answer.