Co-Reference Object Recognition Using ROI-Based Visual Dialogue
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic apparatuses struggle to accurately understand and respond to co-references in user queries, as they fail to clearly identify the objects referred to by co-references in images, leading to inappropriate responses.
Innovation Solution
An electronic apparatus equipped with a microphone, camera, and processor that identifies regions of interest based on co-reference attributes such as distance and number of objects, and provides information on the identified objects by executing instructions stored in its memory, using modules for voice recognition, natural language understanding, and object recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a related-art AI model is used to understand user utterance, then the apparatus can process basic queries, but it cannot clearly understand co-references and provide appropriate responses
Solution Approach 1:
The system segments the co-reference understanding process into distinct modules: co-reference identification module to detect co-reference expressions, region of interest identification module to locate relevant image areas, and object identification module to recognize specific objects. This segmentation allows each module to specialize in one aspect, improving overall accuracy without overwhelming system complexity
Solution Approach 2:
The patent introduces an intermediary region of interest (ROI) identification step between co-reference detection and object recognition. The ROI acts as a mediator that narrows down the search space in the image based on the co-reference context, making the final object identification more accurate while keeping the overall process manageable
2Measurement precision
If the apparatus processes only text input, then the system remains simple, but it cannot identify visual objects referred to by co-references
Solution Approach 1:
The patent merges multiple processing streams: text processing (co-reference identification) and image processing (region of interest identification and object recognition) into a unified system. The co-reference from text is linked with the visual content through the ROI, enabling accurate object identification by combining linguistic and visual information
Solution Approach 2:
The system transitions from one-dimensional text processing to two-dimensional image analysis by introducing region of interest identification. When a co-reference is detected in text, the system maps it to a specific spatial region in the image, adding a visual dimension to the understanding process and enabling precise object identification
Data Source
AI summary
Disclosed is an electronic apparatus providing a reply to a query of a user. The electronic apparatus includes a microphone, a camera, a memory configured to store at least one instruction, and at least one processor, and the processor is configured to execute the at least one instruction to control the electronic apparatus to: identify a region of interest corresponding to a co-reference in an image acquired through the camera based on a co-reference being included in the query, identify an object referred to by the co-reference among at least one object included in the identified region of interest based on a dialogue content that includes the query, and provide information on the identified object as the reply.


