Multimodal Coreference Resolution for Visual Assistant Dialog
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems face challenges in resolving coreferences to entities across multiple input modalities, particularly in ambiguous user queries involving visual data, and in performing multimodal dialog state tracking and action prediction, which affect the accuracy and relevance of user interactions.
Innovation Solution
The assistant system integrates visual and dialog states to resolve coreferences, uses a CV module and scene understanding engine for multimodal dialog state tracking, and employs context information to determine user intent and context, enabling proactive recommendations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the assistant system processes multimodal inputs and performs coreference resolution across multiple modalities, then the accuracy and relevance of user interactions are improved, but the device complexity increases due to integration of CV module, scene understanding engine, and dialog state tracking mechanisms
Solution Approach 1:
The patent combines the CV module, scene understanding engine, and dialog state tracking into a unified assistant system that processes visual and dialog states together. The coreference resolution mechanism integrates multiple modalities (visual and textual) into a single processing pipeline, resolving references across different input types through coordinated analysis of both visual and dialog states.
Solution Approach 2:
The assistant system is designed to handle multiple input modalities universally - both visual inputs (images, video) and dialog inputs (text, speech) are processed through the same coreference resolution framework. The scene understanding engine serves multiple functions including entity recognition, relationship extraction, and context tracking across different modalities.
2Productivity
If the assistant system integrates visual data and dialog state tracking to resolve coreferences, then the user interaction relevance is improved, but the processing time increases due to multimodal analysis requirements
Solution Approach 1:
The system performs preliminary processing of visual data by extracting and storing visual states, entities, and relationships before they are needed for coreference resolution. The scene understanding engine pre-processes visual inputs to create a structured representation that can be quickly queried during dialog interactions, reducing real-time processing requirements.
Solution Approach 2:
The patent introduces an intermediary mechanism that bridges visual processing and dialog processing by creating a unified representation space. The coreference resolution system acts as an intermediary that queries both visual and dialog states to resolve references, efficiently matching between modalities without requiring full re-processing of all inputs.
Data Source
AI summary
In one embodiment, a method includes receiving, at a client system, an audio input, where the audio input comprises a coreference to a target object, accessing visual data from one or more camera associated with the client system, where the visual data comprises images portraying one or more objects, resolving the coreference to the target object from among the one or more objects, resoling the target object to a specific entity, and providing, at the client system, a response to the audio input, where the response comprises information about the specific entity.


