Multimodal Coreference Resolution for Visual Assistant Dialog

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems face challenges in resolving coreferences to entities across multiple input modalities, particularly in ambiguous user queries involving visual data, and in performing multimodal dialog state tracking and action prediction, which affect the accuracy and relevance of user interactions.

Innovation Solution

The assistant system integrates visual and dialog states to resolve coreferences, uses a CV module and scene understanding engine for multimodal dialog state tracking, and employs context information to determine user intent and context, enabling proactive recommendations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the assistant system processes multimodal inputs and performs coreference resolution across multiple modalities, then the accuracy and relevance of user interactions are improved, but the device complexity increases due to integration of CV module, scene understanding engine, and dialog state tracking mechanisms

Engineering Contradiction:
Improvecoreference resolution accuracyVSAvoidsystem architecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines the CV module, scene understanding engine, and dialog state tracking into a unified assistant system that processes visual and dialog states together. The coreference resolution mechanism integrates multiple modalities (visual and textual) into a single processing pipeline, resolving references across different input types through coordinated analysis of both visual and dialog states.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The assistant system is designed to handle multiple input modalities universally - both visual inputs (images, video) and dialog inputs (text, speech) are processed through the same coreference resolution framework. The scene understanding engine serves multiple functions including entity recognition, relationship extraction, and context tracking across different modalities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If the assistant system integrates visual data and dialog state tracking to resolve coreferences, then the user interaction relevance is improved, but the processing time increases due to multimodal analysis requirements

Engineering Contradiction:
Improveuser interaction relevanceVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary processing of visual data by extracting and storing visual states, entities, and relationships before they are needed for coreference resolution. The scene understanding engine pre-processes visual inputs to create a structured representation that can be quickly queried during dialog interactions, reducing real-time processing requirements.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary mechanism that bridges visual processing and dialog processing by creating a unified representation space. The coreference resolution system acts as an intermediary that queries both visual and dialog states to resolve references, efficiently matching between modalities without requiring full re-processing of all inputs.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20260073365A1Multimodal entity and coreference resolution for assistant systems
Publication Date: 2026.03.12 META PLATFORMS TECHNOLOGIES LLC
  • US20260073365A1 patent drawing
  • US20260073365A1 patent drawing
  • US20260073365A1 patent drawing

AI summary

In one embodiment, a method includes receiving, at a client system, an audio input, where the audio input comprises a coreference to a target object, accessing visual data from one or more camera associated with the client system, where the visual data comprises images portraying one or more objects, resolving the coreference to the target object from among the one or more objects, resoling the target object to a specific entity, and providing, at the client system, a response to the audio input, where the response comprises information about the specific entity.