Head-Mounted Assistant Co-Reference for Multimodal Entity Resolution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems face challenges in accurately identifying subjects and attributes from visual input, resolving entities from multimodal user input, and responding appropriately to user queries across different modalities.

Innovation Solution

The assistant system employs machine-learning models for facial recognition and object detection, a co-reference module to link voice and visual analysis, and determines suitable modalities based on contextual information for effective communication.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine-learning models are used for facial recognition and object detection, then measurement precision of visual input is improved, but device complexity increases

Engineering Contradiction:
Improveidentification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a co-reference module as an intermediary component that connects the visual analysis system with the voice input processing system. This module resolves entities by linking visual subjects (identified through machine-learning models) with their corresponding mentions in voice or text inputs, thereby improving overall identification accuracy without requiring the machine-learning models to directly handle all processing complexity themselves

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system is divided into distinct functional modules: visual analysis module (handling image processing and subject identification), co-reference module (handling entity resolution and linking), and response generation module. This segmentation allows each component to specialize in specific tasks, improving measurement precision while managing device complexity through modular architecture

Inventive Principle:
Principle #1Segmentation

2Reliability

If co-reference module is used to link voice and visual analysis, then reliability of entity resolution is improved, but device complexity increases

Engineering Contradiction:
Improveentity resolution accuracyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The co-reference module serves multiple functions: it identifies subjects in visual input, extracts entities from voice and text inputs, links corresponding entities across modalities, and resolves ambiguities. This multi-functionality improves reliability of entity resolution while avoiding the need for separate specialized components for each function, thereby managing device complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If multiple modalities are supported for user input, then adaptability of the system is improved, but device complexity increases

Engineering Contradiction:
Improvemulti-modal supportVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system is designed with universal processing capabilities that handle multiple input modalities (visual, voice, text) through a unified architecture. The co-reference module and response generation system work consistently across all modalities, improving adaptability while avoiding the need for entirely separate processing pipelines for each modality

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically determines which modalities to use based on contextual information and user preferences. The response generation can adaptively select from text, voice, or visual responses, and can switch between active listening modes and proactive task suggestions, providing versatility while managing complexity through dynamic rather than static configuration

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12406316B2Processing multimodal user input for assistant systems
Publication Date: 2025.09.02 META PLATFORMS INC
  • US12406316B2 patent drawing
  • US12406316B2 patent drawing
  • US12406316B2 patent drawing

AI summary

In one embodiment, a method includes receiving at a head-mounted device a speech input from a user and a visual input captured by cameras of the head-mounted device, wherein the visual input comprises subjects and attributes associated with the subjects, and wherein the speech input comprises a co-reference to one or more of the subjects, resolving entities corresponding to the subjects associated with the co-reference based on the attributes and the co-reference, and presenting a communication content responsive to the speech input and the visual input at the head-mounted device, wherein the communication content comprises information associated with executing results of tasks corresponding to the resolved entities.