Head-Mounted Assistant Co-Reference for Multimodal Entity Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems face challenges in accurately identifying subjects and attributes from visual input, resolving entities from multimodal user input, and responding appropriately to user queries across different modalities.
Innovation Solution
The assistant system employs machine-learning models for facial recognition and object detection, a co-reference module to link voice and visual analysis, and determines suitable modalities based on contextual information for effective communication.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If machine-learning models are used for facial recognition and object detection, then measurement precision of visual input is improved, but device complexity increases
Solution Approach 1:
The patent introduces a co-reference module as an intermediary component that connects the visual analysis system with the voice input processing system. This module resolves entities by linking visual subjects (identified through machine-learning models) with their corresponding mentions in voice or text inputs, thereby improving overall identification accuracy without requiring the machine-learning models to directly handle all processing complexity themselves
Solution Approach 2:
The system is divided into distinct functional modules: visual analysis module (handling image processing and subject identification), co-reference module (handling entity resolution and linking), and response generation module. This segmentation allows each component to specialize in specific tasks, improving measurement precision while managing device complexity through modular architecture
2Reliability
If co-reference module is used to link voice and visual analysis, then reliability of entity resolution is improved, but device complexity increases
Solution Approach 1:
The co-reference module serves multiple functions: it identifies subjects in visual input, extracts entities from voice and text inputs, links corresponding entities across modalities, and resolves ambiguities. This multi-functionality improves reliability of entity resolution while avoiding the need for separate specialized components for each function, thereby managing device complexity
3Adaptability or versatility
If multiple modalities are supported for user input, then adaptability of the system is improved, but device complexity increases
Solution Approach 1:
The system is designed with universal processing capabilities that handle multiple input modalities (visual, voice, text) through a unified architecture. The co-reference module and response generation system work consistently across all modalities, improving adaptability while avoiding the need for entirely separate processing pipelines for each modality
Solution Approach 2:
The system dynamically determines which modalities to use based on contextual information and user preferences. The response generation can adaptively select from text, voice, or visual responses, and can switch between active listening modes and proactive task suggestions, providing versatility while managing complexity through dynamic rather than static configuration
Data Source
AI summary
In one embodiment, a method includes receiving at a head-mounted device a speech input from a user and a visual input captured by cameras of the head-mounted device, wherein the visual input comprises subjects and attributes associated with the subjects, and wherein the speech input comprises a co-reference to one or more of the subjects, resolving entities corresponding to the subjects associated with the co-reference based on the attributes and the co-reference, and presenting a communication content responsive to the speech input and the visual input at the head-mounted device, wherein the communication content comprises information associated with executing results of tasks corresponding to the resolved entities.


