Context-Aware NLP for Spoken Commands on Displayed Content
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-enabled electronic devices struggle to accurately interpret spoken commands that reference or relate to displayed content due to the lack of consideration of visual context in natural language understanding systems.
Innovation Solution
A system that processes spoken commands by integrating audio data with contextual data from the display screen, using a neural-network-based model to generate scores for intent and named entities, enhancing the accuracy of natural language understanding by considering both textual and visual inputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voice-enabled electronic devices use traditional natural language understanding systems that process only audio data, then the system complexity remains lower, but the accuracy of interpreting spoken commands that reference displayed content deteriorates
Solution Approach 1:
The patent merges audio data processing and visual context data processing into a unified natural language understanding system. The system combines speech recognition output with display content information (such as displayed text, images, or interface elements) to create a more comprehensive understanding of user commands, thereby improving accuracy without requiring separate independent systems
Solution Approach 2:
The natural language understanding system is designed to handle multiple types of input data (audio and visual) through a single unified processing architecture. The system can interpret commands that reference displayed content by universally processing both audio transcriptions and visual context information, eliminating the need for separate specialized systems
2Measurement precision
If the system processes only audio data for natural language understanding, then the processing time remains shorter, but the accuracy of determining user intent deteriorates when displayed content is relevant
Solution Approach 1:
The system performs preliminary processing of display content into structured contextual data (identifying displayed elements, extracting text, analyzing images) before the audio command is fully processed. This pre-processing of visual context allows the natural language understanding system to quickly integrate visual and audio information without significant delay, as the visual data is already prepared and indexed for rapid retrieval during command processing
3Measurement precision
If the system integrates contextual data from the display screen with audio data, then the accuracy of entity recognition improves, but the device complexity increases
Solution Approach 1:
The patent segments the data processing pipeline into distinct modules: audio processing module for speech recognition, visual context processing module for analyzing display content, and integration module for combining both data types. This segmentation allows each module to specialize in specific tasks, making the overall complex system more manageable and easier to implement while maintaining high entity recognition accuracy through modular architecture
Data Source
AI summary
Multi-modal natural language processing systems are provided. Some systems are context-aware systems that use multi-modal data to improve the accuracy of natural language understanding as it is applied to spoken language input. Machine learning architectures are provided that jointly model spoken language input (“utterances”) and information displayed on a visual display (“on-screen information”). Such machine learning architectures can improve upon, and solve problems inherent in, existing spoken language understanding systems that operate in multi-modal contexts.


