Context-Aware NLP for Spoken Commands on Displayed Content

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice-enabled electronic devices struggle to accurately interpret spoken commands that reference or relate to displayed content due to the lack of consideration of visual context in natural language understanding systems.

Innovation Solution

A system that processes spoken commands by integrating audio data with contextual data from the display screen, using a neural-network-based model to generate scores for intent and named entities, enhancing the accuracy of natural language understanding by considering both textual and visual inputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice-enabled electronic devices use traditional natural language understanding systems that process only audio data, then the system complexity remains lower, but the accuracy of interpreting spoken commands that reference displayed content deteriorates

Engineering Contradiction:
Improveaccuracy of interpreting spoken commandsVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges audio data processing and visual context data processing into a unified natural language understanding system. The system combines speech recognition output with display content information (such as displayed text, images, or interface elements) to create a more comprehensive understanding of user commands, thereby improving accuracy without requiring separate independent systems

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The natural language understanding system is designed to handle multiple types of input data (audio and visual) through a single unified processing architecture. The system can interpret commands that reference displayed content by universally processing both audio transcriptions and visual context information, eliminating the need for separate specialized systems

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If the system processes only audio data for natural language understanding, then the processing time remains shorter, but the accuracy of determining user intent deteriorates when displayed content is relevant

Engineering Contradiction:
Improveaccuracy of determining user intentVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary processing of display content into structured contextual data (identifying displayed elements, extracting text, analyzing images) before the audio command is fully processed. This pre-processing of visual context allows the natural language understanding system to quickly integrate visual and audio information without significant delay, as the visual data is already prepared and indexed for rapid retrieval during command processing

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If the system integrates contextual data from the display screen with audio data, then the accuracy of entity recognition improves, but the device complexity increases

Engineering Contradiction:
Improveaccuracy of entity recognitionVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the data processing pipeline into distinct modules: audio processing module for speech recognition, visual context processing module for analyzing display content, and integration module for combining both data types. This segmentation allows each module to specialize in specific tasks, making the overall complex system more manageable and easier to implement while maintaining high entity recognition accuracy through modular architecture

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12579973B2Natural language processing with contextual data associated with content displayed by a computing device
Publication Date: 2026.03.17 AMAZON TECH INC
  • US12579973B2 patent drawing
  • US12579973B2 patent drawing
  • US12579973B2 patent drawing

AI summary

Multi-modal natural language processing systems are provided. Some systems are context-aware systems that use multi-modal data to improve the accuracy of natural language understanding as it is applied to spoken language input. Machine learning architectures are provided that jointly model spoken language input (“utterances”) and information displayed on a visual display (“on-screen information”). Such machine learning architectures can improve upon, and solve problems inherent in, existing spoken language understanding systems that operate in multi-modal contexts.