Voice Command Context Recognition via Audio Watermarks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice-activated devices lack the ability to accurately recognize context and environmental factors associated with voice commands, leading to reduced accuracy in responding to user inputs.

Innovation Solution

The system determines executable operations by receiving audio data and voice commands, identifying context through audio identifiers such as content identifiers or watermarks, and executing operations based on these identifiers and commands, allowing for enhanced user interaction with content assets and environmental triggers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice-activated devices process only basic voice commands without context analysis, then device complexity remains low, but measurement precision of user intent deteriorates

Engineering Contradiction:
Improveaccuracy of voice command recognitionVSAvoidcomplexity of context processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by detecting audio identifiers (watermarks, content IDs) from playing content before processing the voice command. This pre-detection of context information allows the voice command processor to focus only on interpreting user intent within the already-established context, improving accuracy without proportionally increasing overall system complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary context layer that sits between the raw voice command and the execution system. Audio identifiers from playing content serve as intermediaries that carry contextual information (showing, playing, paused) without requiring the entire content to be processed. This intermediary approach enables precise intent recognition while maintaining manageable system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If voice-activated devices ignore context and environmental factors, then device complexity remains simple, but reliability of command execution deteriorates

Engineering Contradiction:
Improveaccuracy of responding to voice commandsVSAvoidcomplexity of context recognition system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system applies local quality by focusing context processing only on relevant aspects of the playing content. Instead of analyzing entire media files, the system extracts specific audio identifiers (watermarks, content IDs) at key moments. This localized approach to context processing improves reliability of command execution while avoiding the complexity of comprehensive content analysis.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes parameters by transforming complex content context into simplified identifier parameters. Audio watermarks and content IDs serve as compressed representations of complex media content, converting high-dimensional context information into low-dimensional parameters that can be efficiently processed. This parameter transformation maintains reliability while reducing system complexity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240040181A1Determining context to initiate interactivity
Publication Date: 2024.02.01 COMCAST CABLE COMM LLC
  • US20240040181A1 patent drawing
  • US20240040181A1 patent drawing
  • US20240040181A1 patent drawing

AI summary

Methods and systems are disclosed for executing a voice command based on and association of the voice command and one or more identifiers. Audio data associated with a content asset may be received at a user device such as a voice activated device. A voice command may also be received at the user device. One or more identifiers associated with the audio data, such as a content or product identifier, may be determined. The identifiers may be determined based on playback of the content asset or may be received in response to a request generated by the user device. One or more operations capable of being executed by the user device may be determined and initiated or executed by the user device based on the one or more identifiers and the received voice command.