Context-Based Device Arbitration for Voice Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In environments with multiple voice-enabled devices, existing systems often lead to undesirable user experiences due to multiple devices independently processing and responding to voice commands, resulting in unintended actions.

Innovation Solution

A speech processing system that uses contextual information, such as audio signal metrics and device states, to arbitrate which voice-enabled device should respond to a speech utterance, ensuring the most appropriate device performs the intended action.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If multiple voice-enabled devices independently process voice commands, then each device can respond to user input, but multiple devices may perform the same task creating an undesirable user experience

Engineering Contradiction:
ImproveVoice command responsivenessVSAvoidIntended device selection
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

A speech processing system acts as an intermediary between multiple voice-enabled devices and the user. The system receives audio signals from multiple devices, processes them to determine which device the user intended to address, and routes the command to the correct device. This mediator approach prevents multiple devices from independently acting on the same command while maintaining voice command responsiveness.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The speech processing system analyzes audio signals from multiple devices and uses feedback mechanisms to determine which device the user intended to address. By evaluating audio characteristics and device states, the system provides feedback to select the appropriate target device, ensuring reliable device selection while maintaining ease of operation.

Inventive Principle:
Principle #23Feedback

2Reliability

If a speech processing system arbitrates between multiple devices, then the correct device can be selected, but additional processing complexity is introduced

Engineering Contradiction:
ImproveIntended device selectionVSAvoidSpeech processing system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The speech processing system performs multiple functions: receiving audio signals from multiple devices, processing the signals to determine user intent, evaluating device states, and routing commands to the appropriate device. By consolidating these diverse functions into a single multi-functional system, the patent achieves reliable device selection without proportionally increasing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240404521A1Context-based device arbitration
Publication Date: 2024.12.05 AMAZON TECH INC
  • US20240404521A1 patent drawing
  • US20240404521A1 patent drawing
  • US20240404521A1 patent drawing

AI summary

This disclosure describes, in part, context-based device arbitration techniques to select a voice-enabled device from multiple voice-enabled devices to provide a response to a command included in a speech utterance of a user. In some examples, the context-driven arbitration techniques may include determining a ranked list of voice-enabled devices that are ranked based on audio signal metric values for audio signals generated by each voice-enabled device, and iteratively moving through the list to determine, based on device states of the voice-enabled devices, whether one of the voice-enabled devices can perform an action responsive to the command. If the voice-enabled devices that detected the speech utterance are unable to perform the action responsive to the command, all other voice-enabled devices associated with an account may be analyzed to determine whether one of the other voice-enabled devices can perform the action responsive to the command in the speech utterance.