Voice-Enabled Device Latency Reduction via Intermediary Mirroring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice-enabled devices experience high latency and poor user experience when communicating with remote systems in environments with poor connectivity, such as vehicles, due to limitations in local processing capabilities and high-latency connections.

Innovation Solution

Implementing a dual-device system where a lower-latency device mirrors the local speech processing of a higher-latency device, allowing the remote system to predict and initiate processing of utterances, thereby reducing latency and improving response times for higher-latency devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If voice-enabled devices rely on remote systems for speech processing, then processing capabilities are improved, but latency increases in environments with poor connectivity

Engineering Contradiction:
Improvespeech processing capabilityVSAvoidlatency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The lower-latency device performs speech processing in advance and notifies the remote system before the higher-latency device would otherwise need to request processing. This preliminary action reduces the effective latency for the higher-latency device by having work already completed or in progress when needed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary lower-latency device that acts as a mediator between the remote system and the higher-latency device. This intermediary performs speech processing locally and can provide results to the higher-latency device without requiring constant communication with the remote system, thus reducing latency in poor connectivity environments.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If multiple voice-enabled devices are deployed in an environment, then coverage and accessibility are improved, but system complexity increases

Engineering Contradiction:
Improveenvironmental coverageVSAvoidsystem configuration
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent uses a lower-latency device as a copy or mirror of the higher-latency device's speech processing functionality. Instead of requiring each device to have full processing capabilities, the lower-latency device copies the necessary processing functions, simplifying the overall system architecture while maintaining multiple deployment points.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12125489B1Speech recognition using multiple voice-enabled devices
Publication Date: 2024.10.22 AMAZON TECH INC
  • US12125489B1 patent drawing
  • US12125489B1 patent drawing
  • US12125489B1 patent drawing

AI summary

Techniques for using multiple voice-enabled devices in a user environment to reduce the latency for obtaining responses to user utterances from a remote system. The voice-enabled devices may each establish connections with the remote system to have the remote system perform supplemental speech processing for utterances the devices are unable to process locally. One voice-enabled device may have a higher-latency connection to the remote system, and another voice-enabled device may have a lower-latency connection to the remote system. The lower-latency device may send an utterance to the remote system before the higher-latency device is able, and the remote system may begin processing the utterance faster than if the lower-latency device sent the utterance. The remote system may then provide a response for the utterance to the higher-latency device in less time than if the remote system had to wait for the utterance from the higher-latency device.