Distributed Voice Processing for Remote Control Power and Latency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing remote control devices face challenges in accurately processing voice commands due to background noise, requiring increased power consumption and processing capabilities, which shortens battery life and increases network usage and latency.

Innovation Solution

The system distributes speech recognition performance between a remote control device and a cloud-based voice platform, using the remote control device for initial voice command processing and sending unclear commands to the cloud for further analysis, while also allowing integration with multiple digital assistants.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If an audio responsive remote control processes voice input locally with faster processor and increased memory, then voice command recognition accuracy improves, but power consumption increases and battery life decreases

Engineering Contradiction:
Improvevoice command recognition accuracyVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the voice processing task into two parts: local processing for trigger word detection and cloud-based processing for full command analysis. The remote control device performs initial trigger word detection locally using minimal resources, then sends only the detected trigger words to the cloud-based voice service for comprehensive processing. This segmentation allows accurate voice recognition without continuously running high-power local processors.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If an audio responsive remote control sends voice input continuously to cloud voice service, then voice command recognition accuracy improves, but network consumption and latency increase

Engineering Contradiction:
Improvevoice command recognition accuracyVSAvoidnetwork consumption
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent extracts only the essential trigger word detection function to the local device, while sending only the detected trigger words (not the entire voice input stream) to the cloud service. This extraction approach maintains accurate voice recognition by leveraging cloud-based automated speech recognition and natural language processing, while significantly reducing network bandwidth consumption compared to continuous voice streaming.

Inventive Principle:
Principle #2Taking out (Extraction)

3Device complexity

If an audio responsive remote control sends voice input to cloud voice service, then processing power requirements decrease, but response time increases due to network latency

Engineering Contradiction:
Improveprocessor requirementsVSAvoidresponse time
Core Design Contradiction:
Device complexityVSLoss of time

Solution Approach 1:

The patent performs preliminary trigger word detection locally before sending data to the cloud service. By pre-processing the voice input to identify and extract only trigger words, the system reduces the amount of data that needs to be transmitted and processed in the cloud, thereby reducing overall response time despite the network latency inherent in cloud-based processing.

Inventive Principle:
Principle #10Preliminary action

4Adaptability or versatility

If an audio responsive remote control is configured to work with multiple digital assistants, then versatility and task performance improve, but device complexity and configuration requirements increase

Engineering Contradiction:
Improvedigital assistant compatibilityVSAvoidconfiguration complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal interface layer that allows the remote control device to work with multiple different digital assistants through a common communication protocol. The device can select among various digital assistants based on the detected trigger word, enabling versatile task performance across different assistant specialties without requiring separate dedicated devices for each assistant.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3676832B1Audio responsive device with play/stop and tell me something buttons
Publication Date: 2025.04.30 ROKU INC
  • EP3676832B1 patent drawingFigure 1
  • EP3676832B1 patent drawingFigure 2
  • EP3676832B1 patent drawingFigure 3

AI summary

Disclosed herein are embodiments for an audio responsive electronic device. The audio responsive electronic device operates by receiving an indication that a user pressed the play/stop button. The audio responsive electronic device retrieves an intent from an intent queue that is associated with content previously paused. The audio responsive electronic device also retrieves state information associated with the paused content, and then causes content to be played based on the paused content and the state information. The audio responsive electronic device can receive an indication that a user selected tell me something functionality. In response, the audio responsive electronic device determines an identity of the user, a location of the identified user, and accesses information relating to the identified user. Based on this information, the audio responsive electronic device customizes a topic from a topic database and audibly provides the customized topic to the identified user.