Distributed Voice Processing for Remote Control Power and Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing remote control devices face challenges in accurately processing voice commands due to background noise, requiring increased power consumption and processing capabilities, which shortens battery life and increases network usage and latency.
Innovation Solution
The system distributes speech recognition performance between a remote control device and a cloud-based voice platform, using the remote control device for initial voice command processing and sending unclear commands to the cloud for further analysis, while also allowing integration with multiple digital assistants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If an audio responsive remote control processes voice input locally with faster processor and increased memory, then voice command recognition accuracy improves, but power consumption increases and battery life decreases
Solution Approach 1:
The patent segments the voice processing task into two parts: local processing for trigger word detection and cloud-based processing for full command analysis. The remote control device performs initial trigger word detection locally using minimal resources, then sends only the detected trigger words to the cloud-based voice service for comprehensive processing. This segmentation allows accurate voice recognition without continuously running high-power local processors.
2Measurement precision
If an audio responsive remote control sends voice input continuously to cloud voice service, then voice command recognition accuracy improves, but network consumption and latency increase
Solution Approach 1:
The patent extracts only the essential trigger word detection function to the local device, while sending only the detected trigger words (not the entire voice input stream) to the cloud service. This extraction approach maintains accurate voice recognition by leveraging cloud-based automated speech recognition and natural language processing, while significantly reducing network bandwidth consumption compared to continuous voice streaming.
3Device complexity
If an audio responsive remote control sends voice input to cloud voice service, then processing power requirements decrease, but response time increases due to network latency
Solution Approach 1:
The patent performs preliminary trigger word detection locally before sending data to the cloud service. By pre-processing the voice input to identify and extract only trigger words, the system reduces the amount of data that needs to be transmitted and processed in the cloud, thereby reducing overall response time despite the network latency inherent in cloud-based processing.
4Adaptability or versatility
If an audio responsive remote control is configured to work with multiple digital assistants, then versatility and task performance improve, but device complexity and configuration requirements increase
Solution Approach 1:
The patent implements a universal interface layer that allows the remote control device to work with multiple different digital assistants through a common communication protocol. The device can select among various digital assistants based on the detected trigger word, enabling versatile task performance across different assistant specialties without requiring separate dedicated devices for each assistant.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Disclosed herein are embodiments for an audio responsive electronic device. The audio responsive electronic device operates by receiving an indication that a user pressed the play/stop button. The audio responsive electronic device retrieves an intent from an intent queue that is associated with content previously paused. The audio responsive electronic device also retrieves state information associated with the paused content, and then causes content to be played based on the paused content and the state information. The audio responsive electronic device can receive an indication that a user selected tell me something functionality. In response, the audio responsive electronic device determines an identity of the user, a location of the identified user, and accesses information relating to the identified user. Based on this information, the audio responsive electronic device customizes a topic from a topic database and audibly provides the customized topic to the identified user.