Hybrid Voice Command Routing for Privacy and Low-Latency Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice-controlled systems, such as smart speakers, face issues with user privacy due to the streaming of audio voice commands over network-based servers, which can lead to data misappropriation, and they often require physical interaction when driving, making it inconvenient. Additionally, they may experience high false detection rates and latency due to the need for continuous network processing.
Innovation Solution
A hybrid voice recognition system that includes a non-transitory computer readable storage medium with command phrases and a neural network circuit, allowing for continuous monitoring of audio signals to detect command phrases locally without immediate transmission, using a wake word to activate command recognition and optionally transmitting audio signals only when specific commands are given, thereby maintaining user privacy and reducing latency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If audio voice commands are streamed to network-based servers for interpretation, then voice recognition capability is improved, but user privacy deteriorates due to data misappropriation risks
Solution Approach 1:
The system segments voice command processing into two categories: simple commands processed locally on the device and complex commands processed on remote servers. This segmentation allows the system to maintain voice recognition versatility while minimizing privacy risks by limiting network transmission to only necessary cases.
Solution Approach 2:
The device performs self-service by processing simple voice commands locally without requiring network connectivity. The on-device processing capability allows the device to independently handle routine commands, reducing dependency on external servers and protecting user privacy.
2Measurement precision
If voice commands are processed continuously on network servers, then recognition accuracy is improved, but response latency increases
Solution Approach 1:
The system segments commands based on complexity and processing requirements. Simple, frequently-used commands are processed instantly on-device, providing rapid response. Complex commands are routed to servers for enhanced processing, accepting longer latency only when necessary for accuracy.
Solution Approach 2:
The system applies partial processing on-device for immediate response to simple commands, and excessive processing on-servers only when needed for complex commands. This selective approach optimizes the balance between response time and recognition accuracy.
3Adaptability or versatility
If all audio signals are transmitted to external networks for processing, then command recognition capability is improved, but data security deteriorates
Solution Approach 1:
The system segments audio processing into local and remote components. Only audio containing complex commands or uncertain recognition results is transmitted to external networks. Routine commands are processed locally, maintaining data security while preserving overall command recognition capability.
Solution Approach 2:
The device maintains self-service capability by processing commands locally without external network involvement. This self-sufficiency ensures data security by keeping sensitive audio data on-device whenever possible, while still providing comprehensive command recognition through selective cloud processing.
Data Source
AI summary
Systems and methods are presented for recognizing and responding to voice commands at a local system and selectively streaming audio to a network-based computing system to recognize voice commands when the user provides a specific voice command to stream to the network-based computing system and/or when the user provides a voice command that is not recognizable by the local system.


