Locally Distributed Keyword Detection for Private, Faster Voice Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-assisted media playback systems rely heavily on cloud-based voice assistant services for command processing, leading to potential privacy concerns and slower response times due to the need for data transmission and processing.
Innovation Solution
Implementing locally distributed keyword detection using network microphone devices with onboard command-keyword engines and local natural language understanding units to process voice commands without transmitting raw audio data to the cloud, enhancing privacy and response speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If cloud-based voice assistant services are used for command processing, then comprehensive voice processing capability is achieved, but response time increases and user privacy is compromised due to data transmission requirements
Solution Approach 1:
The patent segments the voice processing system into distributed network microphone devices that perform local keyword detection and cloud-based services that handle complex processing. Each device independently detects keywords locally, then transmits only relevant audio segments to the cloud, reducing transmission time and improving overall response speed while maintaining comprehensive processing capability.
Solution Approach 2:
The system performs preliminary keyword detection locally on each network microphone device before cloud transmission. This preliminary action filters out non-relevant audio data, so only audio containing detected keywords is transmitted to the cloud, significantly reducing data transmission time and improving response speed.
2Measurement precision
If cloud-based voice assistant services are used for command processing, then comprehensive voice processing capability is achieved, but user privacy is compromised due to data transmission requirements
Solution Approach 1:
The patent extracts and processes voice data locally on network microphone devices, performing keyword detection and filtering before cloud transmission. Only relevant audio segments containing detected keywords are transmitted to the cloud, minimizing the amount of personal voice data that leaves the local environment and thereby protecting user privacy.
Solution Approach 2:
The system introduces local keyword detection as an intermediary layer between the user's voice and the cloud service. This intermediary filters and processes voice data locally, acting as a privacy-protecting mediator that prevents raw voice data from being continuously transmitted to the cloud while still enabling comprehensive voice processing when needed.
3Loss of time
If local keyword detection is implemented on network microphone devices, then response time is improved and privacy is protected, but false positives increase due to limited local processing capability
Solution Approach 1:
The patent merges local keyword detection on network microphone devices with cloud-based verification services. Local devices perform rapid initial detection to maintain fast response times, while the cloud service performs secondary verification to reduce false positives, combining the advantages of both local and cloud processing.
Solution Approach 2:
The system implements feedback mechanisms where cloud-based services verify local keyword detections and provide correction signals back to the network microphone devices. This feedback loop allows the local devices to learn from cloud verification results, gradually improving their detection accuracy and reducing false positives over time.
Data Source
AI summary
In one aspect, a playback device includes a command-keyword engine having a local natural language unit (NLU). The playback device detects, via the command-keyword engine, a first command keyword in voice input of sound detected by one or more microphones of the playback device. The playback device determines whether the sound input data includes a keyword from a first predetermined library of keywords via a local natural language unit (NLU). The playback device transmits the input sound data to a second playback device over a local area network, the second playback device employing a second local NLU with a second predetermined library of keywords. The playback device receives a response from the second playback device and performs an action based on an intent determined by at least one of the first NLU or the second NLU according to the keywords in the voice input.


