Server Clustering Speech Similarity to Filter False Wakeups
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
False wakeups of AI assistants occur due to noise patterns similar to trigger keywords, leading to unintended activation of multiple devices and resource wastage, causing inconvenience to users.
Innovation Solution
A server with a communication module and processor that generates clusters based on similarities among multiple speech inputs within a preset period, determining whether to respond to each cluster to differentiate between intended and unintended activations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the voice assistant monitors ambient sounds continuously to detect trigger keywords, then the assistant can respond to user speech inputs, but false wakeups occur when noise patterns similar to keywords are detected
Solution Approach 1:
The patent merges multiple speech detection devices (microphones) to form an array that collectively analyzes ambient sounds. By combining signals from multiple microphones and performing joint speech pattern recognition, the system achieves more reliable distinction between actual trigger keywords and similar noise patterns, thereby reducing false wakeups while maintaining accurate detection.
Solution Approach 2:
The system implements feedback mechanisms where speech recognition results from multiple microphones are analyzed collectively. The server receives speech data from multiple devices, performs comprehensive pattern matching, and provides feedback to determine whether activation is warranted. This feedback loop enables the system to learn from multiple data sources and make more accurate wakeup decisions.
2Productivity
If multiple devices simultaneously wake up due to similar noise patterns, then the system responds to all detected speeches, but resource wastage occurs and user convenience deteriorates
Solution Approach 1:
The patent combines speech detection capabilities across multiple electronic devices into a unified system. Instead of each device independently waking up and consuming resources, the system merges speech data from multiple microphones and performs centralized speech pattern recognition. This consolidation ensures that only genuine trigger keywords activate the voice assistant, preventing redundant resource consumption from false wakeups while maintaining comprehensive monitoring coverage.
3Reliability
If the system activates voice assistant for every detected speech pattern, then no false wakeups are missed, but user convenience is reduced due to unintended activations
Solution Approach 1:
The system merges speech recognition results from multiple electronic devices and microphones to achieve more reliable detection. By analyzing combined data from multiple sources, the system can distinguish genuine user intent from environmental noise with higher confidence. This approach maintains high detection sensitivity for actual trigger keywords while significantly reducing false activations that would inconvenience users.
Data Source
AI summary
A server is provided. The server includes a communication circuitry, and at least one processor operatively connected with the communication circuitry. The at least one processor may be configured to, in response to traffic of a plurality of speeches to wake up a voice assistant feature, received within a preset period being a preset value or more, generate a plurality of clusters based on similarities between the plurality of speeches, and determine whether to respond to each of speeches included in each of the plurality of clusters based on similarities between the speeches included in each of the plurality of clusters.


