Server Clustering Speech Similarity to Filter False Wakeups

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

False wakeups of AI assistants occur due to noise patterns similar to trigger keywords, leading to unintended activation of multiple devices and resource wastage, causing inconvenience to users.

Innovation Solution

A server with a communication module and processor that generates clusters based on similarities among multiple speech inputs within a preset period, determining whether to respond to each cluster to differentiate between intended and unintended activations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the voice assistant monitors ambient sounds continuously to detect trigger keywords, then the assistant can respond to user speech inputs, but false wakeups occur when noise patterns similar to keywords are detected

Engineering Contradiction:
Improveaccuracy of wakeup detectionVSAvoidfalse wakeup
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent merges multiple speech detection devices (microphones) to form an array that collectively analyzes ambient sounds. By combining signals from multiple microphones and performing joint speech pattern recognition, the system achieves more reliable distinction between actual trigger keywords and similar noise patterns, thereby reducing false wakeups while maintaining accurate detection.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system implements feedback mechanisms where speech recognition results from multiple microphones are analyzed collectively. The server receives speech data from multiple devices, performs comprehensive pattern matching, and provides feedback to determine whether activation is warranted. This feedback loop enables the system to learn from multiple data sources and make more accurate wakeup decisions.

Inventive Principle:
Principle #23Feedback

2Productivity

If multiple devices simultaneously wake up due to similar noise patterns, then the system responds to all detected speeches, but resource wastage occurs and user convenience deteriorates

Engineering Contradiction:
Improveresponse capabilityVSAvoidresource wastage
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent combines speech detection capabilities across multiple electronic devices into a unified system. Instead of each device independently waking up and consuming resources, the system merges speech data from multiple microphones and performs centralized speech pattern recognition. This consolidation ensures that only genuine trigger keywords activate the voice assistant, preventing redundant resource consumption from false wakeups while maintaining comprehensive monitoring coverage.

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If the system activates voice assistant for every detected speech pattern, then no false wakeups are missed, but user convenience is reduced due to unintended activations

Engineering Contradiction:
Improvedetection sensitivityVSAvoiduser convenience
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system merges speech recognition results from multiple electronic devices and microphones to achieve more reliable detection. By analyzing combined data from multiple sources, the system can distinguish genuine user intent from environmental noise with higher confidence. This approach maintains high detection sensitivity for actual trigger keywords while significantly reducing false activations that would inconvenience users.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11967322B2Server for identifying false wakeup and method for controlling the same
Publication Date: 2024.04.23 SAMSUNG ELECTRONICS CO LTD
  • US11967322B2 patent drawing
  • US11967322B2 patent drawing
  • US11967322B2 patent drawing

AI summary

A server is provided. The server includes a communication circuitry, and at least one processor operatively connected with the communication circuitry. The at least one processor may be configured to, in response to traffic of a plurality of speeches to wake up a voice assistant feature, received within a preset period being a preset value or more, generate a plurality of clusters based on similarities between the plurality of speeches, and determine whether to respond to each of speeches included in each of the plurality of clusters based on similarities between the speeches included in each of the plurality of clusters.