Speech Direction Detection for Voice Command Targeting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional electronic devices equipped with speech recognition struggle to accurately determine whether they are the intended target of a voice command, leading to unintended responses when multiple devices are present in close proximity.
Innovation Solution
The method involves analyzing the direction of departure of speech by calculating a ratio between energy values of high and low frequency ranges or determining spectral flatness, allowing devices to determine if the speech is directed towards themselves and thus decide whether to process voice commands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple electronic devices are equipped with speech recognition function, then user convenience is enhanced, but devices cannot accurately determine whether they are the intended target of a voice command
Solution Approach 1:
The patent introduces a new dimension for target identification by analyzing the directional characteristics of speech. Instead of relying solely on device proximity or signal strength, the system evaluates the direction of departure of speech from the user's mouth by comparing frequency characteristics (high vs. low frequency energy ratios or spectral flatness) of speech received at different locations on the device. This dimensional addition allows multiple devices to coexist while accurately identifying the intended target through directional discrimination.
2Productivity
If devices respond to any received speech, then response coverage is maximized, but false activations occur when speech is not directed at the device
Solution Approach 1:
The patent applies local quality by making different parts of the received speech signal serve different functions. Specifically, speech received at different physical locations on the device (e.g., top vs. bottom, left vs. right) is analyzed with different frequency characteristics. By comparing the high-frequency-to-low-frequency energy ratio or spectral flatness values from different locations, the system determines whether speech is directed toward the device. This localized analysis enables the device to distinguish between intentional commands and ambient speech, reducing false activations while maintaining comprehensive response coverage.
3Measurement precision
If directional analysis is implemented, then target identification accuracy is improved, but device complexity increases
Solution Approach 1:
The patent implements parameter changes by transforming the speech signal into different frequency domains and analyzing specific parameters (high-frequency energy ratio or spectral flatness) to determine speech direction. Instead of using complex multi-microphone arrays or sophisticated signal processing algorithms, the system changes the parameter being measured from general speech presence to frequency-based directional characteristics. This parameter transformation achieves accurate target identification while keeping the device complexity manageable through computationally efficient frequency analysis.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method for controlling an electronic device in response to speech spoken by a user is disclosed. The method may include receiving an input sound by a sound sensor. The method may also detect the speech spoken by the user in the input sound, determine first characteristics of a first frequency range and second characteristics of a second frequency range of the speech in response to detecting the speech in the input sound, and determine whether a direction of departure of the speech spoken by the user is toward the electronic device based on the first and second characteristics.