Speech Direction Detection for Voice Command Targeting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional electronic devices equipped with speech recognition struggle to accurately determine whether they are the intended target of a voice command, leading to unintended responses when multiple devices are present in close proximity.

Innovation Solution

The method involves analyzing the direction of departure of speech by calculating a ratio between energy values of high and low frequency ranges or determining spectral flatness, allowing devices to determine if the speech is directed towards themselves and thus decide whether to process voice commands.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple electronic devices are equipped with speech recognition function, then user convenience is enhanced, but devices cannot accurately determine whether they are the intended target of a voice command

Engineering Contradiction:
Improvespeech recognition functionVSAvoidtarget identification accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces a new dimension for target identification by analyzing the directional characteristics of speech. Instead of relying solely on device proximity or signal strength, the system evaluates the direction of departure of speech from the user's mouth by comparing frequency characteristics (high vs. low frequency energy ratios or spectral flatness) of speech received at different locations on the device. This dimensional addition allows multiple devices to coexist while accurately identifying the intended target through directional discrimination.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If devices respond to any received speech, then response coverage is maximized, but false activations occur when speech is not directed at the device

Engineering Contradiction:
Improveresponse coverageVSAvoidfalse activation rate
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies local quality by making different parts of the received speech signal serve different functions. Specifically, speech received at different physical locations on the device (e.g., top vs. bottom, left vs. right) is analyzed with different frequency characteristics. By comparing the high-frequency-to-low-frequency energy ratio or spectral flatness values from different locations, the system determines whether speech is directed toward the device. This localized analysis enables the device to distinguish between intentional commands and ambient speech, reducing false activations while maintaining comprehensive response coverage.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If directional analysis is implemented, then target identification accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvetarget identification accuracyVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements parameter changes by transforming the speech signal into different frequency domains and analyzing specific parameters (high-frequency energy ratio or spectral flatness) to determine speech direction. Instead of using complex multi-microphone arrays or sophisticated signal processing algorithms, the system changes the parameter being measured from general speech presence to frequency-based directional characteristics. This parameter transformation achieves accurate target identification while keeping the device complexity manageable through computationally efficient frequency analysis.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3274988B1Controlling electronic device based on direction of speech
Publication Date: 2019.08.07 QUALCOMM INC
  • EP3274988B1 patent drawingFigure 1
  • EP3274988B1 patent drawingFigure 2
  • EP3274988B1 patent drawingFigure 3

AI summary

A method for controlling an electronic device in response to speech spoken by a user is disclosed. The method may include receiving an input sound by a sound sensor. The method may also detect the speech spoken by the user in the input sound, determine first characteristics of a first frequency range and second characteristics of a second frequency range of the speech in response to detecting the speech in the input sound, and determine whether a direction of departure of the speech spoken by the user is toward the electronic device based on the first and second characteristics.