Direction-Aware Speech Recognition Without Wake-Up Keywords

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition devices require a wake-up keyword for initiating services, which can be inconvenient for users and may not accurately determine user intentions, especially when multiple sound sources are present.

Innovation Solution

A speech recognition device and method that determine the location and direction of a sound source using microphones and algorithms like GCC-PHAT and SRP-PHAT, allowing for service provision without a wake-up keyword if the sound source is registered, and distinguishing between registered and non-registered directions to improve service accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If wake-up keyword is required for speech recognition service, then service can be initiated intentionally, but user convenience deteriorates and additional operation steps are needed

Engineering Contradiction:
Improveconvenience of initiating speech recognition serviceVSAvoidaccuracy of determining user intention
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system performs preliminary registration of sound sources and their directions before actual speech recognition occurs. During service operation, the pre-registered direction information is used to quickly determine whether to process the speech signal, eliminating the need for wake-up keywords while maintaining reliable user intention detection through the registered direction matching mechanism.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If speech recognition is provided without direction discrimination, then all sound sources are processed equally, but accuracy of determining user intention deteriorates when multiple sound sources are present

Engineering Contradiction:
Improveaccuracy of determining user intentionVSAvoidcomplexity of sound source discrimination
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the acoustic space by dividing it into multiple directional regions and registering specific sound sources in each region. When speech signals are received, the system determines which directional region the sound originates from and selectively processes only those signals from registered directions, thereby accurately discriminating between multiple sound sources without requiring complex analysis.

Inventive Principle:
Principle #1Segmentation

3Reliability

If direction-based service provision is implemented, then service accuracy is improved, but system complexity increases due to direction determination algorithms

Engineering Contradiction:
Improveservice accuracyVSAvoidcomplexity of direction determination system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs preliminary registration of sound sources and their directional information before actual speech recognition occurs. During service operation, the pre-registered direction information is used to quickly determine whether to process the speech signal, eliminating the need for wake-up keywords while maintaining reliable user intention detection through the registered direction matching mechanism.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3676830B1Speech recognition device and speech recognition method based on sound source direction
Publication Date: 2022.08.17 SAMSUNG ELECTRONICS CO LTD
  • EP3676830B1 patent drawingFigure 1
  • EP3676830B1 patent drawingFigure 2
  • EP3676830B1 patent drawingFigure 3

AI summary

A speech recognition device is provided. The speech recognition device includes at least one microphone configured to receive a sound signal from a first sound source, and at least one processor configured to determine a direction of the first sound source based on the sound signal, determine whether the direction of the first sound source is in a registered direction, and based on whether the direction of the first sound source is in the registered direction, recognize a speech from the sound signal regardless of whether the sound signal comprises a wake-up keyword.