Direction-Aware Speech Recognition Without Wake-Up Keywords
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition devices require a wake-up keyword for initiating services, which can be inconvenient for users and may not accurately determine user intentions, especially when multiple sound sources are present.
Innovation Solution
A speech recognition device and method that determine the location and direction of a sound source using microphones and algorithms like GCC-PHAT and SRP-PHAT, allowing for service provision without a wake-up keyword if the sound source is registered, and distinguishing between registered and non-registered directions to improve service accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If wake-up keyword is required for speech recognition service, then service can be initiated intentionally, but user convenience deteriorates and additional operation steps are needed
Solution Approach 1:
The system performs preliminary registration of sound sources and their directions before actual speech recognition occurs. During service operation, the pre-registered direction information is used to quickly determine whether to process the speech signal, eliminating the need for wake-up keywords while maintaining reliable user intention detection through the registered direction matching mechanism.
2Reliability
If speech recognition is provided without direction discrimination, then all sound sources are processed equally, but accuracy of determining user intention deteriorates when multiple sound sources are present
Solution Approach 1:
The system segments the acoustic space by dividing it into multiple directional regions and registering specific sound sources in each region. When speech signals are received, the system determines which directional region the sound originates from and selectively processes only those signals from registered directions, thereby accurately discriminating between multiple sound sources without requiring complex analysis.
3Reliability
If direction-based service provision is implemented, then service accuracy is improved, but system complexity increases due to direction determination algorithms
Solution Approach 1:
The system performs preliminary registration of sound sources and their directional information before actual speech recognition occurs. During service operation, the pre-registered direction information is used to quickly determine whether to process the speech signal, eliminating the need for wake-up keywords while maintaining reliable user intention detection through the registered direction matching mechanism.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A speech recognition device is provided. The speech recognition device includes at least one microphone configured to receive a sound signal from a first sound source, and at least one processor configured to determine a direction of the first sound source based on the sound signal, determine whether the direction of the first sound source is in a registered direction, and based on whether the direction of the first sound source is in the registered direction, recognize a speech from the sound signal regardless of whether the sound signal comprises a wake-up keyword.