Voice Input Processing Using Speaker Model Endpoint Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Electronic devices face challenges in accurately processing voice data due to noise interference from environmental sounds and other voices, which reduces preprocessing and recognition efficiency.
Innovation Solution
An electronic device is equipped with a processor, memory, microphone, and communication interface, using speaker recognition to determine a speaker model and detect the end-point of a user's utterance, thereby isolating and processing the user's voice data effectively.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If the electronic device receives voice data in a noisy environment, then the device can capture more audio information, but the recognition accuracy decreases due to noise interference
Solution Approach 1:
The patent segments the audio signal into multiple frequency bands using a filter bank structure. Each band processes audio information independently, allowing the system to capture comprehensive audio data while identifying and suppressing noise in specific frequency ranges, thus maintaining recognition accuracy despite noisy environmental conditions
Solution Approach 2:
The patent applies different processing characteristics to different frequency bands. By assigning unique filter properties and processing parameters to each frequency band, the system optimizes noise suppression and voice recognition for specific spectral regions while preserving overall audio information quality
2Loss of information
If the electronic device processes all received audio data, then complete voice information is captured, but preprocessing efficiency decreases due to noise data
Solution Approach 1:
The patent extracts and isolates voice signals from the mixed audio stream by analyzing spectral characteristics across frequency bands. The system identifies and extracts relevant voice information while separating and discarding noise components, thereby maintaining complete voice information capture while improving preprocessing efficiency through selective data processing
Solution Approach 2:
The patent performs preliminary noise suppression and voice detection in the frequency domain before final voice recognition processing. By pre-processing the audio signal to remove obvious noise components and identify voice segments, the system reduces the computational burden on subsequent recognition stages, thereby improving overall preprocessing efficiency without losing voice information
3Device complexity
If the electronic device uses traditional end-point detection without speaker recognition, then the processing is simpler, but the detection accuracy decreases in noisy environments
Solution Approach 1:
The patent transitions from traditional time-domain end-point detection to a frequency-domain approach using spectral analysis across multiple frequency bands. By analyzing the spectral characteristics and energy distribution in the frequency dimension, the system achieves more accurate end-point detection in noisy environments while maintaining reasonable processing complexity through efficient spectral processing
Data Source
AI summary
According to an embodiment, disclosed is an electronic device including a speaker, a microphone, a communication interface, a processor operatively connected to the speaker, the microphone, and the communication interface, and a memory operatively connected to the processor. The memory stores instructions that, when executed, cause the processor to receive a first utterance through the microphone, to determine a speaker model by performing speaker recognition on the first utterance, to receive a second utterance through the microphone after the first utterance is received, to detect an end-point of the second utterance, at least partially using the determined speaker model. Besides, various embodiments as understood from the specification are also possible.


