Voice Input Processing Using Speaker Model Endpoint Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Electronic devices face challenges in accurately processing voice data due to noise interference from environmental sounds and other voices, which reduces preprocessing and recognition efficiency.

Innovation Solution

An electronic device is equipped with a processor, memory, microphone, and communication interface, using speaker recognition to determine a speaker model and detect the end-point of a user's utterance, thereby isolating and processing the user's voice data effectively.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If the electronic device receives voice data in a noisy environment, then the device can capture more audio information, but the recognition accuracy decreases due to noise interference

Engineering Contradiction:
Improveaudio informationVSAvoidrecognition accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent segments the audio signal into multiple frequency bands using a filter bank structure. Each band processes audio information independently, allowing the system to capture comprehensive audio data while identifying and suppressing noise in specific frequency ranges, thus maintaining recognition accuracy despite noisy environmental conditions

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing characteristics to different frequency bands. By assigning unique filter properties and processing parameters to each frequency band, the system optimizes noise suppression and voice recognition for specific spectral regions while preserving overall audio information quality

Inventive Principle:
Principle #3Local quality

2Loss of information

If the electronic device processes all received audio data, then complete voice information is captured, but preprocessing efficiency decreases due to noise data

Engineering Contradiction:
Improvevoice information completenessVSAvoidpreprocessing efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent extracts and isolates voice signals from the mixed audio stream by analyzing spectral characteristics across frequency bands. The system identifies and extracts relevant voice information while separating and discarding noise components, thereby maintaining complete voice information capture while improving preprocessing efficiency through selective data processing

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary noise suppression and voice detection in the frequency domain before final voice recognition processing. By pre-processing the audio signal to remove obvious noise components and identify voice segments, the system reduces the computational burden on subsequent recognition stages, thereby improving overall preprocessing efficiency without losing voice information

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If the electronic device uses traditional end-point detection without speaker recognition, then the processing is simpler, but the detection accuracy decreases in noisy environments

Engineering Contradiction:
Improveprocessing simplicityVSAvoidend-point detection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent transitions from traditional time-domain end-point detection to a frequency-domain approach using spectral analysis across multiple frequency bands. By analyzing the spectral characteristics and energy distribution in the frequency dimension, the system achieves more accurate end-point detection in noisy environments while maintaining reasonable processing complexity through efficient spectral processing

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11514890B2Method for user voice input processing and electronic device supporting same
Publication Date: 2022.11.29 SAMSUNG ELECTRONICS CO LTD
  • US11514890B2 patent drawing
  • US11514890B2 patent drawing
  • US11514890B2 patent drawing

AI summary

According to an embodiment, disclosed is an electronic device including a speaker, a microphone, a communication interface, a processor operatively connected to the speaker, the microphone, and the communication interface, and a memory operatively connected to the processor. The memory stores instructions that, when executed, cause the processor to receive a first utterance through the microphone, to determine a speaker model by performing speaker recognition on the first utterance, to receive a second utterance through the microphone after the first utterance is received, to detect an end-point of the second utterance, at least partially using the determined speaker model. Besides, various embodiments as understood from the specification are also possible.