Directional Beamforming for Speech End-Pointing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automatic speech recognition (ASR) devices face challenges in accurately isolating user speech from background noise and determining the end of a spoken utterance due to interference from other audio sources, such as appliances and multiple speakers.

Innovation Solution

The use of beamforming techniques with circular or linear microphone arrays, combined with signal processing and filtering methods, allows for the isolation of desired audio by focusing on specific directions and filtering out unwanted noise, enabling effective end-pointing to identify the termination of user speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional speech recognition processes all audio input, then complete speech data is captured, but background noise and interference from other audio sources degrade recognition accuracy

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidbackground noise interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent applies local quality by creating directionally selective audio processing zones. Different spatial regions around the device are assigned different processing characteristics through beamforming, allowing the system to focus computational resources and attention on speech sources from specific directions while naturally filtering out noise from other directions. This spatially varying quality approach resolves the contradiction by making the system selectively sensitive to desired speech while rejecting background noise.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent introduces beamforming as an intermediary processing layer between the raw audio input and the speech recognition engine. This intermediary component analyzes the spatial characteristics of incoming audio signals and selectively enhances or suppresses signals based on their direction of origin. By placing this intermediary layer, the system can capture complete speech data from target directions while filtering out background noise, thus resolving the accuracy-noise interference contradiction.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of time

If the system processes audio from all directions equally, then no speech is missed, but the system cannot accurately determine when user speech ends due to continuous audio input

Engineering Contradiction:
Improveutterance end detection accuracyVSAvoidcontinuous background audio
Core Design Contradiction:
Loss of timeVSObject-affected harmful factors

Solution Approach 1:

The patent applies local quality by creating directionally selective audio processing zones. Different spatial regions around the device are assigned different processing characteristics through beamforming, allowing the system to focus computational resources and attention on speech sources from specific directions while naturally filtering out noise from other directions. This spatially varying quality approach resolves the contradiction by making the system selectively sensitive to desired speech while rejecting background noise.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent extracts the relevant speech signal from the mixed audio input by using beamforming to isolate sounds from specific directions. By separating the desired speech signal from the continuous background audio based on spatial characteristics, the system can accurately detect utterance boundaries without being confused by ongoing noise from other sources. This extraction approach allows precise end-pointing by focusing only on the isolated speech component.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If multiple microphones are used to capture audio from all directions, then comprehensive audio coverage is achieved, but the complexity of isolating user speech increases

Engineering Contradiction:
Improvespeech isolation reliabilityVSAvoidaudio processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent changes the spatial parameter characteristics of the audio processing system by implementing beamforming with multiple microphones. Instead of processing all audio equally, the system dynamically adjusts the weighting and phase relationships of signals from different microphones based on the desired direction of interest. This parameter change approach transforms the complex multi-directional audio mixture into a simplified directional signal stream, improving speech isolation reliability while managing processing complexity through structured spatial filtering.

Inventive Principle:
Principle #35Parameter changes

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enhances the accuracy of speech recognition by isolating user speech from background noise and accurately determining the end of an utterance, even in complex audio environments with multiple noise sources, improving overall human-computer interaction.

Implementation Method 1

The use of beamforming techniques with circular or linear microphone arrays, combined with signal processing and filtering methods, allows for the isolation of desired audio by focusing on specific directions

Methodology Applied
Scientific EffectBeamforming:

Data Source

PatentUS11978478B2Direction based end-pointing for speech recognition
Publication Date: 2024.05.07 AMAZON TECH INC
  • US11978478B2 patent drawing
  • US11978478B2 patent drawing
  • US11978478B2 patent drawing

AI summary

A speech recognition system utilizing automatic speech recognition techniques such as end-pointing techniques in conjunction with beamforming and/or signal processing to isolate speech from one or more speaking users from multiple received audio signals and to detect the beginning and/or end of the speech based at least in part on the isolation. Audio capture devices such as microphones may be arranged in a beamforming array to receive the multiple audio signals. Multiple audio sources including speech may be identified in different beams and processed.