Multi-Directional Speech Recognition via Beamformer Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face challenges in accurately interpreting user speech in noisy environments, particularly when background noise from multiple directions interferes with the audio signal, making it difficult to isolate and process the desired speech effectively.

Innovation Solution

The implementation of a method using a beamformer and microphone array to process multiple channels of audio simultaneously, allowing the system to isolate and recognize speech from specific directions by performing speech recognition on each channel, thereby improving the accuracy of speech recognition results.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional speech recognition processes audio from all directions equally, then the system is simple to implement, but background noise from multiple directions degrades recognition accuracy

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidbackground noise interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent segments the audio processing by creating multiple independent decoding channels, each focused on a specific spatial direction. The audio signal is divided into direction-specific streams that are processed separately, allowing the system to isolate and enhance speech from particular directions while filtering out noise from other directions. This segmentation resolves the contradiction by organizing processing to improve accuracy without overwhelming complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by optimizing processing for each spatial direction independently. Each decoding channel is tuned to capture speech characteristics from its specific direction, applying direction-aware acoustic models and language models. This allows the system to maintain high recognition accuracy for speech from target directions while naturally suppressing noise from other directions, resolving the accuracy-noise interference contradiction.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If the system processes multiple channels of audio simultaneously, then speech isolation from specific directions improves, but computational complexity increases

Engineering Contradiction:
Improvespeech isolation accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent manages complexity by segmenting the multi-channel processing into independent, parallel decoding channels. Each channel processes a specific direction independently, avoiding the need for complex inter-channel interactions. This segmentation allows the system to achieve superior speech isolation through multi-channel processing while keeping each individual processing path relatively simple and manageable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a spatial dimension to traditional speech recognition by incorporating direction-aware processing. Instead of processing audio as a single monolithic stream, the system organizes processing along the spatial dimension, creating multiple decoding channels corresponding to different directions. This dimensional organization improves speech isolation accuracy while providing a structured framework that manages computational complexity through spatial decomposition.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If direction-specific decoding channels are implemented, then speech from specific directions is enhanced, but the system requires more complex architecture

Engineering Contradiction:
Improvespeech recognition reliabilityVSAvoiddecoder architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent improves reliability by segmenting the decoding process into direction-specific channels, each optimized for speech from its particular direction. This segmentation ensures that speech from target directions is consistently recognized with high reliability, as each channel is专门 tuned for its direction. The modular segmented architecture manages complexity by making each channel independent and manageable while collectively achieving superior overall reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent achieves high recognition reliability through local quality optimization in each decoding channel. Each channel employs direction-aware acoustic models and language models tailored to its specific spatial direction, ensuring optimal recognition performance for speech from that direction. This localized optimization across multiple channels collectively delivers high overall reliability while maintaining manageable architectural complexity through the modular channel structure.

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP3050052B1Speech recognizer with multi-directional decoding
Publication Date: 2018.11.07 AMAZON TECH INC
  • EP3050052B1 patent drawingFigure 1
  • EP3050052B1 patent drawingFigure 2
  • EP3050052B1 patent drawingFigure 3

AI summary

In an automatic speech recognition (ASR) processing system, ASR processing may be configured to process speech based on multiple channels of audio received from a beamformer. The ASR processing system may include a microphone array and the beamformer to output multiple channels of audio such that each channel isolates audio in a particular direction. The multichannel audio signals may include spoken utterances/speech from one or more speakers as well as undesired audio, such as noise from a household appliance. The ASR device may simultaneously perform speech recognition on the multi-channel audio to provide more accurate speech recognition results.