Directional Speech Recognition With Multi-Path Echo Cancellation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing electronic devices struggle to accurately distinguish between multiple audio sources, particularly in noisy environments, leading to challenges in speech recognition and transcription, especially for hearing-impaired users and those with language barriers, and echo cancellation is inadequate.

Innovation Solution

The use of a multi-microphone array with multi-path acoustic echo cancellation and beam forming techniques to differentiate between audio sources, combined with an ASR component trained to recognize and attribute speech in multi-path AEC audio, enhances audio quality and accuracy in speech recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple microphones are used to receive audio from multiple sources, then the ability to distinguish between different audio sources is improved, but the complexity of the system increases

Engineering Contradiction:
Improveaudio source distinction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio processing is segmented into distinct functional components: multi-path AEC for echo removal, beamforming for spatial separation, and ASR for speech recognition. Each component handles a specific aspect of the audio processing pipeline, making the complex system more manageable and effective

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Beamforming acts as an intermediary between the multi-microphone array and the ASR system, transforming raw multi-channel audio into directionally-separated audio streams that the ASR can process more effectively

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If acoustic echo cancellation is applied to remove echoes from audio, then audio quality is improved, but the processing complexity and computational requirements increase

Engineering Contradiction:
Improveaudio qualityVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs preliminary echo cancellation processing on the multi-channel audio before passing it to the beamforming and ASR stages. By removing echoes early in the processing pipeline, subsequent stages work with cleaner audio signals, reducing their computational burden

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The multi-path AEC operates in the multi-channel audio domain, treating echo cancellation as a multi-dimensional problem rather than simple single-channel processing. This allows simultaneous echo removal across multiple audio channels while preserving spatial information

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If beam forming is used to segment audio into sectors, then the ability to distinguish audio sources is improved, but the device complexity increases

Engineering Contradiction:
Improveaudio source attribution accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

Beamforming creates directionally-specific audio streams with enhanced quality for each spatial sector. Each beamformed channel is optimized for capturing sound from a specific direction, providing local quality enhancement that improves source attribution accuracy

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250299678A1Methods, devices, and systems for directional speech recognition with acoustic echo cancellation
Publication Date: 2025.09.25 META PLATFORMS TECHNOLOGIES LLC
  • US20250299678A1 patent drawing
  • US20250299678A1 patent drawing
  • US20250299678A1 patent drawing

AI summary

An example method of providing speech-to-text transcription includes receiving, at an electronic device, multiple channels of audio data from a plurality of microphones, where the multiple channels of audio data comprise speech from a user of the electronic device and speech from one or more other persons. The method also includes generating refined audio data by applying a multi-path acoustic echo cancellation (AEC) technique to the multiple channels of audio data. The method further includes generating directional audio data by applying beamforming to the refined audio data. The method also includes identifying, by inputting the directional audio data to an automatic speech recognizer (ASR), the speech from the user of the electronic device and the speech from the one or more other persons, and generating a textual transcription for the conversation.