Multi-microphone Speech Recognition with Beamforming and Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems face challenges in far-field environments due to additive noise and reverberation, leading to high word error rates (WER) and unnatural user interactions, as they require users to stand near a single microphone, limiting movement and causing errors in device control.

Innovation Solution

The implementation of acoustically distributed yet functionally centralized speech recognition systems that use multiple microphones to concurrently process and combine different versions of an utterance, employing adaptive weighting and deep-neural-networks to improve transcription accuracy and reduce WER, allowing users to move freely while maintaining accurate device control.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If multiple microphones are used to capture far-field speech, then the coverage area and user mobility are improved, but the speech recognition accuracy deteriorates due to additive noise and reverberation

Engineering Contradiction:
Improvecoverage areaVSAvoidspeech recognition accuracy
Core Design Contradiction:
Area of stationary objectVSReliability

Solution Approach 1:

The patent segments the speech signal processing by creating multiple independent recognition streams from different microphone recordings. Each stream processes the speech signal separately through its own acoustic feature extraction and recognition engine, allowing the system to handle far-field conditions more effectively by dividing the processing task across multiple parallel pathways rather than relying on a single processed signal.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges the results from multiple independent recognition streams through a selection mechanism. The selector component combines the transcription candidates from different streams and selects the best transcription, effectively merging the strengths of multiple microphones and processing paths to overcome the weaknesses of individual far-field recordings.

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If users speak towards a single near-field receiver, then the speech recognition accuracy is improved, but the user mobility and natural interaction are restricted

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoiduser mobility
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent creates a universal speech recognition system that functions effectively in both near-field and far-field conditions. By deploying multiple microphones throughout the environment and processing each recording independently, the system becomes universally responsive to speech from various positions and orientations, eliminating the need for users to adopt specific speaking postures or positions.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent transitions from a single-point (0D) or line (1D) recognition approach to a distributed spatial (3D) arrangement of microphones. This dimensional expansion allows the system to capture speech from any location in the room, converting the limitation of single-position recognition into a multi-position spatial recognition capability.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Adaptability or versatility

If independent devices operate autonomously with their own speech recognition, then the device autonomy is improved, but the misinterpretation of user intent increases

Engineering Contradiction:
Improvedevice autonomyVSAvoiduser intent accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces a central hub as an intermediary that coordinates speech recognition across multiple devices. The hub receives recordings from various devices, manages the independent recognition streams, and selects the appropriate transcription and target device. This intermediary layer maintains device autonomy while preventing misinterpretation by centralizing the decision-making process for intent recognition.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10614812B2Multi-microphone speech recognition systems and related techniques
Publication Date: 2020.04.07 APPLE INC
  • US10614812B2 patent drawing
  • US10614812B2 patent drawing
  • US10614812B2 patent drawing

AI summary

A speech recognition system for resolving impaired utterances can have a speech recognition engine configured to receive a plurality of representations of an utterance and concurrently to determine a plurality of highest-likelihood transcription candidates corresponding to each respective representation of the utterance. The recognition system can also have a selector configured to determine a most-likely accurate transcription from among the transcription candidates. As but one example, the plurality of representations of the utterance can be acquired by a microphone array, and beamforming techniques can generate independent streams of the utterance across various look directions using output from the microphone array.