Multi-microphone Speech Recognition via Concurrent Utterance Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems face challenges in far-field environments due to additive noise and reverberation, leading to high word error rates and unnatural user interactions, especially when users are distant from microphones, as they require users to stand near a single receiver and speak directly into it.

Innovation Solution

The development of acoustically distributed yet functionally centralized speech recognition systems that use multiple microphones to concurrently process and combine different versions of an utterance, employing adaptive weighting and fusion techniques to generate a highest-probability transcription, allowing for natural interaction with devices regardless of user position or orientation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a single microphone receiver is used for speech recognition, then the device complexity is reduced, but the speech recognition reliability deteriorates in far-field environments due to additive noise and reverberation

Engineering Contradiction:
Improvespeech recognition reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The speech recognition system is segmented into multiple independent microphone receivers distributed throughout the room, each capturing the utterance independently. This segmentation allows the system to process multiple versions of the same utterance from different spatial positions, improving reliability in far-field environments by reducing the impact of additive noise and reverberation at any single location.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Multiple versions of the utterance captured by different microphones are merged and processed concurrently by a speech recognition engine. The system combines the information from multiple independent streams to determine the most likely transcription, thereby improving speech recognition reliability through ensemble processing while managing system complexity through centralized recognition logic.

Inventive Principle:
Principle #5Merging (Combining)

2Ease of operation

If multiple microphones are distributed throughout the room, then the ease of operation improves by allowing natural user movement, but the device complexity increases due to multiple receivers and processing requirements

Engineering Contradiction:
Improveuser interaction naturalnessVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

Multiple microphones are distributed throughout the room to create a universal speech recognition system that can handle user interactions from any position or orientation. This multi-functional arrangement allows users to naturally move and speak without needing to position themselves specifically, as any microphone can capture the utterance effectively.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

A centralized speech recognition engine acts as an intermediary that receives and processes versions of utterances from multiple distributed microphones. This intermediary consolidates the complexity of handling multiple independent receivers, managing the concurrent processing and selection of the most likely transcription, thereby reducing the operational complexity exposed to the user.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If conventional speech recognition systems are used in far-field environments, then the device complexity remains low, but the measurement precision deteriorates with word error rates increasing to 60-80%

Engineering Contradiction:
Improvetranscription accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The speech recognition system dynamically processes multiple concurrent versions of the utterance from different microphones, adapting to the varying quality and characteristics of each stream. The recognition engine evaluates multiple transcription candidates and dynamically selects the most likely transcription based on the combined information, thereby improving measurement precision in far-field environments.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameter of processing by concurrently analyzing multiple independent streams of the same utterance rather than processing a single stream. This parameter change allows the system to leverage the diverse information from multiple microphones, improving transcription accuracy by reducing the impact of noise and reverberation that affect individual streams.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9865265B2Multi-microphone speech recognition systems and related techniques
Publication Date: 2018.01.09 APPLE INC
  • US9865265B2 patent drawing
  • US9865265B2 patent drawing
  • US9865265B2 patent drawing

AI summary

A speech recognition system for resolving impaired utterances can have a speech recognition engine configured to receive a plurality of representations of an utterance and concurrently to determine a plurality of highest-likelihood transcription candidates corresponding to each respective representation of the utterance. The recognition system can also have a selector configured to determine a most-likely accurate transcription from among the transcription candidates. As but one example, the plurality of representations of the utterance can be acquired by a microphone array, and beamforming techniques can generate independent streams of the utterance across various look directions using output from the microphone array.