Multi-Microphone Speech Dialog System Spatial Zone Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multi-microphone speech systems face challenges in detecting the spatial zone where a speech utterance originates, leading to difficulties in distinguishing between desired and interfering speech components and experiencing latency during transitions between broad and selective listening modes.

Innovation Solution

The system sends spatial zone activity information alongside audio signals to the ASR, allowing it to recognize utterances and perform zone-dedicated speech dialogs by determining the origin of the utterance, and provides multiple audio streams for seamless transition through buffering to resume recognition in the relevant zone.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the SE sends spatial zone activity information to the ASR, then the ASR can identify the spatial zone where speech originates, but the system complexity increases

Engineering Contradiction:
Improvespatial zone detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary component that extracts and transmits spatial zone activity information from the SE to the ASR. This intermediary layer decouples the complexity of spatial zone detection from the ASR module, allowing the ASR to benefit from precise zone identification without directly implementing the complex detection algorithms. The SE module continues to handle the complex spatial processing while providing simplified zone activity indicators to the ASR.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If the system transitions from broad listening to selective listening mode, then the system can focus on a specific spatial zone, but latency occurs during the transition

Engineering Contradiction:
Improvelistening mode flexibilityVSAvoidtransition latency
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent implements preliminary action by pre-processing and buffering audio data from all spatial zones continuously, even before a specific zone is selected for selective listening. This allows the system to have audio data ready from any zone, enabling rapid switching between broad and selective listening modes without the latency that would result from starting processing only after mode transition is initiated.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If multiple audio streams are provided for seamless transition, then the transition between listening modes becomes smooth, but the CPU load increases

Engineering Contradiction:
Improvemode transition smoothnessVSAvoidCPU load
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent applies partial action by maintaining multiple audio streams but selectively processing them based on the current listening mode. During broad listening, all streams are monitored at a lower processing level. When transitioning to selective listening, the system activates full processing for only the relevant zone's stream while reducing or suspending processing for other zones. This partial processing approach enables smooth transitions without sustaining the full CPU load of processing all streams at maximum depth continuously.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11367437B2Multi-microphone speech dialog system for multiple spatial zones
Publication Date: 2022.06.21 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11367437B2 patent drawing
  • US11367437B2 patent drawing
  • US11367437B2 patent drawing

AI summary

There is provided a speech dialog system that includes a first microphone, a second microphone, a processor and a memory. The first microphone captures first audio from a first spatial zone, and produces a first audio signal. The second microphone captures second audio from a second spatial zone, and produces a second audio signal. The processor receives the first audio signal and the second audio signal, and the memory contains instructions that control the processor to perform operations of a speech enhancement module, an automatic speech recognition module, and a speech dialog module that performs a zone-dedicated speech dialog.