Multi-Microphone Speech Dialog System Spatial Zone Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multi-microphone speech systems face challenges in detecting the spatial zone where a speech utterance originates, leading to difficulties in distinguishing between desired and interfering speech components and experiencing latency during transitions between broad and selective listening modes.
Innovation Solution
The system sends spatial zone activity information alongside audio signals to the ASR, allowing it to recognize utterances and perform zone-dedicated speech dialogs by determining the origin of the utterance, and provides multiple audio streams for seamless transition through buffering to resume recognition in the relevant zone.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the SE sends spatial zone activity information to the ASR, then the ASR can identify the spatial zone where speech originates, but the system complexity increases
Solution Approach 1:
The patent introduces an intermediary component that extracts and transmits spatial zone activity information from the SE to the ASR. This intermediary layer decouples the complexity of spatial zone detection from the ASR module, allowing the ASR to benefit from precise zone identification without directly implementing the complex detection algorithms. The SE module continues to handle the complex spatial processing while providing simplified zone activity indicators to the ASR.
2Adaptability or versatility
If the system transitions from broad listening to selective listening mode, then the system can focus on a specific spatial zone, but latency occurs during the transition
Solution Approach 1:
The patent implements preliminary action by pre-processing and buffering audio data from all spatial zones continuously, even before a specific zone is selected for selective listening. This allows the system to have audio data ready from any zone, enabling rapid switching between broad and selective listening modes without the latency that would result from starting processing only after mode transition is initiated.
3Adaptability or versatility
If multiple audio streams are provided for seamless transition, then the transition between listening modes becomes smooth, but the CPU load increases
Solution Approach 1:
The patent applies partial action by maintaining multiple audio streams but selectively processing them based on the current listening mode. During broad listening, all streams are monitored at a lower processing level. When transitioning to selective listening, the system activates full processing for only the relevant zone's stream while reducing or suspending processing for other zones. This partial processing approach enables smooth transitions without sustaining the full CPU load of processing all streams at maximum depth continuously.
Data Source
AI summary
There is provided a speech dialog system that includes a first microphone, a second microphone, a processor and a memory. The first microphone captures first audio from a first spatial zone, and produces a first audio signal. The second microphone captures second audio from a second spatial zone, and produces a second audio signal. The processor receives the first audio signal and the second audio signal, and the memory contains instructions that control the processor to perform operations of a speech enhancement module, an automatic speech recognition module, and a speech dialog module that performs a zone-dedicated speech dialog.


