Vertical and Horizontal Microphone Arrays for Echo Suppression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video endpoint designs fail to provide high-quality, hands-free speech pickup in compact integrated systems due to short microphone-loudspeaker distance, leading to degraded Acoustic Echo Cancellation and interference from table reflections and nearby noise sources.

Innovation Solution

A video endpoint equipped with a vertical microphone array and a horizontal microphone array processes audio signals to determine degrees of arrival, adjusting gain to minimize noise and maximize target sound source audio, using beamforming and non-linear suppression techniques to improve Echo-to-Near-end Ratio and frequency response.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Volume of moving object

If conventional video endpoint designs use short microphone-loudspeaker distance for compact integration, then device size is reduced, but acoustic echo cancellation performance degrades and noise interference increases

Engineering Contradiction:
Improvedevice sizeVSAvoidacoustic echo cancellation performance
Core Design Contradiction:
Volume of moving objectVSReliability

Solution Approach 1:

The patent transitions from a conventional single-plane microphone arrangement to a three-dimensional spatial arrangement with vertical and horizontal arrays. This dimensional expansion enables the system to capture audio from multiple directions and depths, creating spatial separation between microphones and loudspeakers even within a compact form factor, thereby maintaining echo cancellation performance while reducing device size

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The microphone system is divided into separate vertical and horizontal arrays, each performing specialized functions. The vertical array primarily captures near-end speech, while the horizontal array captures far-end audio and noise. This segmentation allows independent optimization of each array's position and processing, enabling compact integration while maintaining acoustic performance

Inventive Principle:
Principle #1Segmentation

2Device complexity

If conventional video endpoint designs use single-plane microphone arrangement, then device structure is simplified, but ability to suppress horizontally-displaced noise sources is insufficient

Engineering Contradiction:
Improvemicrophone arrangement structureVSAvoidhorizontally-displaced noise interference
Core Design Contradiction:
Device complexityVSObject-affected harmful factors

Solution Approach 1:

The system adds a vertical dimension to the traditional horizontal microphone array, creating a two-dimensional array structure. This vertical dimension provides the necessary spatial baseline to triangulate and suppress noise sources that are horizontally displaced from the device, as such noise arrives at vertical array elements with distinct time and phase differences that can be exploited for suppression

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If video endpoint uses vertical and horizontal microphone arrays with beamforming, then audio quality and noise suppression are improved, but signal processing complexity increases

Engineering Contradiction:
Improveaudio qualityVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The signal processing is segmented into distinct functional blocks: a near-end processor that handles vertical array signals for target speech enhancement, and a far-end processor that handles horizontal array signals for noise suppression. This segmentation allows each processor to be optimized for its specific function, reducing overall processing complexity while maintaining high audio quality

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different processing strategies are applied to different spatial regions: the vertical array receives aggressive beamforming and gain adjustment for near-end speech, while the horizontal array receives processing optimized for far-end noise suppression. This local quality approach allows tailored processing for each spatial domain, improving audio quality without uniformly increasing complexity across all processing paths

Inventive Principle:
Principle #3Local quality

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The solution enhances audio quality by suppressing unwanted noise and echo, ensuring clear hands-free speech pickup in compact settings, even in open office environments.

Implementation Method 1

obtains, from the vertical microphone array, a first audio signal including audio from a target sound source and audio from a horizontally-displaced sound source

Methodology Applied
Scientific EffectSound wave propagation: Sound

Implementation Method 2

based on the second audio signal and the third audio signal, determines at least one of a first degree of arrival of the audio from the target sound source or a second degree of arrival of the audio from the horizontally-displaced sound source

Methodology Applied
Scientific EffectTime difference of arrival: Time of Flight

Implementation Method 3

using beamforming and non-linear suppression techniques to improve Echo-to-Near-end Ratio and frequency response

Methodology Applied
Scientific EffectBeamforming: Focusing

Implementation Method 4

Conventional video endpoint designs fail to provide high-quality, hands-free speech pickup in compact integrated systems due to short microphone-loudspeaker distance, leading to degraded Acoustic Echo Cancellation

Methodology Applied
Scientific EffectAcoustic echo cancellation: Echo

Data Source

PatentEP4052481B1Audio signal processing based on microphone arrangement
Publication Date: 2024.08.14 CISCO TECHNOLOGY INC
  • EP4052481B1 patent drawingFigure 1
  • EP4052481B1 patent drawingFigure 2A
  • EP4052481B1 patent drawingFigure 2B

AI summary

In one example, a video endpoint obtains, from a vertical microphone array, a first audio signal including audio from a target sound source and audio from a horizontally- displaced sound source. The video endpoint obtains, from a horizontal microphone array, a second audio signal and a third audio signal both including the audio from the target sound source and the audio from the horizontally-displaced sound source. Based on the second audio signal and the third audio signal, the video endpoint determines at least one of a first degree of arrival of the audio from the target sound source or a second degree of arrival of the audio from the horizontally-displaced sound source. Based on the at least one of the first degree of arrival or the second degree of arrival, the video endpoint adjusts a gain of the first audio signal.