Microphone Array Beamforming for Video Conference Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In video conferencing, participants in larger group settings often struggle to position microphones optimally for clear audio, leading to difficulties in differentiating between multiple speakers and dealing with ambient noise, which degrades the overall audio quality.

Innovation Solution

Participants can interact with the video conferencing application to select a region of interest within a video stream, allowing the client device to adjust the microphone or microphone array to focus on that area, using techniques like beamforming to improve audio capture.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If multiple microphones are used to capture audio from different speakers in a group setting, then the coverage area is improved, but the ability to differentiate between multiple speakers and filter ambient noise deteriorates

Engineering Contradiction:
Improveaudio capture coverage areaVSAvoidaudio differentiation quality
Core Design Contradiction:
Area of stationary objectVSLoss of information

Solution Approach 1:

The patent segments the audio capture function by dividing the microphone array into multiple independent microphone units, each capable of being independently controlled to focus on different spatial regions. This allows the system to capture audio from multiple speakers simultaneously while maintaining the ability to differentiate between them by directing specific microphones to specific speakers.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by enabling different microphones to have different directional characteristics and focus patterns. Each microphone can be optimized to capture audio from a specific direction or region, allowing the system to maintain high audio differentiation quality across the entire coverage area by assigning specialized capture characteristics to each microphone.

Inventive Principle:
Principle #3Local quality

2Loss of information

If microphones are positioned to focus on specific speakers, then audio differentiation between speakers is improved, but the coverage area and ability to capture ambient audio deteriorates

Engineering Contradiction:
Improveaudio differentiation qualityVSAvoidaudio capture coverage area
Core Design Contradiction:
Loss of informationVSArea of stationary object

Solution Approach 1:

The patent implements dynamics by making the microphone focus directions adjustable and changeable in real-time. The system can dynamically reposition the directional patterns of microphones to follow different speakers as they move or become the active speaker, while maintaining comprehensive coverage through coordinated adjustment of multiple microphones.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies universality by designing the microphone array system to perform multiple functions simultaneously: it can capture audio from multiple specific speakers individually, capture ambient audio from the overall environment, and adapt to different meeting scenarios. This multi-functionality allows the system to maintain both focused audio differentiation and broad coverage.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Device complexity

If a single microphone captures all audio in a room, then the device complexity is reduced, but the audio quality and ability to filter ambient noise deteriorates

Engineering Contradiction:
Improvemicrophone system complexityVSAvoidambient noise interference
Core Design Contradiction:
Device complexityVSObject-affected harmful factors

Solution Approach 1:

The patent segments the audio capture function by dividing the microphone array into multiple independent microphone units, each capable of being independently controlled to focus on different spatial regions. This allows the system to capture audio from multiple speakers simultaneously while maintaining the ability to differentiate between them by directing specific microphones to specific speakers.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent converts the potentially harmful effect of ambient noise into a benefit by using the spatial distribution of multiple microphones to distinguish between desired speech signals and unwanted ambient noise. Through beamforming and spatial filtering, the system can identify and suppress ambient noise while preserving speaker audio, effectively turning the presence of multiple sound sources into an opportunity for improved noise rejection.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enhances audio quality by allowing participants to clearly hear each other, even in noisy environments, by physically adjusting or beamforming microphones to target specific regions within the room.

Implementation Method 1

using techniques like beamforming to improve audio capture

Methodology Applied
Scientific EffectBeamforming:

Data Source

PatentUS11729354B2Remotely adjusting audio capture during video conferences
Publication Date: 2023.08.15 ZOOM VIDEO COMM INC
  • US11729354B2 patent drawing
  • US11729354B2 patent drawing
  • US11729354B2 patent drawing

AI summary

One example method includes joining, by a first client device, a videoconferencing meeting hosted by a video conference provider, the videoconference meeting including a plurality of participants; providing an audio stream and a video stream to a video conference provider; receiving, from a second client device, an audio focus area associated with a video stream provided the first client device; determining, based on the audio focus area, a bounding region within an environment shown in the video stream; directing a microphone array to capture audio from the bounding region; and providing the captured audio as an audio stream to the video conference provider.