Smart Audio Depth-Based Microphone Tuning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In virtual communication sessions, changes in the distance of a speaking user to a communication device can lead to reduced speech quality of audio data, as existing systems fail to effectively adapt microphone settings based on the depth of the audio source within the camera's field of view.

Innovation Solution

A communication system that optimizes audio reception by selecting an appropriate audio mode (near-field or far-field) based on the depth of the audio source relative to the device, adjusting tuning parameters such as automatic gain control, noise suppression, and echo cancellation to enhance sound quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If the user moves closer to or farther from the communication device during a virtual communication session, then the audio capture coverage is maintained, but the speech quality of audio data is reduced

Engineering Contradiction:
Improveaudio capture coverageVSAvoidspeech quality
Core Design Contradiction:
Area of stationary objectVSManufacturing precision

Solution Approach 1:

The system dynamically adjusts microphone tuning parameters (such as gain, noise suppression, and beamforming) based on the real-time depth of the audio source detected by the camera. This allows the audio capture characteristics to change adaptively as the user moves closer or farther from the device, maintaining optimal speech quality across varying distances while preserving audio capture coverage.

Inventive Principle:
Principle #15Dynamics

2Manufacturing precision

If the microphone tuning parameters are adjusted to optimize speech quality for a specific distance, then the speech quality is improved, but the system cannot adapt to users at different distances

Engineering Contradiction:
Improvespeech qualityVSAvoiddistance adaptability
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The system employs a feedback mechanism where the camera continuously tracks the depth of the audio source, and this depth information is fed back to adjust the microphone tuning parameters in real-time. This closed-loop control enables the system to automatically adapt to users at different distances, optimizing speech quality for each specific distance without requiring manual reconfiguration.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system changes multiple microphone tuning parameters simultaneously based on the detected audio source depth, including gain adjustments, noise suppression levels, and beamforming patterns. By coordinating changes across multiple parameters, the system maintains optimal speech quality across a wide range of distances, enhancing both speech quality and distance adaptability.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11202148B1Smart audio with user input
Publication Date: 2021.12.14 META PLATFORMS INC
  • US11202148B1 patent drawing
  • US11202148B1 patent drawing
  • US11202148B1 patent drawing

AI summary

A communication system establishes a communication session with a communication device via a network and captures, via one or more cameras, a local area within a field of view of the one or more cameras. The local area includes one or more audio sources. The communication system receives, from the communication device, a selection of a respective audio source from the one or more audio source in the local area. The selection indicates that audio originating from the respective audio source is to be prioritized over other audio sources in the local area. The communication system tunes one or more microphones based on a depth of the respective audio source relative to the communication system to optimize reception of audio originating at the respective audio source by the one or more microphones.