Videoconference Audio Focus Control via Camera State

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

During videoconferences, the existing technologies face challenges in accurately focusing audio signals on the speaker, leading to unnatural noise perception when noise sources other than the speaker are not properly differentiated from the speaker's direction, and transitioning between monophonic and stereophonic signals does not align with the video camera's focus.

Innovation Solution

A computing system determines the direction of the video camera's focus and generates beamformed signals in the directions of both the speaker and noise sources, preferentially weighting the speaker's signal to create a clear and natural audio experience by transitioning between monophonic and stereophonic signals based on the camera's focus.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If audio signals are focused on the speaker using beamforming, then the speaker's voice clarity is improved, but noise from other directions cannot be properly differentiated and appears to come from the same direction as the speaker

Engineering Contradiction:
Improvespeaker voice clarityVSAvoidnoise direction information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The audio signal processing is segmented into multiple independent beamformed signals, each focused on a different direction (speaker direction and noise source direction). This segmentation allows the system to preserve directional information for both speaker and noise sources separately, rather than combining them into a single monophonic signal that loses spatial cues.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from monophonic (single-channel) audio to stereophonic (multi-channel) audio, adding a spatial dimension to the audio output. This dimensional change enables the reproduction of noise from different directions, preserving the spatial relationship between speaker and noise sources that was lost in traditional beamforming approaches.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If a monophonic signal is transmitted when the video system aims at the speaker, then the speaker's audio is clear, but the audio does not match the visual context when multiple people are shown

Engineering Contradiction:
Improveaudio focus accuracyVSAvoidaudio-visual synchronization
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system dynamically switches between monophonic and stereophonic signal transmission based on the video camera's focus state. When the camera aims at a single speaker, a monophonic signal is transmitted for clear audio focus. When the camera shows multiple people or pans away from the speaker, the system transitions to stereophonic signals to maintain audio-visual consistency and provide spatial context.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses feedback from the video camera's aiming state to control the audio signal processing mode. The camera's focus information serves as feedback that triggers the appropriate audio mode (monophonic or stereophonic), ensuring that the audio output always matches the visual context being displayed to the user.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If beamforming is applied in the speaker's direction only, then the speaker's signal is enhanced, but the natural spatial perception of noise sources is lost

Engineering Contradiction:
Improvespeaker signal enhancementVSAvoidnatural spatial perception
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The audio processing is segmented into multiple beamformed signals pointing in different directions (speaker direction and noise source directions). Each segment preserves the spatial characteristics of its respective source, allowing the system to enhance the speaker signal while simultaneously preserving the natural spatial perception of noise sources through separate directional processing.

Inventive Principle:
Principle #1Segmentation

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This solution ensures that the audio signals are clearly focused on the speaker and noise is perceived as coming from a different direction, mimicking a face-to-face meeting experience, thereby enhancing the naturalness of the audio output during videoconferences.

Implementation Method 1

generating a first beamformed signal based on beamforming, in the first direction, multiple first direction audio signals received by the array of microphones

Methodology Applied
Scientific EffectBeamforming:

Data Source

PatentUS10805575B2Controlling focus of audio signals on speaker during videoconference
Publication Date: 2020.10.13 GOOGLE LLC
  • US10805575B2 patent drawing
  • US10805575B2 patent drawing
  • US10805575B2 patent drawing

AI summary

A non-transitory computer-readable storage medium may include instructions stored thereon. When executed by at least one processor, the instructions may be configured to cause a computing system to determine that a video system is aiming at a single speaker of a plurality of people, receive audio signals from a plurality of microphones, the received audio signals including audio signals generated by the single speaker, based on determining that the video system is aiming at the single speaker, transmit a monophonic signal, the monophonic signal being based on the received audio signals, determine that the video system is not aiming at the single speaker, and based on the determining that the video system is not aiming at the single speaker, transmit a stereophonic signal, the stereophonic signal being based on the received audio signals.