Videoconference Audio Focus Control via Camera State
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
During videoconferences, the existing technologies face challenges in accurately focusing audio signals on the speaker, leading to unnatural noise perception when noise sources other than the speaker are not properly differentiated from the speaker's direction, and transitioning between monophonic and stereophonic signals does not align with the video camera's focus.
Innovation Solution
A computing system determines the direction of the video camera's focus and generates beamformed signals in the directions of both the speaker and noise sources, preferentially weighting the speaker's signal to create a clear and natural audio experience by transitioning between monophonic and stereophonic signals based on the camera's focus.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If audio signals are focused on the speaker using beamforming, then the speaker's voice clarity is improved, but noise from other directions cannot be properly differentiated and appears to come from the same direction as the speaker
Solution Approach 1:
The audio signal processing is segmented into multiple independent beamformed signals, each focused on a different direction (speaker direction and noise source direction). This segmentation allows the system to preserve directional information for both speaker and noise sources separately, rather than combining them into a single monophonic signal that loses spatial cues.
Solution Approach 2:
The system transitions from monophonic (single-channel) audio to stereophonic (multi-channel) audio, adding a spatial dimension to the audio output. This dimensional change enables the reproduction of noise from different directions, preserving the spatial relationship between speaker and noise sources that was lost in traditional beamforming approaches.
2Measurement precision
If a monophonic signal is transmitted when the video system aims at the speaker, then the speaker's audio is clear, but the audio does not match the visual context when multiple people are shown
Solution Approach 1:
The system dynamically switches between monophonic and stereophonic signal transmission based on the video camera's focus state. When the camera aims at a single speaker, a monophonic signal is transmitted for clear audio focus. When the camera shows multiple people or pans away from the speaker, the system transitions to stereophonic signals to maintain audio-visual consistency and provide spatial context.
Solution Approach 2:
The system uses feedback from the video camera's aiming state to control the audio signal processing mode. The camera's focus information serves as feedback that triggers the appropriate audio mode (monophonic or stereophonic), ensuring that the audio output always matches the visual context being displayed to the user.
3Measurement precision
If beamforming is applied in the speaker's direction only, then the speaker's signal is enhanced, but the natural spatial perception of noise sources is lost
Solution Approach 1:
The audio processing is segmented into multiple beamformed signals pointing in different directions (speaker direction and noise source directions). Each segment preserves the spatial characteristics of its respective source, allowing the system to enhance the speaker signal while simultaneously preserving the natural spatial perception of noise sources through separate directional processing.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This solution ensures that the audio signals are clearly focused on the speaker and noise is perceived as coming from a different direction, mimicking a face-to-face meeting experience, thereby enhancing the naturalness of the audio output during videoconferences.
Implementation Method 1
generating a first beamformed signal based on beamforming, in the first direction, multiple first direction audio signals received by the array of microphones
Data Source
AI summary
A non-transitory computer-readable storage medium may include instructions stored thereon. When executed by at least one processor, the instructions may be configured to cause a computing system to determine that a video system is aiming at a single speaker of a plurality of people, receive audio signals from a plurality of microphones, the received audio signals including audio signals generated by the single speaker, based on determining that the video system is aiming at the single speaker, transmit a monophonic signal, the monophonic signal being based on the received audio signals, determine that the video system is not aiming at the single speaker, and based on the determining that the video system is not aiming at the single speaker, transmit a stereophonic signal, the stereophonic signal being based on the received audio signals.


