Spectral Audio Gain Adjustment for Off-Axis Speech
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
During videoconferences, audio quality suffers when participants move away from the audio input or become distracted, leading to reduced volume and missed details due to the directional nature of audio inputs like condenser microphones.
Innovation Solution
Adjusting the gain of audio input spectrally for a subset of audible frequencies between 1,000 and 5,000 Hz when the mouth is oriented off-axis relative to the audio input, specifically avoiding adjustments to low and high frequencies to improve speech intelligibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the gain of audio input is adjusted when the mouth is oriented off-axis, then speech intelligibility is improved, but overall volume consistency deteriorates
Solution Approach 1:
The patent applies different gain adjustments to different frequency bands based on the mouth's angular position. Specifically, frequencies between 1-5 kHz receive enhanced gain when the mouth is off-axis, while other frequency ranges maintain their original gain levels. This selective frequency-based adjustment improves speech intelligibility without causing noticeable overall volume changes, resolving the contradiction between intelligibility enhancement and volume consistency.
2Measurement precision
If spectral adjustment is applied only to midrange frequencies (1-5 kHz), then speech intelligibility is enhanced, but frequency response uniformity deteriorates
Solution Approach 1:
The invention selectively adjusts gain only in the midrange frequency band (1-5 kHz) where speech intelligibility is most critical, while leaving other frequency bands unchanged. This localized adjustment improves speech understanding without significantly altering the overall frequency response characteristics, as the human ear is less sensitive to volume changes in non-speech frequency ranges.
Solution Approach 2:
The patent dynamically changes the gain parameter specifically for the 1-5 kHz frequency band based on the detected mouth angle. When the mouth moves off-axis, the system increases gain in this specific frequency range to compensate for reduced acoustic energy capture, while maintaining original gain settings for other frequencies. This targeted parameter modification achieves intelligibility improvement with minimal impact on overall frequency response uniformity.
3Measurement precision
If gain adjustment is made based on mouth orientation detection, then audio quality is improved, but system complexity increases
Solution Approach 1:
The patent replaces complex mechanical or acoustic solutions with an electronic/digital approach. Instead of using multiple physical microphones or complex acoustic waveguides to maintain audio quality for off-axis speakers, the system uses a single audio input combined with image processing to detect mouth orientation, then applies computational gain adjustment. This substitution of mechanical complexity with electronic control achieves audio quality improvement while keeping the physical system relatively simple.
Solution Approach 2:
The invention introduces an intermediary processing stage between audio capture and audio output. The system first captures audio through a standard microphone, then uses image data as an intermediary to determine mouth orientation, and finally applies gain adjustment based on this orientation information. This intermediary approach allows the system to maintain simple hardware while achieving improved audio quality through intelligent signal processing.
Data Source
AI summary
An electronic device includes an imager capturing one or more images of a subject engaging the electronic device and an audio input receiving acoustic signals having audible frequencies from the mouth of the subject engaging the electronic device. One or more processors determine from the one or more images of the subject whether the mouth of the subject is oriented on-axis relative to the audio input or off-axis relative to the audio input. The one or more processors adjust a gain of the audio input associated with a subset of the audible frequencies when the mouth of the subject is oriented off-axis relative to the audio input.


