Call Downlink Audio Limiting for Consistent Earphone Volume
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
During phone calls using earphones, near-end users often face challenges in maintaining consistent volume due to variations in the far-end speaker's voice level, leading to poor call experiences.
Innovation Solution
An audio limiting method and system that utilize an audio limiting module in the near-end device to receive far-end sound, perform voice activity detection, and adjust the volume of each audio frame to a target level, ensuring consistent playback volume while suppressing noise during non-speech segments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If far-end sound is played directly without volume control, then the original voice characteristics are preserved, but the near-end user experiences volume surges and inconsistent volume when the far-end speaker changes volume level
Solution Approach 1:
The system performs voice activity detection and calculates gain values in advance for each audio frame before playback. By pre-processing the audio frames to determine appropriate gain values based on voice presence and volume levels, the system prepares volume adjustment parameters ahead of time, ensuring consistent volume delivery without adding complex real-time processing during playback.
Solution Approach 2:
The system dynamically adjusts the gain applied to each audio frame based on voice activity detection results and current volume levels. The gain value changes adaptively frame by frame, allowing the system to respond to varying voice characteristics while maintaining overall volume consistency for the near-end user.
2Measurement precision
If volume is increased to make soft speech audible, then speech intelligibility improves, but background noise and sudden volume surges also increase
Solution Approach 1:
The system applies different gain values to different audio frames based on local voice activity characteristics. Frames containing speech receive appropriate gain to ensure intelligibility, while frames containing only background noise receive minimal or no gain. This localized processing ensures that noise is amplified only when necessary for speech intelligibility, not continuously.
Solution Approach 2:
The system performs voice activity detection on each audio frame and uses the detection results to determine the appropriate gain value. This feedback mechanism allows the system to continuously monitor the audio content and adjust volume accordingly, ensuring speech remains intelligible while minimizing background noise amplification when no speech is present.
3Manufacturing precision
If voice activity detection is performed on each audio frame, then volume adjustment accuracy improves, but processing time and computational load increase
Solution Approach 1:
The system divides the audio stream into discrete frames and performs voice activity detection and gain calculation on each frame independently. This segmentation allows for efficient parallel processing and reduces the computational complexity compared to analyzing the entire audio stream as a single unit, while still achieving precise frame-level volume adjustment.
Solution Approach 2:
The system performs voice activity detection on each audio frame, which may result in some computational redundancy. However, this excessive action ensures that no speech segment is missed and provides robust volume control. The computational overhead is acceptable given the improved precision in volume adjustment and the ability to handle varying voice characteristics effectively.
Data Source
AI summary
A system for audio limiting in a call downlink algorithm during a call is provided, the audio limiter enables automatic amplification of the low pitch and suppression of the high pitch on the earphone side and prevents distortion of the high pitch, so that the audio limiter can ensure that an earphone wearer can clearly understand the voice content and maintain a consistent volume level regardless of whether the sound of a far-end speaker is small or large.


