Dynamic Call Audio Mixing via Voice Activity Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional call audio mixing methods often result in unsatisfactory voice mixing due to interference from ambient noise, leading to poor call quality, as they are inflexible in route selection and may exclude participants with low sound recording and collecting volume, causing them to be unheard.
Innovation Solution
A call audio mixing processing method that performs voice analysis on call audios to determine voice activity, adjusts audio volumes based on this activity, and mixes the adjusted audios to reduce interference, allowing all participants to contribute effectively without limiting the number of participants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If selective mixing manner is used to reduce interference from ambient noise, then interference on the speaker is reduced, but participants with low sound recording and collecting volume are excluded from mixing, resulting in low call quality
Solution Approach 1:
The patent implements dynamic route selection that adapts to real-time call conditions. The mixing module dynamically determines whether to mix audio from each participant based on voice activity detection and audio quality assessment, rather than using a fixed selective mixing policy. This allows the system to flexibly include or exclude participants based on current call state, resolving the contradiction between reducing noise interference and maintaining call quality.
Solution Approach 2:
The system changes the parameter of audio routing decisions based on multiple factors including voice activity level, signal-to-noise ratio, and participant count. By dynamically adjusting these parameters, the system can optimize the balance between reducing ambient noise interference and ensuring all relevant participants are heard, thereby improving call quality without excessive interference.
2Reliability
If full mixing mode is used to include all participants, then all voices can be heard, but interference from ambient noise increases, resulting in unsatisfactory voice mixing effect
Solution Approach 1:
The patent applies different mixing strategies to different participants based on their individual audio characteristics. Rather than uniformly mixing all audio or excluding certain participants, the system performs local optimization by analyzing each participant's voice activity and audio quality separately, then applying appropriate mixing decisions to each. This localized approach allows the system to include participants with good signal quality while excluding those with high ambient noise, resolving the contradiction between comprehensive coverage and interference reduction.
Solution Approach 2:
The system implements feedback mechanisms where the mixing module continuously monitors audio quality metrics and voice activity levels, then adjusts mixing decisions accordingly. This closed-loop control allows the system to respond to changing call conditions in real-time, optimizing the balance between including all participants and reducing ambient noise interference through continuous adaptation.
3Object-affected harmful factors
If route selection is based on volume levels, then some sounds with low volume are excluded to reduce interference, but participants with high background noise are more likely to be selected, resulting in low call quality
Solution Approach 1:
The patent introduces an intermediary evaluation mechanism that assesses multiple factors beyond just volume levels, including voice activity detection results, signal-to-noise ratio estimates, and participant context. This intermediary layer between raw audio input and mixing decision prevents the system from making suboptimal routing decisions based solely on volume, allowing it to correctly identify and exclude participants with high background noise while maintaining those with low volume but clear speech.
Data Source
Figure 1~2
Figure 3
Figure 4~5
AI summary
A call audio mixing processing method is executed by a server. Said method comprises: acquiring call audios sent by terminals of call members participating in a call; performing voice analysis on each of the call audios to determine voice popularity corresponding to each of the call member terminals, the voice popularity being used to reflect the popularity of the call member participating in the call; determining, according to the voice popularity, voice adjustment parameters corresponding to the call member terminals respectively; adjusting, according to the voice adjustment parameters respectively corresponding to the call member terminals, corresponding call audios to obtain adjusted audios, and performing mixing processing on the adjusted audios to obtain a mixed audio.