Voice Mixing Strategy via Control Information
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice processing systems face challenges in multi-party voice communication, particularly with increased background noise and output overflow when mixing voices from multiple channels, and struggle to meet high quality requirements for audio/video conferences due to resource constraints.
Innovation Solution
A method and system where voice bit streams and control information are sent to a voice server to determine and implement voice-mixing strategies, allowing for dynamic selection and processing of voice streams, either on the server or terminal, to optimize voice mixing quality and resource usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If voice mixing is performed on the server side, then voice mixing quality is improved, but hardware resource consumption increases significantly
Solution Approach 1:
The voice mixing process is segmented into two parts: voice bit stream transmission (performed by terminals) and voice mixing strategy determination (performed by server). This segmentation allows the server to focus only on control operations while terminals handle data processing, reducing server resource consumption while maintaining mixing quality.
Solution Approach 2:
Voice control information acts as an intermediary between the voice bit streams and the voice mixing process. The server processes this control information to determine mixing strategies, rather than directly processing all voice bit streams, thereby reducing hardware resource consumption while preserving mixing quality.
2Device complexity
If voices from multiple channels are directly added together, then the voice mixing process is simplified, but background noise increases and output overflow occurs
Solution Approach 1:
Voice control information is extracted and processed in advance before the actual voice mixing occurs. This preliminary action allows the server to determine the optimal mixing strategy (including channel selection and mixing coefficients) before combining voice signals, preventing noise and overflow issues while maintaining process simplicity.
Solution Approach 2:
The system changes parameters such as the number of channels participating in mixing and the mixing coefficients dynamically based on voice control information. This allows the mixing process to adapt to different conditions, minimizing noise and overflow while keeping the process relatively simple through automated parameter adjustment.
3Object-affected harmful factors
If a small number of channels are selected for voice mixing, then background noise and output overflow are minimized, but voice mixing quality decreases
Solution Approach 1:
The voice mixing strategy is made dynamic rather than static. The server determines the mixing strategy based on real-time voice control information, allowing the number of channels and mixing coefficients to change dynamically. This enables the system to maintain high mixing quality by selecting optimal channels while minimizing noise and overflow through adaptive parameter adjustment.
Data Source
AI summary
Methods, apparatus, and systems for voice processing are provided herein. An exemplary method can be implemented by a terminal. A voice bit stream to be sent can be obtained. Voice control information corresponding to the voice bit stream to be sent can be obtained. The voice control information can be used for a voice server to determine a voice-mixing strategy. The voice bit stream and the voice control information can be sent to the voice server. At least one voice bit stream, returned by the voice server based on the voice-mixing strategy, can be received. The at least one voice bit stream can be outputted.


