Dynamic Audio Channel Segmentation for Remote Conferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current videotelephony conference software merges audio and video into a single channel, making it difficult for participants to localize sound, separate sounds from different speakers, and conduct private conversations without leaving the main conference room.
Innovation Solution
The system separates audio and video streams into different channels, allowing participants to simulate in-person conversations by positioning speakers in specific regions of the interface and dynamically adjusting audio outputs based on speaker positions, enabling private conversations within the main conference room.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If audio from all conference participants is merged into a single audio stream, then system efficiency is improved and target quality of service is achieved, but participants cannot easily determine direction of sound or separate sounds from different speakers
Solution Approach 1:
The patent segments the single merged audio stream into multiple separate audio channels, each corresponding to a specific participant or spatial region. This allows participants to selectively attend to different speakers while maintaining system efficiency through structured channel management rather than complete merging.
Solution Approach 2:
The patent applies different audio processing characteristics to different spatial regions or participant channels. Each region can have customized audio properties (volume, filtering, spatial positioning) that match the local conversation context, enabling natural sound localization while maintaining overall system efficiency.
2Productivity
If video from all conference participants is merged into a single video channel, then system efficiency is improved, but participants cannot easily determine direction of sight or separate visual information from different speakers
Solution Approach 1:
The patent segments the single merged video stream into multiple separate video channels, each associated with a specific participant or spatial region. This enables participants to visually locate and separate different speakers while maintaining system efficiency through organized channel structure.
Solution Approach 2:
The patent applies different video processing characteristics to different spatial regions or participant channels, allowing customized visual properties for each region that enhance the ability to distinguish and localize different speakers while maintaining overall system efficiency.
3Productivity
If all audio and video are merged into a single data transmission channel, then network optimization is achieved, but participants cannot conduct private conversations without leaving the main conference room
Solution Approach 1:
The patent segments the single data transmission channel into multiple logical channels carrying different audio and video streams. This enables the system to optimize the overall network transmission while allowing participants to selectively combine or separate channels to conduct private conversations within the main conference room.
Solution Approach 2:
The patent dynamically configures channel combinations based on participant needs, allowing flexible formation of private conversation groups while maintaining the optimized single physical transmission channel to the network. Participants can dynamically switch between different channel configurations without leaving the main conference.
4Device complexity
If uniform audio filters are applied to all conference audio, then processing simplicity is maintained, but participants cannot efficiently separate sounds from different simultaneous speakers
Solution Approach 1:
The patent segments the uniform audio filtering approach into region-specific or participant-specific filtering channels. Each channel can apply appropriate filtering characteristics while maintaining overall processing simplicity through systematic channel management rather than complex individual processing.
Data Source
AI summary
Apparatus and methods for enhancing a videotelephony conference experience by generating dynamic audio channels. Audio outputs may be provided to listeners over different channels. The multiple channels may simulate live, in-person conversation with conference participants. For example, listeners may also conduct separate, private conversations with other participants of the conference without leaving the general conference conversation. The videotelephony conference interface may coordinate presentation of participants to reflect the audio channels provided to a listener. For example, actively speaking participants may be positioned in different regions of the interface.


