Spatial Audio Multiplexing for Seamless Conference Continuity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Audio conference systems face challenges in maintaining perceptual continuity and reducing background noise when multiple endpoints with different audio capabilities, such as monophonic and soundfield terminals, are connected, leading to undesirable noise and spatial scene complexity.
Innovation Solution
A method for multiplexing continuous input audio signals from multiple endpoints, prioritizing the most recent active soundfield and adjusting gains to minimize background noise, while maintaining spatial presence, by determining talk activity and applying time-dependent gains to ensure seamless transitions and reduced switching between soundfields.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple soundfield endpoints are simultaneously mixed in the audio conference system, then the spatial presence and immersion are improved, but the background noise and spatial scene complexity increase
Solution Approach 1:
The patent extracts and processes individual soundfield signals from multiple endpoints separately before selective mixing. Each soundfield is analyzed for talker presence and spatial characteristics, allowing the system to extract only the necessary spatial information from active endpoints while filtering out background noise from inactive ones.
Solution Approach 2:
The mixing configuration dynamically adjusts based on the activity state of each endpoint. The system continuously monitors talker presence and spatial characteristics, reconfiguring the mix in real-time to include only actively speaking endpoints, thereby maintaining spatial presence while minimizing background noise accumulation.
2Adaptability or versatility
If multiple soundfield endpoints are simultaneously mixed, then the spatial immersion is improved, but the spatial scene complexity increases
Solution Approach 1:
The patent extracts spatial characteristics from individual soundfield signals and processes them separately through dedicated analysis modules. This extraction approach allows complex spatial information to be handled in a structured manner, reducing overall system complexity while preserving spatial immersion.
Solution Approach 2:
The audio processing system is segmented into independent modules for each endpoint, with separate talker presence detection and spatial characteristic analysis. This segmentation allows each module to handle one endpoint's spatial data independently, reducing the complexity of managing multiple simultaneous soundfields while maintaining spatial immersion.
3Reliability
If soundfield signals are continuously transmitted without selective multiplexing, then the perceptual continuity is improved, but the background noise increases
Solution Approach 1:
The system performs preliminary detection of talker presence and spatial characteristics for each endpoint before including that endpoint's soundfield in the mix. This preliminary action ensures that only endpoints with active talkers are included, maintaining perceptual continuity through seamless transitions while preventing background noise from inactive endpoints from degrading audio quality.
Solution Approach 2:
The system continuously monitors the activity state of each endpoint and provides feedback to the mixing configuration. This feedback mechanism ensures that the mix is dynamically adjusted to maintain perceptual continuity - when a talker becomes active, their soundfield is seamlessly integrated; when inactive, their contribution is reduced or eliminated, preventing background noise accumulation.
4Adaptability or versatility
If transitions between different soundfields are frequent, then the adaptability to active talkers is improved, but the unnatural shifts and artifacts increase
Solution Approach 1:
The system performs preliminary analysis of spatial characteristics and talker presence for all endpoints before making transition decisions. This preliminary action allows the system to anticipate necessary transitions and execute them smoothly, reducing unnatural shifts and artifacts while maintaining adaptability to active talkers.
Solution Approach 2:
The system prepares transition paths in advance by maintaining readiness to switch between soundfields based on predicted talker activity. This beforehand cushioning ensures that transitions are smooth and natural, preventing abrupt switching artifacts while maintaining the ability to quickly adapt to changing talker activity patterns.
Data Source
Figure 1A~1B
Figure 1C
Figure 2
AI summary
The present document relates to audio conference systems. In particular, the present document relates to improving the perceptual continuity within an audio conference system. According to an aspect, a method for multiplexing first and second continuous input audio signals is described, to yield a multiplexed output audio signal which is to be rendered to a listener. The first and second input audio signals (123) are indicative of sounds captured by a first and a second endpoint (120, 170), respectively. The method comprises determining a talk activity (201, 202) in the first and second input audio signals (123), respectively; and determining the multiplexed output audio signal based on the first and/or second input audio signals (123) and subject to one or more multiplexing conditions. The one or more multiplexing conditions comprise: at a time instant, when there is talk activity (201) in the first input audio signal (123), determining the multiplexed output audio signal at least based on the first input audio signal (123); at a time instant, when there is talk activity (202) in the second input audio signal (123), determining the multiplexed output audio signal at least based on the second input audio signal (123); and at a silence time instant, when there is no talk activity (201, 202) in the first and in the second input audio signals (123), determining the multiplexed output audio signal based on only one of the first and second input audio signals (123).