Audio Scene Manager for Spatial Teleconference Sound
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current teleconference systems fail to effectively simulate the spatial distribution of sound sources, leading to difficulties in distinguishing between multiple speakers, as sound origins appear from the same spatial position, making it unpleasant for participants to follow conversations.
Innovation Solution
A method and device for group sound telecommunication that processes signals to position sound sources at different angles relative to the listener, adjusting angle separation based on sound activity, allowing sounds from active speakers to be perceived as emanating from distinct angles, thereby enhancing the spatial separation and naturalness of the virtual meeting environment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If participants are positioned evenly around a round table in 3D audio teleconference, then spatial distribution of sound sources is achieved, but listeners cannot turn their virtual head to follow talkers, resulting in poor sound separation when multiple persons speak simultaneously
Solution Approach 1:
The patent applies dynamics by making the virtual head position adjustable and movable. The system allows the virtual head to be turned towards active talkers dynamically, enabling listeners to follow conversations naturally. This resolves the contradiction by transforming the static virtual head position into a dynamic one that can adapt to different speaking situations.
Solution Approach 2:
The patent changes the parameter of virtual head orientation angle based on sound activity detection. When a talker is identified, the system adjusts the head turning angle to orient towards that talker, optimizing sound reception. This parameter change enables the listener to focus on active speakers without manual intervention.
2Measurement precision
If the listener actively turns the virtual head towards talkers, then sound reception is maximized, but concentration is stolen from what persons are actually saying
Solution Approach 1:
The system applies self-service by automatically detecting which participant is talking and autonomously turning the virtual head towards that talker. The listener does not need to manually control head orientation; the system serves itself by monitoring sound activity and adjusting head position accordingly, freeing the listener to concentrate on the conversation content.
Solution Approach 2:
The patent implements feedback by continuously monitoring sound activity levels from different participants and using this information to adjust virtual head orientation. The system receives feedback about who is speaking and automatically responds by orienting the virtual head towards the active talker, creating a closed-loop control system.
3Measurement precision
If advanced positioning equipment is provided to automatically detect head direction, then accurate sound positioning is achieved, but device complexity increases
Solution Approach 1:
The patent replaces complex mechanical positioning equipment with a software-based virtual head turning mechanism. Instead of using physical sensors and actuators to detect and adjust head direction, the system uses signal processing and virtual reality techniques to simulate head orientation changes, significantly reducing hardware complexity while maintaining functionality.
Data Source
Figure 1A~1B
Figure 2
Figure 3A~3D
AI summary
A method of audio scene management in a teleconference or other group sound telecommunication is presented, in which teleconference at least a first transmitting party, a second transmitting party and a receiving party participates. The method comprises receiving (210) of signals representing sound of the first transmitting party and sound of the second transmitting party. The method further comprises obtaining (212) of measures of sound activity for the first and second transmitting parties, respectively and selecting (214) a first angle and/ or a second angle based on the obtained measures of sound activity. The method further comprises processing (216) of the received signals into processed signals such that sound from the first transmitting party is experienced by the receiving party as emanating from the first angle while sound from the second transmitting party is experienced as emanating from the second angle, with respect to the receiving party. Finally signals representing the processed signals are outputted (218).