Spatial Audio Rendering for Screen-Correlated Video Conferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio and video-conferencing systems fail to effectively provide spatial cues, making it difficult for participants to perceive sounds as coming from the direction of remotely located participants displayed on the screen, limiting effective collaboration in geographically distributed settings.
Innovation Solution
A multimedia-conferencing system that uses stereo loudspeakers positioned on either side of a display to play sounds in a way that simulates the location of remote participants on the screen by calculating and adjusting interaural time and intensity differences, creating the illusion that sounds are emanating from the remote participant's image.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If loudspeaker based methods are used to address cross talk cancellation, then audio can be played through loudspeakers, but the audio output does not correlate with the location of participants displayed on the screen
Solution Approach 1:
The audio signal from each remote participant is segmented and assigned to specific loudspeakers based on their displayed location. The system divides the audio output into multiple channels corresponding to different spatial positions on the display, enabling independent control of audio from each direction.
Solution Approach 2:
Different audio processing characteristics are applied to different loudspeakers based on the location of participants on the display. Each loudspeaker receives audio with specific interaural time and intensity differences tailored to simulate sound coming from the corresponding direction relative to the participant's position.
2Loss of information
If HRTF based methods are used for stereo spatial audio rendering, then spatial audio cues can be provided, but participants have to wear headphones
Solution Approach 1:
The system uses loudspeakers as an intermediary device to deliver spatial audio cues without requiring headphones. By positioning loudspeakers around the display and applying directional audio processing, the system mediates between the audio signal and the participant's ears, creating spatial perception through the loudspeaker array rather than direct ear canal insertion.
3Illumination intensity
If conventional video-conferencing systems are used, then participants can see each other on the screen, but the audio does not correlate with the location of the participants presented on the display
Solution Approach 1:
The system merges the visual display system with the audio output system by spatially correlating participant images on the display with audio playback from corresponding directions. The audio-visual information is combined in terms of spatial positioning, creating a unified perceptual experience where sound and image originate from the same location.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables local participants to perceive sounds as originating from the direction of the remote participant's image on the display, enhancing the realism and effectiveness of remote collaboration by correlating audio with visual cues.
Implementation Method 1
adjusting interaural time and intensity differences to simulate the sound emanating from the remote participant's location on the display
Implementation Method 2
adjusting interaural time and intensity differences to simulate the sound emanating from the remote participant's location on the display
Data Source
AI summary
Disclosed herein are multimedia-conferencing systems and methods enabling local participants to hear remote participants from the direction the remote participants are rendered on a display. In one aspect, a method includes a computing device receives a remote participant's image and sound information collected at a remote site. The remote participant's image is rendered on a display at a local site. When the local participant is in close proximity to the display, sounds generated by the remote participant are played over stereo loudspeakers so that the local participant perceives the sounds as emanating from the remote participant's location rendered on the display.


