3D Sound Conferencing Spatial Audio Voice Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing teleconferencing and distance learning systems fail to effectively distinguish and separate multiple voices speaking simultaneously, leading to a one-dimensional communication experience where participants cannot easily differentiate between speakers.

Innovation Solution

A 3D Sound Conferencing system is introduced, where each participant is associated with a position in a virtual room, and sound cues are transformed to simulate the direction and distance of voices, allowing listeners to perceive voices as originating from specific locations, enhancing the ability to distinguish between multiple speakers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If multiple voices are mixed together in existing teleconferencing systems, then all participants can hear all speakers, but participants cannot distinguish or understand multiple voices speaking simultaneously

Engineering Contradiction:
ImproveVoice distinction capabilityVSAvoidAudio processing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies spatial audio technology to add a spatial dimension to voice separation. By assigning different spatial positions to different speakers in a virtual environment, the system enables participants to distinguish multiple simultaneous voices through directional cues and spatial localization, transforming a one-dimensional audio mixing problem into a multi-dimensional spatial audio experience.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If distance learning systems use virtual classrooms, then participants can engage in group discussions, but listeners cannot readily differentiate between speakers

Engineering Contradiction:
ImproveSpeaker identificationVSAvoidListener effort to distinguish voices
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent employs visual indicators (analogous to color changes) such as avatar highlighting, name tags, and visual cues that change based on which speaker is currently active or whose voice is being emphasized. This visual-auditory linkage makes it easier for listeners to differentiate between speakers during group discussions in virtual classrooms.

Inventive Principle:
Principle #32Color changes

3Adaptability or versatility

If existing teleconferencing systems mix voices, then communication can occur, but the experience is relatively one dimensional

Engineering Contradiction:
ImproveCommunication dimensionalityVSAvoidSpatial audio cues
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent restores spatial audio cues by implementing a virtual environment where each participant has a defined position and orientation. Audio signals are processed to preserve and enhance spatial characteristics such as direction, distance, and reverberation, transforming flat one-dimensional audio mixing into immersive multi-dimensional spatial audio that mimics real-world acoustic environments.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentEP3039677B1Multidimensional virtual learning system and method
Publication Date: 2019.09.25 GLEIM CONFERENCING
  • EP3039677B1 patent drawingFigure 1
  • EP3039677B1 patent drawingFigure 2
  • EP3039677B1 patent drawingFigure 3

AI summary

A process and system for generating three dimensional sound conferencing includes generating a virtual map with a plurality of positions, each participant selecting one of the positions, determining a direction from each position to each other position on the map, determining a distance from each position to each other position on the map, receiving sound from each participant, mixing the received sound, transforming the mixed sound into binaural audio, and directing the binaural audio sound to each participant via a speaker associated with the virtual position of the speaking participant. The result is a clarified sound that gives to the listening participant a sense of where the speaking participant is positioned relative to the listening participant.