3D Spatial Audio Engine for Contact Center Voice Positioning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In contact center transactions, it is challenging for agents to manage and distinguish between multiple participants' voices, as existing monaural voice communications do not allow for spatial localization of voices, leading to difficulties in managing conferences and maintaining private conversations.

Innovation Solution

A contact center media server equipped with a 3D spatial audio engine generates multi-channel voice signals that position each participant's voice at a specific aural location based on agent-designated positions, allowing agents to use a stereo headset to discern speakers and control who hears whom through a user interface.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If monaural voice communications are used in contact center transactions, then device complexity is reduced, but the ability to distinguish and localize participant voices deteriorates

Engineering Contradiction:
Improvecommunication system complexityVSAvoidvoice location discrimination
Core Design Contradiction:
Device complexityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies spatial audio positioning to add a dimensional aspect to voice communication by assigning three-dimensional spatial coordinates to different participants. This allows the contact center agent to distinguish between multiple participants based on their aural positions in space, transforming a monaural flat communication into a spatially-aware multi-dimensional experience that enhances voice localization without requiring complex multi-channel hardware for each participant.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If multiple participants are connected in a single conference, then communication efficiency is improved, but the ability to manage private conversations and control who hears whom deteriorates

Engineering Contradiction:
Improvecommunication efficiencyVSAvoidconference management control
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent segments the audio output to different participants by assigning unique spatial positions to each participant in the conference. This allows the contact center agent to control which participants can hear each other by manipulating their aural positions and applying selective muting, while maintaining all participants in a single conference call. The segmentation creates virtual audio channels that provide management control without requiring separate physical conference calls.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If stereo headsets or multiple speakers are used, then voice localization capability is improved, but device complexity and cost increase

Engineering Contradiction:
Improvevoice position discriminationVSAvoidaudio output equipment
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the contact center agent's existing audio equipment multi-functional by implementing software-based spatial audio positioning that works with standard monaural headsets or speakers. The system assigns three-dimensional spatial coordinates to different participants and processes audio signals to create directional perception, allowing a single audio output device to serve multiple participants with distinct spatial identities without requiring separate stereo headsets or speaker systems for each participant.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8363810B2Method and system for aurally positioning voice signals in a contact center environment
Publication Date: 2013.01.29 AVAYA INC
  • US8363810B2 patent drawing
  • US8363810B2 patent drawing
  • US8363810B2 patent drawing

AI summary

A contact center media server for aurally positioning participants of a contact center transaction at aural positions designated by a contact center agent. The media server includes a communications interface coupled to a controller and adapted to interface with a plurality of voice paths. Each of the voice paths is associated with one of a plurality of participants in a contact center transaction. A three-dimensional (3D) spatializer engine is coupled to the controller and can receive incoming voice signals received over voice paths and corresponding aural position data. The 3D spatializer engine processes the incoming voice signals and generates outgoing voice signals that include signal characteristics that aurally position the first outgoing voice signals at an aural position with respect to the contact center agent indicated by the aural position data.