Spatial Audio Positioning for Multi-Party Call Speaker Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In modern call centers, agents face information overload during multi-party communications, as they struggle to identify active speakers and manage multiple participants in conference calls without conventional technology providing clear contextual information.

Innovation Solution

The implementation of a system that creates separate virtual audio locations for each call participant, allowing agents to perceive audio streams as coming from distinct positions in three-dimensional space, enhancing auditory localization and reducing errors by providing positional audio outputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If multiple audio streams are mixed in conventional telephony, then all participants can hear all speakers, but agents cannot identify which participant is actively speaking

Engineering Contradiction:
Improvespeaker identification informationVSAvoidaudio processing system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies spatial audio positioning to add a dimensional attribute to audio streams. Each participant's audio is assigned a specific spatial location in three-dimensional space, allowing agents to identify speakers through their positional cues rather than through complex audio mixing or visual interfaces. This transforms the audio experience from a flat mixed signal to a spatially distributed soundscape.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If agents monitor multiple communications channels simultaneously, then productivity increases, but information overload and difficulty identifying active speakers increases

Engineering Contradiction:
Improveagent multi-tasking capabilityVSAvoidagent cognitive load
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent segments the mixed audio stream into spatially separated individual participant audio streams. Each participant occupies a distinct spatial location, allowing agents to mentally organize and track multiple speakers without cognitive overload. This segmentation occurs in the auditory domain, enabling natural spatial grouping of information sources.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If visual interfaces are used to show active speakers, then speaker identification is clear, but agents cannot understand audio context without viewing the interface

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidaudio contextual information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent replaces the visual interface mechanism with an auditory spatial positioning mechanism. Instead of requiring agents to visually check screens to identify speakers, the system uses binaural audio and spatial cues to provide the same identification information through sound alone. This substitution maintains information accessibility while allowing agents to maintain audio context without visual distraction.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11632627B2Systems and methods for distinguishing audio using positional information
Publication Date: 2023.04.18 INCONTACT INC
  • US11632627B2 patent drawing
  • US11632627B2 patent drawing
  • US11632627B2 patent drawing

AI summary

A separate virtual (e.g. aural) location for one or more interaction or telephony call participants may provide an indication or clue for at least one of the call participants of who is speaking at any one time, reducing errors and misunderstandings during the call. Auditory localization may be used so that participants are heard from separate virtual locations. An audible user interface (AUI) may be produced such that audio presented to the listening user is location-specific, the location being relevant to the user, just as information presented in a graphical user interface (GUI) might be relevant. For example, a plurality of audio streams which are part of an interaction between communicating parties may be accepted, and based on the audio streams, a plurality of audio outputs may be provided, each located at a different location in three-dimensional space.