Speaker Source Classification in Summed Audio Signals

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio analysis methods face challenges in effectively separating and identifying speaker signals in summed audio interactions, leading to erroneous results and the need for tools to differentiate between agent and customer signals for quality monitoring and customer experience analysis.

Innovation Solution

A method for classifying audio signals into agent and customer signals using feature extraction, universal background modeling, and projection matrices to segment summed audio and associate signals correctly, enabling analysis such as emotion detection and quality monitoring.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If audio analysis tools are activated on summed audio signal, then analysis can be performed without separation, but the results are erroneous and unreliable

Engineering Contradiction:
Improveanalysis speedVSAvoidanalysis accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the summed audio signal into separate speaker signals using speaker diarization techniques. This segmentation allows subsequent audio analysis tools to be applied to individual speaker signals rather than the mixed summed signal, thereby maintaining both analysis speed and improving accuracy by ensuring each tool processes clean, isolated speaker input.

Inventive Principle:
Principle #1Segmentation

2Reliability

If speaker separation is performed on summed audio, then separate signals can be obtained, but the signals may contain non-continuous segments due to speech overlap

Engineering Contradiction:
Improvesignal separation accuracyVSAvoidsignal continuity
Core Design Contradiction:
ReliabilityVSStability of the object's composition

Solution Approach 1:

The patent performs preliminary speaker diarization and segmentation before applying audio analysis tools. By pre-separating the summed signal into speaker-specific segments and identifying speech overlap regions, the system can then apply appropriate handling strategies (such as selecting dominant speaker or marking ambiguous regions) to maintain signal continuity while preserving separation accuracy.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If audio is recorded as summed signal, then recording equipment is simplified, but speaker identification and source separation become difficult

Engineering Contradiction:
Improverecording system complexityVSAvoidspeaker identification difficulty
Core Design Contradiction:
Device complexityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent introduces speaker diarization and classification algorithms as intermediary processing steps between the summed audio recording and the final analysis. These intermediary tools extract speaker characteristics, separate speech segments by speaker, and identify speaker roles (customer vs. agent), thereby solving the speaker identification problem while maintaining the simplicity of the original summed recording approach.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Adaptability or versatility

If different analyses are applied to different sides of interaction, then analysis relevance is improved, but verification of signal-side association is required

Engineering Contradiction:
Improveanalysis customizationVSAvoidsignal verification complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements automatic speaker-side classification that enables the system to self-identify which speaker signal corresponds to the customer side and which corresponds to the agent side. By using acoustic features, speech patterns, and contextual information to automatically label and categorize speaker signals, the system eliminates the need for manual verification while enabling customized analysis applications tailored to each speaker role.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS8306814B2Method for speaker source classification
Publication Date: 2012.11.06 NICE LTD
  • US8306814B2 patent drawing
  • US8306814B2 patent drawing
  • US8306814B2 patent drawing

AI summary

A method for classifying a pair of audio signals into an agent audio signal and a customer audio signal. One embodiment relates to unsupervised training, in which the training corpus comprises a multiplicity of audio signal pairs, wherein each pair comprises an agent signal and a customer signal, and wherein it is unknown for each signal if it is by the agent or by the customer. Training is based on the agent signals being more similar to one another than the customer signals. An agent cluster and a customer cluster are determined. The input signals are associated with the agent or the customer according to the higher score combination of the input signals and the clusters. Another embodiment relates to supervised training, wherein an agent model is generated, and the input signal that yields higher score against the model is the agent signal, while the other is the customer signal.