Speaker Source Classification in Summed Audio Signals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio analysis methods face challenges in effectively separating and identifying speaker signals in summed audio interactions, leading to erroneous results and the need for tools to differentiate between agent and customer signals for quality monitoring and customer experience analysis.
Innovation Solution
A method for classifying audio signals into agent and customer signals using feature extraction, universal background modeling, and projection matrices to segment summed audio and associate signals correctly, enabling analysis such as emotion detection and quality monitoring.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If audio analysis tools are activated on summed audio signal, then analysis can be performed without separation, but the results are erroneous and unreliable
Solution Approach 1:
The patent segments the summed audio signal into separate speaker signals using speaker diarization techniques. This segmentation allows subsequent audio analysis tools to be applied to individual speaker signals rather than the mixed summed signal, thereby maintaining both analysis speed and improving accuracy by ensuring each tool processes clean, isolated speaker input.
2Reliability
If speaker separation is performed on summed audio, then separate signals can be obtained, but the signals may contain non-continuous segments due to speech overlap
Solution Approach 1:
The patent performs preliminary speaker diarization and segmentation before applying audio analysis tools. By pre-separating the summed signal into speaker-specific segments and identifying speech overlap regions, the system can then apply appropriate handling strategies (such as selecting dominant speaker or marking ambiguous regions) to maintain signal continuity while preserving separation accuracy.
3Device complexity
If audio is recorded as summed signal, then recording equipment is simplified, but speaker identification and source separation become difficult
Solution Approach 1:
The patent introduces speaker diarization and classification algorithms as intermediary processing steps between the summed audio recording and the final analysis. These intermediary tools extract speaker characteristics, separate speech segments by speaker, and identify speaker roles (customer vs. agent), thereby solving the speaker identification problem while maintaining the simplicity of the original summed recording approach.
4Adaptability or versatility
If different analyses are applied to different sides of interaction, then analysis relevance is improved, but verification of signal-side association is required
Solution Approach 1:
The patent implements automatic speaker-side classification that enables the system to self-identify which speaker signal corresponds to the customer side and which corresponds to the agent side. By using acoustic features, speech patterns, and contextual information to automatically label and categorize speaker signals, the system eliminates the need for manual verification while enabling customized analysis applications tailored to each speaker role.
Data Source
AI summary
A method for classifying a pair of audio signals into an agent audio signal and a customer audio signal. One embodiment relates to unsupervised training, in which the training corpus comprises a multiplicity of audio signal pairs, wherein each pair comprises an agent signal and a customer signal, and wherein it is unknown for each signal if it is by the agent or by the customer. Training is based on the agent signals being more similar to one another than the customer signals. An agent cluster and a customer cluster are determined. The input signals are associated with the agent or the customer according to the higher score combination of the input signals and the clusters. Another embodiment relates to supervised training, wherein an agent model is generated, and the input signal that yields higher score against the model is the agent signal, while the other is the customer signal.


