Speaker Verification Mono Audio Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Call centers face challenges in accurately verifying customer identities using mono audio recordings, as they lack stereo recording technology and have limited high-quality data for customers, leading to unreliable customer voiceprints.
Innovation Solution
The system employs speaker segmentation and verification by using a blind segmentation program to isolate customer speech from agent speech in mono recordings, leveraging a reliable voiceprint of the known agent to create or verify the customer's voiceprint, and utilizing an extraction module to remove agent speech and verify the customer's identity based on their voiceprint.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If stereo recording technology is used to separate customer and agent speech, then speaker verification accuracy is improved, but device complexity and cost increase
Solution Approach 1:
The patent applies segmentation by dividing the mono audio signal into separate customer speech and agent speech segments using blind source separation algorithms. This allows the system to achieve stereo-like verification accuracy without requiring stereo recording hardware, effectively segmenting the mixed audio signal into distinct speaker components for accurate voiceprint comparison.
Solution Approach 2:
The patent replaces the mechanical/stereo recording system with a computational signal processing approach. Instead of using separate microphones or stereo recording hardware to capture distinct audio channels, the system uses algorithmic separation of mono audio signals, substituting physical recording complexity with computational processing to achieve the same separation goal.
2Reliability
If more high-quality customer data is collected to improve voiceprint reliability, then verification accuracy is improved, but loss of time and increased processing complexity occur
Solution Approach 1:
The system applies self-service by using the agent's voiceprint (which is already reliable and readily available) to automatically identify and isolate customer speech segments. This eliminates the need for extensive manual customer data collection, as the system leverages existing agent data to serve the purpose of creating reliable customer voiceprints through automatic segmentation and extraction of customer speech portions.
Solution Approach 2:
The patent applies preliminary action by pre-processing the audio signals to separate customer and agent speech before voiceprint extraction. By performing blind source separation and speech segmentation in advance, the system prepares clean, isolated customer speech data that can be directly used for voiceprint creation, eliminating the need for time-consuming manual data collection and preprocessing steps later.
3Measurement precision
If blind segmentation is used to separate speakers in mono recordings, then verification accuracy is improved, but measurement precision of speaker isolation decreases
Solution Approach 1:
The patent introduces an intermediary approach by using blind source separation algorithms as a mediator between the mixed mono audio signal and the final speaker verification. These algorithms act as an intermediate processing step that separates the mixed signals into distinct speaker components, enabling accurate verification without requiring direct, perfect isolation of each speaker's audio stream.
Solution Approach 2:
The system applies feedback by using the known agent voiceprint to verify and refine the segmentation of agent speech from the mixed audio. This feedback mechanism allows the system to iteratively improve the accuracy of speaker isolation by comparing segmented segments against the known agent voiceprint and adjusting the separation process accordingly, thereby improving overall measurement precision.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
In many scenarios, speaker verification systems can be given a single-channel audio with recordings of multiple speakers. To perform accurate speaker verification, a system can isolate the speech of a speaker. In one embodiment, a method, and corresponding system, of speaker verification includes extracting a target speaker's speech, using a known speaker voiceprint, from an audio recording that includes the target speaker's speech and the known speaker's speech. The known speaker voiceprint can correspond to the known speaker. Extracting the target speaker's speech can include determining portions of the audio recording where the known speaker voiceprint matches the known speaker's speech above a particular threshold, and extracting the target speaker's speech from other portions of the audio recording. In this manner, speaker verification is performed on the target speaker's speech without interference from the known speaker's speech and allows for a more accurate verification.