Speaker Verification via Agent Filtering and Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speaker verification techniques are unreliable due to imperfect audio segmentation in multi-party conversations, leading to false negatives and uncertainties when comparing voice characteristics with voice prints, especially in scenarios where agent and caller speech are mixed in a single audio recording.
Innovation Solution
The method involves decomposing multi-party single-channel audio into segments corresponding to individual speakers using automatic segmentation techniques and comparing these segments with both agent and user voice prints to determine speaker identity, employing 'agent filtering' to improve accuracy by attributing speech segments reliably.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If audio segmentation is performed in multi-party conversations, then speaker identification becomes possible, but segmentation imperfections cause false negatives and verification uncertainty
Solution Approach 1:
The audio stream is divided into multiple segments using automatic segmentation techniques, with each segment attributed to a specific speaker. This segmentation enables independent voiceprint comparison for each speaker, improving the precision of speaker identification while the system handles segmentation imperfections through multiple comparison points
Solution Approach 2:
Voiceprint comparisons are performed preliminarily on segmented audio before making a final verification decision. The system compares each segmented speaker audio against both the agent voiceprint and user voiceprint in advance, allowing it to identify potential matches and handle segmentation errors through redundant comparisons
2Reliability
If conventional voiceprint comparison is used without agent filtering, then processing is simpler, but false negatives increase and verification accuracy decreases
Solution Approach 1:
The agent voiceprint serves as an intermediary reference in the verification process. By first comparing segmented audio against the agent voiceprint, the system can filter out segments that belong to the agent, thereby reducing false negatives and improving overall verification accuracy without requiring complex additional hardware
Data Source
AI summary
Techniques for automatically identifying a speaker in a conversation as a known person based on processing of audio of the speaker's voice to extract characteristics of that voice and on an automated comparison of those characteristics to known characteristics of the known person's voice. A speaker segmentation process may be performed on audio of the conversation to produce, for each speaker in the conversation, a segment that includes the audio of that speaker. Audio of each of the segments may then be processed to extract characteristics of that speaker's voice. The characteristics derived from each segment (and thus for multiple speakers) may then be compared to characteristics of the known person's voice to determine whether the speaker for that segment is the known person. For each segment, a degree of match between the voice characteristics of the speaker and the voice characteristics of the known person may be calculated.


