Speaker Verification via Agent Filtering and Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speaker verification techniques are unreliable due to imperfect audio segmentation in multi-party conversations, leading to false negatives and uncertainties when comparing voice characteristics with voice prints, especially in scenarios where agent and caller speech are mixed in a single audio recording.

Innovation Solution

The method involves decomposing multi-party single-channel audio into segments corresponding to individual speakers using automatic segmentation techniques and comparing these segments with both agent and user voice prints to determine speaker identity, employing 'agent filtering' to improve accuracy by attributing speech segments reliably.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If audio segmentation is performed in multi-party conversations, then speaker identification becomes possible, but segmentation imperfections cause false negatives and verification uncertainty

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidverification reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The audio stream is divided into multiple segments using automatic segmentation techniques, with each segment attributed to a specific speaker. This segmentation enables independent voiceprint comparison for each speaker, improving the precision of speaker identification while the system handles segmentation imperfections through multiple comparison points

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Voiceprint comparisons are performed preliminarily on segmented audio before making a final verification decision. The system compares each segmented speaker audio against both the agent voiceprint and user voiceprint in advance, allowing it to identify potential matches and handle segmentation errors through redundant comparisons

Inventive Principle:
Principle #10Preliminary action

2Reliability

If conventional voiceprint comparison is used without agent filtering, then processing is simpler, but false negatives increase and verification accuracy decreases

Engineering Contradiction:
Improveverification accuracyVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The agent voiceprint serves as an intermediary reference in the verification process. By first comparing segmented audio against the agent voiceprint, the system can filter out segments that belong to the agent, thereby reducing false negatives and improving overall verification accuracy without requiring complex additional hardware

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9728191B2Speaker verification methods and apparatus
Publication Date: 2017.08.08 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9728191B2 patent drawing
  • US9728191B2 patent drawing
  • US9728191B2 patent drawing

AI summary

Techniques for automatically identifying a speaker in a conversation as a known person based on processing of audio of the speaker's voice to extract characteristics of that voice and on an automated comparison of those characteristics to known characteristics of the known person's voice. A speaker segmentation process may be performed on audio of the conversation to produce, for each speaker in the conversation, a segment that includes the audio of that speaker. Audio of each of the segments may then be processed to extract characteristics of that speaker's voice. The characteristics derived from each segment (and thus for multiple speakers) may then be compared to characteristics of the known person's voice to determine whether the speaker for that segment is the known person. For each segment, a degree of match between the voice characteristics of the speaker and the voice characteristics of the known person may be calculated.