Cross-Lingual Speaker Recognition with Language-Adaptive Thresholds

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice biometric systems struggle to provide consistent speaker recognition across different languages, generating distinct results when speakers switch between languages.

Innovation Solution

A machine-learning architecture with an embedding extraction engine and a multi-class language classifier is employed to determine speaker similarity and language likelihood scores, using cross-lingual quality measures to adjust speaker verification scores and fine-tune the model with flip signal augmentation to compensate for language differences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional voice biometric systems are used, then speaker recognition can be performed, but the recognition results vary when speakers switch between languages

Engineering Contradiction:
Improvespeaker recognition consistencyVSAvoidlanguage adaptability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary language identification on enrollment and verification audio signals before speaker recognition. Based on the detected languages, it pre-adjusts the verification threshold to compensate for cross-lingual variations, ensuring consistent recognition results across different language scenarios

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically changes the verification threshold parameter based on language detection results. When different languages are detected between enrollment and verification signals, the threshold is adjusted to maintain reliable speaker recognition across language boundaries

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If language compensation functions are added to compensate for language differences, then cross-lingual speaker recognition accuracy improves, but system complexity increases

Engineering Contradiction:
Improvespeaker recognition accuracyVSAvoidsystem architecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the speaker recognition process into distinct modules: language identification module, quality measure calculation module, and threshold adjustment module. This segmentation allows each component to be optimized independently while maintaining overall system manageability and clarity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces a language identification result as an intermediary element that mediates between the audio signals and the verification threshold. This intermediary enables automatic adaptation to cross-lingual scenarios without requiring complex manual configuration or intervention

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If cross-lingual quality measures are used to adjust verification scores, then speaker verification accuracy across languages improves, but computational requirements increase

Engineering Contradiction:
Improveverification accuracyVSAvoidcomputational energy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system applies quality measure adjustment selectively based on language detection results. Rather than always performing full cross-lingual compensation, it only adjusts verification thresholds when different languages are detected, reducing unnecessary computational overhead while maintaining accuracy when needed

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12451138B2Cross-lingual speaker recognition
Publication Date: 2025.10.21 PINDROP SECURITY INC
  • US12451138B2 patent drawing
  • US12451138B2 patent drawing
  • US12451138B2 patent drawing

AI summary

Disclosed are systems and methods including computing-processes executing machine-learning architectures for voice biometrics, in which the machine-learning architecture implements one or more language compensation functions. Embodiments include an embedding extraction engine (sometimes referred to as an “embedding extractor”) that extracts speaker embeddings and determines a speaker similarity score for determine or verifying the likelihood that speakers in different audio signals are the same speaker. The machine-learning architecture further includes a multi-class language classifier that determines a language likelihood score that indicates the likelihood that a particular audio signal includes a spoken language. The features and functions of the machine-learning architecture described herein may implement the various language compensation techniques to provide more accurate speaker recognition results, regardless of the language spoken by the speaker.