Speaker-Independent Audio Embeddings for Spoof-Resistant Verification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automatic speaker verification systems are susceptible to voice spoofing and background noise, and rely on unreliable telephony identifiers, necessitating a means to verify audio sources independent of speaker voice characteristics.

Innovation Solution

Systems and methods employing machine learning models, including Gaussian Mixture Models and neural networks, to extract speaker-independent characteristics from audio signals, generating deep-phoneprint vectors for authentication and downstream operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If speaker-dependent voice biometrics are used for authentication, then authentication can be performed, but the system becomes susceptible to voice spoofing and background noise

Engineering Contradiction:
Improveauthentication reliabilityVSAvoidvoice spoofing and background noise susceptibility
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent segments the audio signal analysis into two independent components: speaker-dependent voice biometrics and speaker-independent device characteristics. By analyzing device type, microphone type, geographical location, and network information separately from voice characteristics, the system can authenticate audio sources without being susceptible to voice spoofing attacks. This segmentation allows the speaker-independent component to remain effective even when voice biometrics are compromised.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces speaker-independent device characteristics as an intermediary authentication mechanism. Instead of directly relying on voice prints that can be spoofed, the system uses device characteristics (device type, microphone type, geographical location, network information) as an intermediary layer that is much more difficult to fake. This intermediary provides an additional authentication dimension that complements traditional voice biometrics.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If telephony identifiers such as caller ID are used for verification, then verification can be performed, but the identifiers become unreliable due to spoofing

Engineering Contradiction:
Improveverification reliabilityVSAvoidtelephony identifier reliability
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent extracts and analyzes speaker-independent characteristics from the audio signal itself, separating these characteristics from traditional telephony identifiers. By taking out device type, microphone type, geographical location, and network information directly from the audio analysis, the system creates a verification mechanism that does not depend on potentially spoofed telephony metadata. This extraction approach allows the system to verify audio sources based on actual acoustic characteristics rather than potentially falsified identifiers.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces reliance on telephony identifier systems with acoustic signal analysis. Instead of trusting metadata that can be easily spoofed, the system uses machine learning models to analyze the acoustic characteristics of the audio signal to determine device and location information. This substitution of mechanical/protocol-based verification with acoustic analysis provides more robust verification that is difficult to fake.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If speaker-independent characteristics are analyzed, then authentication becomes more secure against spoofing, but the system complexity increases

Engineering Contradiction:
Improvesecurity against spoofingVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a universal audio processing framework that can handle multiple authentication tasks simultaneously. The same speaker-independent characteristic extraction and machine learning models serve multiple purposes: authentication, verification, and characterization of audio sources. By making the system multi-functional, the patent reduces the need for separate specialized systems for each task, thereby managing complexity while enhancing security capabilities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent transforms the analysis approach by changing from speaker-dependent parameters (voice prints) to speaker-independent parameters (device type, microphone type, geographical location, network information). This parameter transformation allows the system to maintain security against spoofing while managing complexity through standardized extraction and classification processes that can be applied consistently across different audio sources and scenarios.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12437751B2Systems and methods of speaker-independent embedding for identification and verification from audio
Publication Date: 2025.10.07 PINDROP SECURITY INC
  • US12437751B2 patent drawing
  • US12437751B2 patent drawing
  • US12437751B2 patent drawing

AI summary

Embodiments described herein provide for audio processing operations that evaluate characteristics of audio signals that are independent of the speaker's voice. A neural network architecture trains and applies discriminatory neural networks tasked with modeling and classifying speaker-independent characteristics. The task-specific models generate or extract feature vectors from input audio data based on the trained embedding extraction models. The embeddings from the task-specific models are concatenated to form a deep-phoneprint vector for the input audio signal. The DP vector is a low dimensional representation of the each of the speaker-independent characteristics of the audio signal and applied in various downstream operations.