Embedding Convertors for Cross-Channel Speaker Verification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional automatic speaker verification (ASV) systems face incompatibility issues due to differences in machine-learning architectures and sampling rates, leading to cumbersome and costly processes for users to update enrollment voiceprints across various systems.

Innovation Solution

A computing device executes software routines for speaker recognition using a machine-learning architecture with embedding extractors and convertors that map voiceprint embeddings from one type to another, enabling cross-compatibility and backward compatibility across different systems and channel requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If different machine-learning architectures are used for speaker verification, then system functionality and adaptability are improved, but compatibility between systems deteriorates

Engineering Contradiction:
Improvesystem functionalityVSAvoidcompatibility
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces embedding convertors as intermediary components that translate speaker embeddings between different machine-learning architectures (e.g., x-vector to i-vector). These convertors act as mediators that enable compatibility between systems using different architectures, allowing speaker verification to function across heterogeneous systems without requiring users to re-enroll in multiple systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter representation of speaker embeddings by applying transformation models that map embeddings from one feature space to another. This parameter transformation allows the same speaker representation to be adapted across different architectural paradigms, maintaining functionality while ensuring compatibility.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If different sampling rates are used for voice processing, then channel versatility is improved, but embedding compatibility deteriorates

Engineering Contradiction:
Improvechannel versatilityVSAvoidembedding compatibility
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent employs embedding convertors as intermediaries that handle embeddings extracted from audio signals with different sampling rates. The convertors normalize and transform these embeddings into a compatible format, enabling speaker verification across diverse communication channels (telephone, VoIP, mobile) without requiring uniform sampling rates.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If multiple enrollment voiceprints are required for different systems, then cross-system compatibility is improved, but user convenience and operational efficiency deteriorate

Engineering Contradiction:
Improvecross-system compatibilityVSAvoiduser convenience
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent creates a universal speaker verification system where a single enrollment voiceprint can be converted and used across multiple different ASV systems. The embedding convertors enable one voiceprint to serve multiple functions and systems, eliminating the need for separate enrollments and improving user convenience while maintaining cross-system compatibility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20230005486A1Speaker embedding conversion for backward and cross-channel compatability
Publication Date: 2023.01.05 PINDROP SECURITY INC
  • US20230005486A1 patent drawing
  • US20230005486A1 patent drawing
  • US20230005486A1 patent drawing

AI summary

Embodiments include a computer executing voice biometric machine-learning for speaker recognition. The machine-learning architecture includes embedding extractors that extract embeddings for enrollment or for verifying inbound speakers, and embedding convertors that convert enrollment voiceprints from a first type of embedding to a second type of embedding. The embedding convertor maps the feature vector space of the first type of embedding to the feature vector space of the second type of embedding. The embedding convertor takes as input enrollment embeddings of the first type of embedding and generates as output converted enrolled embeddings that are aggregated into a converted enrolled voiceprint of the second type of embedding. To verify an inbound speaker, a second embedding extractor generates an inbound voiceprint of the second type of embedding, and scoring layers determine a similarity between the inbound voiceprint and the converted enrolled voiceprint, both of which are the second type of embedding.