Embedding Convertors for Cross-Channel Speaker Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional automatic speaker verification (ASV) systems face incompatibility issues due to differences in machine-learning architectures and sampling rates, leading to cumbersome and costly processes for users to update enrollment voiceprints across various systems.
Innovation Solution
A computing device executes software routines for speaker recognition using a machine-learning architecture with embedding extractors and convertors that map voiceprint embeddings from one type to another, enabling cross-compatibility and backward compatibility across different systems and channel requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If different machine-learning architectures are used for speaker verification, then system functionality and adaptability are improved, but compatibility between systems deteriorates
Solution Approach 1:
The patent introduces embedding convertors as intermediary components that translate speaker embeddings between different machine-learning architectures (e.g., x-vector to i-vector). These convertors act as mediators that enable compatibility between systems using different architectures, allowing speaker verification to function across heterogeneous systems without requiring users to re-enroll in multiple systems.
Solution Approach 2:
The patent changes the parameter representation of speaker embeddings by applying transformation models that map embeddings from one feature space to another. This parameter transformation allows the same speaker representation to be adapted across different architectural paradigms, maintaining functionality while ensuring compatibility.
2Adaptability or versatility
If different sampling rates are used for voice processing, then channel versatility is improved, but embedding compatibility deteriorates
Solution Approach 1:
The patent employs embedding convertors as intermediaries that handle embeddings extracted from audio signals with different sampling rates. The convertors normalize and transform these embeddings into a compatible format, enabling speaker verification across diverse communication channels (telephone, VoIP, mobile) without requiring uniform sampling rates.
3Reliability
If multiple enrollment voiceprints are required for different systems, then cross-system compatibility is improved, but user convenience and operational efficiency deteriorate
Solution Approach 1:
The patent creates a universal speaker verification system where a single enrollment voiceprint can be converted and used across multiple different ASV systems. The embedding convertors enable one voiceprint to serve multiple functions and systems, eliminating the need for separate enrollments and improving user convenience while maintaining cross-system compatibility.
Data Source
AI summary
Embodiments include a computer executing voice biometric machine-learning for speaker recognition. The machine-learning architecture includes embedding extractors that extract embeddings for enrollment or for verifying inbound speakers, and embedding convertors that convert enrollment voiceprints from a first type of embedding to a second type of embedding. The embedding convertor maps the feature vector space of the first type of embedding to the feature vector space of the second type of embedding. The embedding convertor takes as input enrollment embeddings of the first type of embedding and generates as output converted enrolled embeddings that are aggregated into a converted enrolled voiceprint of the second type of embedding. To verify an inbound speaker, a second embedding extractor generates an inbound voiceprint of the second type of embedding, and scoring layers determine a similarity between the inbound voiceprint and the converted enrolled voiceprint, both of which are the second type of embedding.


