Voiceprint Frequency Mapping for Speaker Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speaker recognition systems face challenges in accurately matching voice data sampled at different frequencies, particularly when dealing with mixed bandwidth conditions, as they either lose information in upsampled signals or fail to exploit the richer quality of wideband speech samples.
Innovation Solution
A system that maps speaker recognition voiceprints from narrowband to wideband by obtaining voice vectors from signals sampled at different frequencies and uses a machine learning model to compare them, allowing for improved accuracy by leveraging the additional information in wideband signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If wideband data is downsampled to narrowband to use with narrowband speaker recognition model, then compatibility is improved, but information loss occurs in the upper frequency bands
Solution Approach 1:
Instead of downsampling wideband data to narrowband (the conventional approach), the patent inverts the process by upsampling narrowband data to wideband. A neural network model generates missing high-frequency components by learning from pairs of narrowband and wideband speech data, thereby preserving information while achieving compatibility across different sampling rates
Solution Approach 2:
The patent replaces the mechanical downsampling operation with a neural network-based upsampling system. The neural network learns the complex mapping between narrowband and wideband representations and generates synthetic high-frequency components, substituting a simple mechanical process with an intelligent system that preserves information
2Adaptability or versatility
If narrowband data is upsampled to wideband to use with wideband speaker recognition model, then bandwidth is improved, but accuracy deteriorates due to lack of information in upper frequency bands
Solution Approach 1:
The patent replaces simple mechanical upsampling with a neural network-based synthesis system. The neural network learns to generate realistic high-frequency components by training on wideband speech data, substituting a lossy mechanical process with an intelligent generation process that preserves accuracy
Solution Approach 2:
The patent changes the sampling rate parameter from narrowband (8kHz) to wideband (16kHz) through neural network processing. The network learns the statistical relationships between frequency bands and generates appropriate high-frequency components, transforming the signal parameters while maintaining information integrity
3Device complexity
If a single narrowband speaker recognition system is used for all applications, then simplicity is improved, but the ability to exploit wideband speech quality is lost
Solution Approach 1:
The patent introduces a dynamic upsampling process that adapts to the input narrowband data. The neural network dynamically generates wideband representations based on the specific characteristics of each speech sample, allowing the system to exploit wideband quality when available while maintaining simplicity through a unified architecture
Solution Approach 2:
The patent introduces a neural network upsampling module as an intermediary between the narrowband input and the wideband speaker recognition model. This intermediary transforms narrowband data into wideband representations, allowing the use of a single narrowband system interface while still exploiting wideband information through the learned transformation
Data Source
AI summary
There is provided a method that includes (a) obtaining a first voice vector that was derived from a signal of a voice that was sampled at a first sampling frequency, (b) obtaining a second voice vector that was derived from a signal of a voice that was sampled at a second sampling frequency, (c) mapping the second voice vector into a mapped voice vector in accordance with a machine learning model, and (d) comparing the first voice vector to the mapped voice vector to yield a score that indicates a probability that the first voice vector and the second voice vector originated from a same person.


