Voiceprint Frequency Mapping for Speaker Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speaker recognition systems face challenges in accurately matching voice data sampled at different frequencies, particularly when dealing with mixed bandwidth conditions, as they either lose information in upsampled signals or fail to exploit the richer quality of wideband speech samples.

Innovation Solution

A system that maps speaker recognition voiceprints from narrowband to wideband by obtaining voice vectors from signals sampled at different frequencies and uses a machine learning model to compare them, allowing for improved accuracy by leveraging the additional information in wideband signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If wideband data is downsampled to narrowband to use with narrowband speaker recognition model, then compatibility is improved, but information loss occurs in the upper frequency bands

Engineering Contradiction:
ImprovecompatibilityVSAvoidinformation loss
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

Instead of downsampling wideband data to narrowband (the conventional approach), the patent inverts the process by upsampling narrowband data to wideband. A neural network model generates missing high-frequency components by learning from pairs of narrowband and wideband speech data, thereby preserving information while achieving compatibility across different sampling rates

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent replaces the mechanical downsampling operation with a neural network-based upsampling system. The neural network learns the complex mapping between narrowband and wideband representations and generates synthetic high-frequency components, substituting a simple mechanical process with an intelligent system that preserves information

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If narrowband data is upsampled to wideband to use with wideband speaker recognition model, then bandwidth is improved, but accuracy deteriorates due to lack of information in upper frequency bands

Engineering Contradiction:
ImprovebandwidthVSAvoidaccuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent replaces simple mechanical upsampling with a neural network-based synthesis system. The neural network learns to generate realistic high-frequency components by training on wideband speech data, substituting a lossy mechanical process with an intelligent generation process that preserves accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the sampling rate parameter from narrowband (8kHz) to wideband (16kHz) through neural network processing. The network learns the statistical relationships between frequency bands and generates appropriate high-frequency components, transforming the signal parameters while maintaining information integrity

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If a single narrowband speaker recognition system is used for all applications, then simplicity is improved, but the ability to exploit wideband speech quality is lost

Engineering Contradiction:
ImprovesimplicityVSAvoidwideband information
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent introduces a dynamic upsampling process that adapts to the input narrowband data. The neural network dynamically generates wideband representations based on the specific characteristics of each speech sample, allowing the system to exploit wideband quality when available while maintaining simplicity through a unified architecture

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent introduces a neural network upsampling module as an intermediary between the narrowband input and the wideband speaker recognition model. This intermediary transforms narrowband data into wideband representations, allowing the use of a single narrowband system interface while still exploiting wideband information through the learned transformation

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20230267936A1Frequency mapping in the voiceprint domain
Publication Date: 2023.08.24 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20230267936A1 patent drawing
  • US20230267936A1 patent drawing
  • US20230267936A1 patent drawing

AI summary

There is provided a method that includes (a) obtaining a first voice vector that was derived from a signal of a voice that was sampled at a first sampling frequency, (b) obtaining a second voice vector that was derived from a signal of a voice that was sampled at a second sampling frequency, (c) mapping the second voice vector into a mapped voice vector in accordance with a machine learning model, and (d) comparing the first voice vector to the mapped voice vector to yield a score that indicates a probability that the first voice vector and the second voice vector originated from a same person.