Voiceprint Extraction Using Neural Networks for Short Segments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice processing methods require long voice segments and cannot incorporate additional speaker information like age, gender, and language for effective identity verification.

Innovation Solution

A neural network-based voiceprint extractor is used to extract a holographic voiceprint from a short voice segment, which includes auxiliary information such as phoneme sequence, gender, age, language, and emotion, and then concatenated with a pre-stored voiceprint for verification using a pre-trained classification model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If joint factor analysis is used to extract i-vector voiceprint, then satisfactory verification performance is achieved, but long voice segment (20 to 30 seconds) is required

Engineering Contradiction:
Improveverification performanceVSAvoidvoice segment length
Core Design Contradiction:
ReliabilityVSLength of moving object

Solution Approach 1:

The patent changes the parameter of voice segment length from 20-30 seconds to 3-5 seconds by introducing a new deep neural network-based extraction method that learns hierarchical features automatically, eliminating the need for long segments while maintaining verification performance

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the conventional joint factor analysis method with a deep neural network-based voiceprint extraction method, substituting traditional statistical processing with learned hierarchical feature representation that achieves better performance with shorter input

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If conventional voiceprint extraction method is used, then verification is performed, but additional speaker information (age, gender, language) cannot be incorporated

Engineering Contradiction:
Improveverification capabilityVSAvoidinformation integration capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent makes the voiceprint extraction system multi-functional by designing the deep neural network to simultaneously extract voiceprint features and speaker attribute features (age, gender, language), allowing a single system to perform both identity verification and attribute analysis

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges the extraction of voiceprint features and speaker attribute features into a unified deep neural network model, combining multiple information types (phoneme sequence, gender, age, language, emotion) into an integrated holographic voiceprint representation

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP3346463B1Identity verification method and apparatus based on voiceprint
Publication Date: 2019.09.04 FUJITSU LTD
  • EP3346463B1 patent drawingFigure 1
  • EP3346463B1 patent drawingFigure 2
  • EP3346463B1 patent drawingFigure 3A

AI summary

An identity verification method and an identity verification apparatus based on a voiceprint are provided. The identity verification method based on a voiceprint includes: receiving an unknown voice; extracting a voiceprint of the unknown voice using a neural network-based voiceprint extractor which is obtained through pre-training; concatenating the extracted voiceprint with a pre-stored voiceprint to obtain a concatenated voiceprint; and performing judgment on the concatenated voiceprint using a pre-trained classification model, to verify whether the extracted voiceprint and the pre-stored voiceprint are from a same person. With the identity verification method and the identity verification apparatus, a holographic voiceprint of the speaker can be extracted from a short voice segment, such that the verification result is more robust.