Voiceprint Extraction Using Neural Networks for Short Segments
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice processing methods require long voice segments and cannot incorporate additional speaker information like age, gender, and language for effective identity verification.
Innovation Solution
A neural network-based voiceprint extractor is used to extract a holographic voiceprint from a short voice segment, which includes auxiliary information such as phoneme sequence, gender, age, language, and emotion, and then concatenated with a pre-stored voiceprint for verification using a pre-trained classification model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If joint factor analysis is used to extract i-vector voiceprint, then satisfactory verification performance is achieved, but long voice segment (20 to 30 seconds) is required
Solution Approach 1:
The patent changes the parameter of voice segment length from 20-30 seconds to 3-5 seconds by introducing a new deep neural network-based extraction method that learns hierarchical features automatically, eliminating the need for long segments while maintaining verification performance
Solution Approach 2:
The patent replaces the conventional joint factor analysis method with a deep neural network-based voiceprint extraction method, substituting traditional statistical processing with learned hierarchical feature representation that achieves better performance with shorter input
2Reliability
If conventional voiceprint extraction method is used, then verification is performed, but additional speaker information (age, gender, language) cannot be incorporated
Solution Approach 1:
The patent makes the voiceprint extraction system multi-functional by designing the deep neural network to simultaneously extract voiceprint features and speaker attribute features (age, gender, language), allowing a single system to perform both identity verification and attribute analysis
Solution Approach 2:
The patent merges the extraction of voiceprint features and speaker attribute features into a unified deep neural network model, combining multiple information types (phoneme sequence, gender, age, language, emotion) into an integrated holographic voiceprint representation
Data Source
Figure 1
Figure 2
Figure 3A
AI summary
An identity verification method and an identity verification apparatus based on a voiceprint are provided. The identity verification method based on a voiceprint includes: receiving an unknown voice; extracting a voiceprint of the unknown voice using a neural network-based voiceprint extractor which is obtained through pre-training; concatenating the extracted voiceprint with a pre-stored voiceprint to obtain a concatenated voiceprint; and performing judgment on the concatenated voiceprint using a pre-trained classification model, to verify whether the extracted voiceprint and the pre-stored voiceprint are from a same person. With the identity verification method and the identity verification apparatus, a holographic voiceprint of the speaker can be extracted from a short voice segment, such that the verification result is more robust.