Voiceprint recognition model construction and recognition method based on deep learning
By combining deep learning models of convolutional neural networks, recurrent neural networks and long and short-term memory networks, and combining joint training of feature enhancement networks and speaker embedding networks, the problem of low recognition accuracy of voiceprint recognition in a reverberation environment is solved, the robustness and recognition accuracy of the model are improved, and the needs of real-time and low-power devices are met.
Patent Information
- Application Number
- CN202510407531.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-06-27
AI Technical Summary
The existing voiceprint recognition technology has low recognition accuracy in reverberation environments, limited computing resources, and cannot meet real-time requirements, and has poor performance in scenarios where noise, reverberation and multi-speaker exist simultaneously.
The voiceprint recognition model based on deep learning is adopted, combining convolutional neural networks, recurrent neural networks and long-term memory networks, and the joint training of the network and speaker embedding network through multi-level feature extraction and feature enhancement network, optimize and enhance the parameters of the network layer and speaker embedding network layer, and use asynchronous sub-region optimization methods and feature connection technology to improve the robustness and recognition accuracy of the model.
It significantly improves the recognition accuracy in noisy environments, improves the model's adaptability to complex environments, reduces the computing complexity and storage requirements, enables the model to run efficiently on low-power devices, meets real-time requirements, and achieves efficient end-to-end voiceprint recognition on multiple devices.
Smart Images

Figure CN120220692A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of deep learning and artificial intelligence, and specifically to a construction and recognition method of a voiceprint recognition model based on deep learning. Background Art
[0002] With the rapid development of biometric authentication technology, voiceprint recognition is increasingly becoming an important technical means in the field of identity authentication due to its unique advantages. As a behavior feature authentication method based on voice signals, voiceprint recognition not only has the characteristics of non-contact, low cost, and high convenience, but also has high anti-counterfeiting security due to the complexity and dynamics of biometric features. Compared with other biometric features such as face and fingerprint, voiceprint recognition shows significant advantages in terms of collection convenience, user privacy protection, and system deployment cost. This technology extracts personalized features in voice signals through a deep learning model, maps the acoustic characteristics of a speaker into an embedded vector of a fixed dimension, and this end-to-end feature representation method effectively improves the accuracy and reliability of recognition. In terms of model optimization, researchers have developed a variety of improved loss functions to enhance the discriminative ability of features, and at the same time adopted voice noise augmentation technology. By simulating various noise and reverberation environments during the training stage, the robustness of the system in real complex scenarios has been significantly improved.
[0003] Currently, voiceprint recognition technology still faces many challenges in practical applications: environmental noise will significantly interfere with the quality of voice signals, especially in noisy scenarios, the system recognition performance will significantly decline; the problem of voice overlap in a multi-speaker environment increases the difficulty of separating and recognizing the target speaker; the limited acoustic features under short speech conditions make it more challenging to extract speaker information; the difference in feature distribution caused by different recording devices and acoustic environments affects the cross-scene robustness of the system; the mixing of direct sound and reflected sound in a reverberant environment reduces the recognition accuracy of existing technologies. In addition, traditional signal processing methods have problems of high computational complexity and low resource utilization, and it is difficult to meet the requirements of mobile deployment; although deep learning-based methods have excellent performance, they still face the problem of insufficient generalization ability caused by the mismatch between training and testing environments, as well as the contradiction between model complexity and real-time requirements; end-to-end systems are limited by the lack of effective acoustic prior knowledge embedding and have limited adaptability to complex acoustic environments, especially in scenarios where noise, reverberation, and multiple speakers exist simultaneously.
[0004] In the prior art, such as the publication number CN114913860A, in terms of data processing and model training, data augmentation is performed on small-sample speech data through a generative adversarial network, and the convolutional layer parameters of the generative adversarial network model are migrated to the speaker recognition model to accelerate the training convergence rate and improve the accuracy; in terms of feature extraction and model structure, feature extraction is mainly performed through traditional preprocessing methods, and the convolutional layer structures of the speaker recognition model and the generative adversarial network model are the same; in terms of loss function and optimization method, optimization is mainly performed by updating network parameters; in terms of system architecture and application scenarios, it focuses on the construction and training process of the speaker recognition model, and the application scenarios are described relatively broadly. Therefore, there is insufficient environmental adaptability, resulting in poor robustness of the model in complex acoustic environments. Such as the publication number KR102294638B1, the core technology focuses on using a combination of a feature enhancement model based on a deep neural network and a modified loss function for combined learning to improve the robustness of speaker recognition in a noisy environment. Its innovation lies in combining the feature enhancement model with the speaker feature vector extraction model and optimizing the entire system through joint learning. Therefore, there are limitations in the model architecture, affecting the expression ability of complex acoustic features. The joint learning mechanism may be relatively simple, and the combination method of the feature enhancement model and the feature extraction model may not be tight enough. Such as the publication number US12067989B2, a combined learning method based on a deep neural network for feature enhancement and an improved loss function is proposed. This method optimizes the overall performance by jointly training the feature enhancement model and the speaker feature vector extraction model. The feature enhancement model learns by minimizing the mean square error (MSE) to convert the acoustic features of degraded speech data into the acoustic features of clean speech data. The speaker feature vector extraction model uses the x-vector model to extract speaker-related information through a time delay neural network (TDNN) layer and calculates the mean and standard deviation in the statistical feature extraction layer to generate a fixed-length speaker feature vector. In the joint training, by modifying the loss function, a margin of subtracting a specific constant value from the output value of the target speaker is added to increase the posterior probability of the speaker, thereby improving the recognition performance of the model in a noisy environment. Therefore, there are limitations in the feature enhancement method, restricting the expression ability of complex acoustic features, and the model complexity may be difficult to meet the deployment requirements of low-power devices. Such as the publication number WO2020204525A1, it focuses on the combined learning method of feature enhancement and modified loss function. Its core lies in combining the feature enhancement model based on a deep neural network with the speaker feature vector extraction model and optimizing the overall performance through joint learning. The feature enhancement model learns by minimizing the mean square error to convert the acoustic features of degraded speech data into the acoustic features of clean speech data. The speaker feature vector extraction model uses the x-vector model to extract speaker-related information through a time delay neural network layer and calculates the mean and standard deviation in the statistical feature extraction layer to generate a fixed-length speaker feature vector.During the collaborative learning process, the posterior probability of the speaker is increased by modifying the loss function, that is, subtracting a margin of a specific constant value from the output value of the target speaker, thereby improving the recognition performance of the model in a noisy environment. Therefore, there are limitations in signal preprocessing, unable to effectively handle non-linear distortion and complex acoustic interference, the convergence and generalization ability of the model are limited, only focusing on the feature enhancement module, lacking end-to-end optimization from signal preprocessing to feature extraction. Summary of the Invention
[0005] The present invention solves the problems in the above voiceprint recognition process, such as low recognition accuracy in a reverberant environment, limited computing resources, and inability to meet real-time requirements. A voiceprint recognition model construction and recognition method based on deep learning is proposed. By combining end-to-end models such as convolutional neural networks, recurrent neural networks, and long short-term memory networks, effective extraction of voiceprint features and accurate identification of identities are achieved, solving the above technical problems, especially problems such as performance degradation in reverberant and noisy environments. The technical solutions are as follows:
[0006] Solution 1.
[0007] A method for constructing a voiceprint recognition model based on deep learning, comprising the following steps:
[0008] S1, input the original speech signal, and output a spectrogram through noise augmentation and preprocessing;
[0009] S2, sequentially pass the spectrogram through multi-level feature extraction of shallow physical features, middle-layer vocal tract features, and deep embedding features to generate the first speaker embedding feature;
[0010] S3, input the first speaker embedding feature into a hybrid neural network to extract local feature information and temporal dependence relationships, and generate a global speech feature representation;
[0011] S4, pass the speech feature representation through a feature enhancement network and a speaker embedding network to generate a second speaker embedding feature;
[0012] S5, perform model training on the second speaker embedding feature with joint optimization of the loss function, divide the network into different regions, and coordinate the influence of different loss functions on network optimization through an asynchronous sub-region optimization method, optimize the enhanced network layer to obtain network parameters that minimize the denoising error, and optimize the speaker embedding network layer to obtain network parameters that maximize the identity discrimination degree.
[0013] Further, the noise augmentation in step S1 includes:
[0014] Randomly select noise signals from the noise database and reverberation signals from the reverberation database; combine the noise signals with the reverberation signals to generate noise samples with different reverberation characteristics; superimpose the noise samples on the original speech signal to generate noisy speech data.
[0015] Further, the preprocessing in step S1 includes: performing at least frame segmentation and windowing operations on the original speech signal or the noisy speech data after noise augmentation.
[0016] Segment the speech signal into frames of a fixed length, perform windowing processing on each frame of the signal; convert the preprocessed speech signal into a spectrogram.
[0017] Further, the shallow physical feature extraction in step S2 includes: extracting the spectrogram, converting the time-domain signal into a frequency-domain signal, calculating the amplitude spectrum of each frame to generate a spectrogram, and simultaneously calculating the short-time energy of the speech signal.
[0018] Further, the middle-level vocal tract feature extraction in step S2 includes: extracting vocal tract position information, inferring the vocal tract position information by analyzing the spectral characteristics of the speech signal, calculating the LPC linear prediction coefficients, and extracting formant frequency and bandwidth parameters.
[0019] Further, the deep embedding feature extraction in step S2 includes: constructing a multi-layer neural network structure, introducing non-linear factors using non-linear activation functions, the network learning the feature mapping relationship, and generating the first speaker embedding feature at the output layer of the network.
[0020] The multi-layer neural network structure includes: fully connected layers and convolutional layers.
[0021] Further, step S3 includes:
[0022] S31, the convolutional neural network extracts local patterns from the spectrogram and recognizes features including prosody and formants; and / or
[0023] S32, the recurrent neural network or long short-term memory network captures the dependencies between time steps in the speech signal based on the results of step S31 to generate a global speech feature representation.
[0024] Further, step S4 includes:
[0025] S41, the feature enhancement network performs denoising and dereverberation processing on the input speech features to generate enhanced features;
[0026] S42, input the enhanced features into the speaker embedding network, extract discriminative features through deep convolutional operations to generate the second speaker embedding feature.
[0027] Further, the discriminative ability of the second speaker embedding feature is greater than that of the first speaker embedding feature.
[0028] Solution 2.
[0029] A voiceprint recognition method based on deep learning, comprising: using the voiceprint recognition model described in Solution 1 to perform the following steps:
[0030] S1. Preprocessing of speech signals: converting the original speech signal into an input format that can be processed by the model, and mapping it end-to-end to a speaker label; the original speech signal includes: phrases, keywords or continuous speech segments;
[0031] S2. Deep feature extraction network: performing multi-level feature extraction on the speech signal through a deep neural network to generate a speaker embedding feature;
[0032] S3. Matching the speaker embedding feature with the identity label in the pre-stored database, and outputting a user authentication result;
[0033] Among them, the multi-level feature extraction includes: shallow physical features, middle-level vocal tract features and deep embedding features.
[0034] The present invention has the following beneficial effects:
[0035] 1. The voiceprint recognition model construction and recognition method based on deep learning according to the present invention can effectively simulate various noise conditions in the real environment by adding noise and reverberation processing by real-time selection of noise data, and enhance the adaptability of the model to complex environments. This method improves the adaptability of the model to background noise and reverberation, and significantly improves the recognition accuracy in a noisy environment; combines multiple deep learning models such as convolutional neural network, recurrent neural network and long short-term memory network, and makes full use of its feature extraction ability to extract deeper voiceprint features. This multi-model fusion method can better capture complex speech features compared with traditional single models; introduces a discriminative loss function, and adopts a joint training framework of a feature enhancement network and a speaker embedding network, which can more effectively optimize the similarity and difference between samples, thereby improving the discrimination ability of the model and enhancing the robustness and accuracy of the voiceprint recognition system. The flexible system architecture simplifies the user interaction process and improves the reliability and flexibility of the system. Only necessary parameters need to be provided, and there is no need to directly face a complex decoder.
[0036] 2. The method for constructing and recognizing a voiceprint recognition model based on deep learning according to the present invention realizes efficient end-to-end recognition in a complex environment through model optimization, robustness improvement, and real-time guarantee. In terms of model optimization, a feature enhancement network and a speaker embedding network are adopted, reducing the computational complexity and storage requirements, enabling the model to operate efficiently on low-power devices. To improve the robustness of the system, a real-time noise speech augmentation technique is introduced during the training phase. Noise samples are randomly selected from a noise database and combined with reverberation to add noise to the original audio in real time, enhancing the model's adaptability to complex environments. At the same time, a joint training framework of the feature enhancement network and the speaker verification network is adopted. The enhancement network is optimized through the least mean square error, and specific modules of the embedding network and the enhancement network are optimized using the speaker loss function, improving the model's denoising ability and recognition accuracy.
[0037] 3. The method for constructing and recognizing a voiceprint recognition model based on deep learning according to the present invention uses an asynchronous sub-region optimization method to update different regions of the network separately in response to the gradient difference between the enhancement loss and the speaker loss, effectively solving the conflict of optimization directions. In addition, a feature connection technique is introduced in the joint training, connecting the original features and the enhanced features in the channel dimension, which not only retains the fine structure of the original signal but also learns denoising from the enhanced features, avoiding artifacts and distortions in high signal-to-noise ratio scenarios. This optimization measure not only reduces the load on the network entity but also improves the robustness and recognition accuracy of the system in complex environments, and at the same time ensures that the system meets the real-time requirements through asynchronous updates and feature connection techniques.
[0038] 4. The method for constructing and recognizing a voiceprint recognition model based on deep learning according to the present invention provides a voiceprint recognition system that can still maintain high accuracy in reverberant and noisy environments. The system has high computational performance and low resource occupancy. After optimization, the system can operate efficiently in practical applications, especially suitable for resource-constrained devices and scenarios that require fast and accurate identity verification. For example, mobile devices or embedded systems, and can achieve end-to-end voiceprint recognition and speech recognition on platforms with low computing power, while ensuring that the feedback time is less than 1 second. This makes the system very suitable for application scenarios that require fast and accurate identity verification, such as mobile device unlocking, remote identity verification, security access control systems, etc. It can be implemented on a variety of devices, including mobile devices and embedded systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The following will combine the embodiments of the present invention and the appended Figure 1, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] Embodiment 1.
[0042] A method for constructing a voiceprint recognition system model based on deep learning, characterized by comprising the following steps:
[0043] S1: Input the original voice signal, and output a spectrogram through noise augmentation and preprocessing.
[0044] First, perform noise augmentation. Randomly select noise signals from the noise database, such as the MUSAN and VOiCEs databases, and randomly select reverberation signals from the reverberation database, such as selecting street noise and café noise types. Combine the noise signals and the reverberation signals to generate noise samples with different reverberation characteristics, and superimpose the noise samples on the original voice signal to generate noisy voice data. Then, preprocess the noisy voice data, including splitting the voice signal into frames of a fixed length, usually the frame length is 20 - 30 milliseconds, performing windowing on each frame signal, and common window functions include Hamming window, Hanning window, etc., to reduce the discontinuity at the frame edges, and convert the preprocessed voice signal into a spectrogram.
[0045] S2: Pass the spectrogram through multi-level feature extraction of shallow physical features, middle-level vocal tract features, and deep embedding features in sequence to generate the first speaker embedding feature. Among them, the shallow physical feature extraction is to extract the spectrogram, convert the time-domain signal into a frequency-domain signal, calculate the amplitude spectrum of each frame to generate the spectrogram, and at the same time calculate the short-time energy of the voice signal; the middle-level vocal tract feature extraction is to extract the vocal tract position information, infer the vocal tract position information by analyzing the spectral characteristics of the voice signal, calculate the LPC linear prediction coefficients, and extract the formant frequency and bandwidth parameters; the deep embedding feature extraction is to construct a multi-layer neural network structure, including fully connected layers and convolutional layers, introduce non-linear factors using non-linear activation functions, the network learns the feature mapping relationship, and generates the first speaker embedding feature at the output layer of the network.
[0046] S3: Input the first speaker embedding feature into a hybrid neural network, extract local feature information and temporal dependence relationship, and generate a global voice feature representation. Specifically, input the first speaker embedding feature into a hybrid neural network. The convolutional neural network extracts local patterns from the spectrogram, identifies features including prosody and formant features, and the recurrent neural network captures the dependence relationship between time steps in the voice signal on this basis to generate a global voice feature representation.
[0047] S4: Pass the speech feature representation through a feature enhancement network and a speaker embedding network. The feature enhancement network performs denoising and dereverberation processing on the input speech features to generate enhanced features, and inputs the enhanced features into the speaker embedding network. Discriminative features are extracted through deep convolutional operations to generate a second speaker embedding feature.
[0048] S5: Conduct model training with joint optimization of the loss function for the second speaker embedding feature. Divide the network into different regions, and coordinate the influence of different loss functions on network optimization through an asynchronous sub-region optimization method. Optimize the enhanced network layer to obtain network parameters that minimize the denoising error, and optimize the speaker embedding network layer to obtain network parameters that maximize the identity discrimination degree.
[0049] A deep learning-based voiceprint recognition method using the above voiceprint recognition model includes: using corresponding optimization parameters and performing the following steps:
[0050] S1. Speech signal preprocessing: Convert the original speech signal into an input format that can be processed by the model, and map it end-to-end to a speaker label; the original speech signal includes: phrases, keywords, or continuous speech segments.
[0051] S2. Deep feature extraction network: Perform multi-level feature extraction on the speech signal through a deep neural network to generate a speaker embedding feature.
[0052] S3. Match the speaker embedding feature with the identity label in the pre-stored database and output the user identity authentication result.
[0053] Among them, the multi-level feature extraction includes: shallow physical features, middle vocal tract features, and deep embedding features.
[0054] Embodiment 2.
[0055] A method for constructing a deep learning-based voiceprint recognition system model, which is characterized by including the following steps:
[0056] S1: Input the original speech signal, and output a spectrogram through noise augmentation and preprocessing.
[0057] First, perform preprocessing, at least perform frame segmentation and windowing operations on the speech signal, divide the speech signal into frames of a fixed length, perform windowing processing on each frame of the signal, and then convert the preprocessed speech signal into a spectrogram. Then, perform noise augmentation. Randomly select noise signals from the noise database and randomly select reverberation signals from the reverberation database; combine the noise signals and the reverberation signals to generate noise samples with different reverberation characteristics; inject the noise samples into the spectrogram to generate noisy speech data, and convert the noisy speech data into a new spectrogram.
[0058] S2: Sequentially perform multi-level feature extraction on the new spectrogram through shallow physical features, middle-level vocal tract features, and deep embedding features to generate the first speaker embedding feature. Among them, the shallow physical feature extraction is to extract the spectrogram, which shows the distribution of the speech signal in the frequency domain. The time-domain signal is converted into the frequency-domain signal through the short-time Fourier transform, and the amplitude spectrum of each frame is calculated to generate the spectrogram. At the same time, the short-time energy of the speech signal is calculated. The middle-level vocal tract feature extraction is to extract the vocal tract position information. Different vocal tract positions generate different speech features, such as the larynx, oral cavity, nasal cavity, etc. And the vocal tract position information is inferred by analyzing the spectral characteristics of the speech signal, and methods such as linear predictive coding are used to estimate the formant frequency and bandwidth. The deep embedding feature extraction is to construct a multi-layer neural network structure, including fully connected layers and convolutional layers, to further process and extract the shallow physical features and the middle-level vocal tract features. Nonlinear activation functions, such as ReLU, Sigmoid, etc., are used to introduce nonlinear factors, so that the network can learn complex feature mapping relationships, and a fixed-dimensional first speaker embedding feature is generated at the output layer of the network.
[0059] S3: Input the first speaker embedding feature into a hybrid neural network, a recurrent neural network, or a long short-term memory network to capture the dependencies between time steps in the speech signal based on the spectrogram and generate a global speech feature representation.
[0060] S4: Pass the speech feature representation through a feature enhancement network and a speaker embedding network. The feature enhancement network denoises the input speech features to generate enhanced features, and the enhanced features are input into the speaker embedding network. Distinguishing features are extracted through deep convolutional operations to generate the second speaker embedding feature. The discriminative ability of the second speaker embedding feature is greater than that of the first speaker embedding feature.
[0061] S5: Perform model training with joint optimization of the loss function on the second speaker embedding feature. Divide the network into different regions, and coordinate the influence of different loss functions on network optimization through an asynchronous sub-region optimization method. Optimize the enhanced network layer to obtain network parameters that minimize the denoising error, and optimize the speaker embedding network layer to obtain network parameters that maximize the identity discrimination.
[0062] A deep learning-based voiceprint recognition method using the above voiceprint recognition model includes: using corresponding optimized parameters and performing the following steps:
[0063] S1. Speech signal preprocessing: Convert the original speech signal into an input format that can be processed by the model, and map it end-to-end to the speaker label; the original speech signal includes: phrases, keywords, or continuous speech segments.
[0064] S2. Deep feature extraction network: Perform multi-level feature extraction on the speech signal through a deep neural network to generate speaker embedding features.
[0065] S3. Match the speaker embedding features with the identity labels in the pre-stored database and output the user authentication result.
[0066] Among them, the multi-level feature extraction includes: shallow physical features, middle-level vocal tract features, and deep embedding features.
[0067] Embodiment III.
[0068] A method for constructing a voiceprint recognition system model based on deep learning, characterized by including the following steps:
[0069] S1: Input the original speech signal, and output a spectrogram through noise augmentation and preprocessing.
[0070] First, perform noise augmentation. Randomly select noise signals from the noise database, and these noise signals cover various common background noise types, such as street noise, café noise, mechanical noise, etc. Randomly select reverberation signals from the reverberation database to simulate the noise interference in the actual complex environment. Combine the noise signals and the reverberation signals to generate noise samples with different reverberation characteristics. Superimpose the noise samples on the original speech signal to generate noisy speech data, which not only avoids overfitting in training but also improves the model's adaptability to complex factors such as background noise and echo. Then, preprocess the noisy speech data, including splitting the noisy speech data into frames of a fixed length, performing windowing processing on each frame signal, converting the preprocessed speech signal into a spectrogram, and then normalizing the spectrogram so that its amplitude values are within a certain range.
[0071] S2: Pass the normalized spectrogram through multi-level feature extraction of shallow physical features, middle-level vocal tract features, and deep embedding features in sequence to generate the first speaker embedding feature. Among them, the shallow physical feature extraction is to extract the spectrogram, convert the time-domain signal into a frequency-domain signal, calculate the amplitude spectrum of each frame to generate the spectrogram, and at the same time calculate the short-time energy of the speech signal; the middle-level vocal tract feature extraction is to extract the vocal tract position information, infer the vocal tract position information by analyzing the spectral characteristics of the speech signal, calculate the LPC linear prediction coefficients, and extract the formant frequency and bandwidth parameters; the deep embedding feature extraction is to construct a multi-layer neural network structure, including fully connected layers and convolutional layers, introduce non-linear factors using non-linear activation functions, the network learns the feature mapping relationship, and generate the first speaker embedding feature at the output layer of the network. These features not only have high discrimination but also can effectively cope with the changes in noise and speech content.
[0072] S3: Input the first speaker embedding feature into the hybrid neural network. The convolutional neural network extracts local patterns from the spectrogram and identifies features including prosody and formant through convolution operations. The specific steps include: constructing a multi-layer convolutional neural network, with each layer containing multiple convolutional kernels for extracting local features of different scales and directions; performing a convolution operation on the spectrogram by sliding the convolutional kernel on the spectrogram and calculating the convolution values to obtain a feature map; applying a pooling operation, such as max pooling or average pooling, to downsample the feature map, reducing the feature dimension while retaining important feature information; using the feature map after convolution and pooling operations as the input for subsequent modules to provide local pattern information for the feature representation of the speech signal.
[0073] Based on this, the recurrent neural network or long short-term memory network captures the dependencies between time steps in the speech signal to generate a global speech feature representation. The specific steps include: constructing a recurrent neural network or long short-term memory network. The recurrent neural network can process sequential data through a recurrent structure to capture the forward and backward dependencies in the time series; the long short-term memory introduces a gating mechanism on the basis of the recurrent neural network, which can effectively solve the long-term dependency problem and avoid the phenomena of gradient vanishing and gradient explosion. Input the feature sequence extracted by the convolutional neural module into the recurrent neural network or long short-term memory network. The network updates the hidden state at each time step and gradually learns the time dependency of the speech signal.
[0074] Output the global speech feature representation at the last time step of the network. This feature representation integrates the information of the speech signal in the time dimension and provides global feature support for subsequent speaker recognition. By combining these networks, the system can extract discriminative features for the speaker in both the time series and frequency domains.
[0075] S4: To further improve the discriminability of the features, pass the speech feature representation through a feature enhancement network combined with a bidirectional long short-term memory network and a speaker embedding network based on a residual network. The feature enhancement network performs dereverberation processing on the input speech features to generate enhanced features, and inputs the enhanced features into the speaker embedding network. Discriminative features are extracted through deep convolutional operations to generate the second speaker embedding feature. The specific steps include: constructing a bidirectional long short-term memory network as the feature enhancement module. The bidirectional long short-term memory network contains a forward and a backward long short-term memory network. The forward long short-term memory network processes the input sequence in chronological order to capture the temporal dependencies from front to back; the backward long short-term memory network processes the input sequence in reverse chronological order to capture the temporal dependencies from back to front. By combining the information from both directions, the enhancement network can more comprehensively capture the time dependencies in the speech signal, perform denoising and dereverberation processing on the input speech features, and generate clearer and more robust speech features.
[0076] Construct a speaker embedding network based on the residual network. By introducing residual connections in the residual network, the problem of gradient vanishing and gradient explosion during the training of deep neural networks is solved, enabling the construction of a deeper network structure to extract more complex features. The enhanced features output by the residual network are input into the residual network. Through multi-layer convolutional operations and residual connections, the residual network further extracts the discriminative features in the speech signal, and finally generates the second speaker embedding feature with high discriminative ability for speaker identity recognition.
[0077] In the speaker recognition task, the differences between samples may be very subtle, making it difficult for traditional loss functions to effectively distinguish the speech features of different speakers. Therefore, loss functions specifically designed to improve discrimination, such as angular-softmax and Triplet Loss, are adopted. Triplet Loss enhances the discriminative ability of the model by optimizing the distance between positive and negative samples. Specifically, given an anchor sample, the distance between it and the positive sample of the same speaker should be less than the distance between it and the negative sample of a different speaker, thereby enhancing the model's sensitivity to the subtle differences in speech details.
[0078] S5: Conduct model training with joint optimization of the loss function for the second speaker embedding feature. Divide the network into different regions and coordinate the influence of different loss functions on network optimization through an asynchronous sub-region optimization method to output speech features. Use the mean squared error loss function to optimize the enhancement network layer, focusing on denoising and enhancing speech features, and obtain the network parameters that minimize the denoising error, enabling the model to effectively remove noise and retain valid speaker information, generating clearer and more robust speech features. Use the Softmax or Angular-Softmax loss function to optimize the speaker embedding network layer to ensure that the tasks at all levels of the network can be fully optimized and obtain the network parameters that maximize the identity discrimination. The Softmax loss function is used for multi-classification tasks and measures the error by calculating the cross-entropy between the probability distribution output by the model and the true label; the Angular-Softmax loss function further improves the discrimination ability of the model by optimizing the angular information. During the optimization process, adjust the parameters of the speaker embedding network so that the model can learn more discriminative second speaker embedding features and improve the accuracy of speaker identity recognition.
[0079] In order to effectively coordinate the influence of different loss functions on network optimization during the joint training process, this solution adopts an asynchronous sub-region optimization method. This method divides the network into different regions and optimizes them using different loss functions according to task requirements. Through this method, different loss functions can act on different layers of the network respectively, avoiding conflicts in the optimization direction and improving the stability and performance of the model. During the training process, by optimizing the TripletLoss function and adjusting the model parameters, the model can learn more discriminative feature representations, improve the sensitivity to speech detail differences, and thus enhance the discriminative ability of the model. By optimizing the Angular-Softmax loss function and adjusting the model parameters, the model can learn more discriminative feature representations, improve the discrimination ability in classification tasks, and further optimize the performance of the model.
[0080] Subsequently, during the testing process, the model is verified on an independent test dataset that covers various noise environments and different speakers to truly reflect the performance of the model in actual applications. According to the test results, the model structure or optimization strategy is further adjusted to improve the generalization ability and adaptability of the model, ensuring that the system can stably and accurately complete the speaker recognition task in various complex environments.
[0081] A deep learning-based speaker recognition method using the above-mentioned speaker recognition model includes: using corresponding optimization parameters and performing the following steps:
[0082] S1. Speech signal preprocessing: Convert the original speech signal into an input format that can be processed by the model and map it end-to-end to a speaker label; the original speech signal includes: phrases, keywords, or continuous speech segments.
[0083] S2. Deep feature extraction network: Perform multi-level feature extraction on the speech signal through a deep neural network to generate speaker embedding features.
[0084] S3. Match the speaker embedding features with the identity labels in the pre-stored database and output the user authentication result.
[0085] Among them, the multi-level feature extraction includes: shallow physical features, middle-level vocal tract features, and deep embedding features.
[0086] The above are only embodiments of the present invention. The selection of the embodiment solutions is only for better understanding the content of the invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, is equally included in the patent protection scope of the present invention.
Claims
1. A method for constructing a voiceprint recognition model based on deep learning, characterized in that: The following steps are involved: S1, inputs the original speech signal, and outputs the spectrogram through noise augmentation and preprocessing; S2, extracting the spectrogram in turn through a multi-level feature extraction process of shallow physical features, middle vocal tract features, and deep embedding features to generate a first embedding feature of the speaker; S3, inputting the first embedded feature of the speaker into a hybrid neural network, extracting local feature information and temporal dependency, and generating a global speech feature representation; S4, the speech feature representation is passed through a feature enhancement network and a speaker embedding network to generate a second speaker embedding feature; S5, performing model training of joint optimization of loss functions on the second embedding features of the speaker, dividing the network into different regions, coordinating the effects of different loss functions on network optimization through an asynchronous sub-region optimization method, optimizing the enhanced network layer to obtain network parameters that minimize denoising errors, and optimizing the speaker embedding network layer to obtain network parameters that maximize identity differentiation.
2. The method for constructing a voiceprint recognition model based on deep learning according to claim 1, characterized in that: The noise augmentation in step S1 includes: A noise signal is randomly selected from a noise database, and a reverberation signal is randomly selected from a reverberation database; the noise signal is combined with the reverberation signal to generate noise samples with different reverberation characteristics; and the noise sample is superimposed with an original speech signal to generate noisy speech data.
3. The method for constructing a voiceprint recognition model based on deep learning according to any one of claims 1 or 2, characterized in that: The preprocessing in step S1 includes: performing at least frame division and windowing operations on the original speech signal or the noisy speech data after noise augmentation; The speech signal is divided into frames of fixed length, and each frame signal is windowed; the preprocessed speech signal is converted into a spectrogram.
4. The method for constructing a voiceprint recognition model based on deep learning according to claim 1, characterized in that: The shallow physical feature extraction in step S2 includes: extracting the spectrogram, converting the time domain signal into a frequency domain signal, calculating the amplitude spectrum of each frame to generate a spectrogram, and calculating the short-time energy of the speech signal.
5. The method for constructing a voiceprint recognition model based on deep learning according to claim 1, characterized in that: The middle vocal tract feature extraction in step S2 includes: extracting the vocalization position information, inferring the vocalization position information by analyzing the spectral characteristics of the speech signal, calculating the LPC linear prediction coefficient, and extracting the formant frequency and bandwidth parameters.
6. The method for constructing a voiceprint recognition model based on deep learning according to claim 1, characterized in that: The deep embedding feature extraction in step S2 includes: constructing a multi-layer neural network structure, using a nonlinear activation function to introduce nonlinear factors, the network learning feature mapping relationship, and generating the first embedding feature of the speaker in the output layer of the network; The multi-layer neural network structure includes: a fully connected layer and a convolutional layer.
7. The method for constructing a voiceprint recognition model based on deep learning according to claim 1, characterized in that: The S3 steps include: S31, a convolutional neural network extracts local patterns from the spectrogram to identify features including rhythm and formant; and / or S32, a recurrent neural network or a long short-term memory network captures the dependencies between time steps in the speech signal based on step S31 and generates a global speech feature representation.
8. The method for constructing a voiceprint recognition model based on deep learning according to claim 1, characterized in that: The S4 step includes: S41, the feature enhancement network performs denoising and dereverberation processing on the input speech features to generate enhanced features; S42, inputting the enhanced features into the speaker embedding network, extracting the distinguishing features through deep convolution operations, and generating the second speaker embedding features.
9. The method for constructing a voiceprint recognition model based on deep learning according to any one of claims 1 or 8, characterized in that: The second speaker embedded feature has a greater discriminative ability than the first speaker embedded feature.
10. The voiceprint recognition method based on deep learning is characterized by: The method comprises: using the voiceprint recognition model according to claim 1, performing the following steps: S1. Speech signal preprocessing: converting the original speech signal into an input format that can be processed by the model, and mapping it to the speaker label end-to-end; the original speech signal includes: phrases, keywords or continuous speech segments; S2. Deep feature extraction network: performing multi-level feature extraction on the speech signal through a deep neural network to generate speaker embedding features; S3. Match the speaker embedded features with the identity tags pre-stored in the database and output the user identity authentication result; The multi-level feature extraction includes: shallow physical features, mid-level vocal tract features and deep embedding features.
Citation Information
Patent Citations
Combined learning method and apparatus using deepening neural network based feature enhancement and modified loss function for speaker recognition robust to noisy environments
KR102294638B1
Combined learning method and apparatus using deepening neural network based feature enhancement and modified loss function for speaker recognition robust to noisy environments
US12067989B2
Combined learning method and device using transformed loss function and feature enhancement based on deep neural network for speaker recognition that is robust in noisy environment
WO2020204525A1