Multi-language voiceprint recognition method based on pre-training voice model

By integrating the pre-trained speech models WavLM and ECAPA-TDNN, combined with multilayer perceptron and attention mechanisms, the problem of scarce multilingual voiceprint recognition data is solved, achieving high efficiency and accuracy in multilingual voiceprint recognition and adapting to identity recognition in different language environments.

CN120913569APending Publication Date: 2025-11-07BEIJING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510095338.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Most existing voiceprint recognition technologies only support a single language and cannot meet the needs of multilingual environments. Furthermore, the lack of multilingual voiceprint recognition datasets results in poor robustness and generalization ability of the models, making it difficult to extract the global commonalities and local individual features of speakers from different languages.

Method used

We employ a fusion approach that combines the pre-trained speech model WavLM with the traditional speaker recognition model ECAPA-TDNN. This approach integrates convolutional encoders and transformer encoders to extract deep features, introduces a multilayer perceptron module and an attention mechanism, and optimizes the model parameters to adapt to the characteristics of different languages ​​through fine-tuning strategies on multilingual datasets.

Benefits of technology

It improves the accuracy and robustness of multilingual voiceprint recognition, enabling more effective identification of different speakers and facilitating simple and efficient identity recognition for multilingual populations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913569A_ABST
    Figure CN120913569A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-language voiceprint recognition method based on a pre-training voice model. The pre-training voice model WavLM and a traditional voiceprint recognition model ECAPA-TDNN are fused. According to the method, a multi-layer perceptron (MLP) module is introduced for further refining and converting the features extracted by the WavLM, so that the features are more suitable for the input requirement of an ECAPA-TDNN model, and the abstraction and expression ability of the model to the features is enhanced. In the aspect of multilingual voiceprint recognition, the model is finely adjusted by using a small amount of voice data sets, and the method comprises the following basic steps of: firstly, freezing parameters of a pre-trained voice model WavLM, so that the pre-trained voice model WavLM keeps learned knowledge; and then, parameters of the MLP module and the ECAPA-TDNN model are continuously adjusted in training, so that the multilingual voiceprint recognition capability is learned. In the application, a to-be-recognized voice passes through the fusion model to obtain a feature vector, and after the vector is subjected to judgment and decision making, a voiceprint recognition result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to deep learning, transfer learning, model fusion, fine-tuning strategy, in particular to a multilingual voiceprint recognition method based on a pre-trained speech model. BACKGROUND

[0002] Voiceprint recognition technology research began in the 1960s, which is a technology that identifies the identity of the speaker according to the voice waveform, and has the advantages of convenience, security and low cost. After years of development, voiceprint recognition technology has made great progress and has been applied in identity authentication, criminal investigation, customer service, smart home and other fields. Voiceprint recognition technology generally includes four main parts: audio feature extraction, model construction, model training and scoring decision. Traditional voiceprint recognition methods (such as GMM probability model and MFCC feature-based model) are simple and easy to implement, but the recognition accuracy is not high. Deep learning methods such as ECAPA-TDNN capture the timing information in the speech signal through a time delay neural network (TDNN), and introduce a channel attention mechanism and a compression activation (SE) module to enhance the model's ability to perceive speech features. On the other hand, pre-trained speech models such as WavLM are trained on a large amount of unsupervised data through the Denoising Masked Speech Modeling framework, and have shown excellent performance on multiple speech tasks, including speaker recognition. However, existing voiceprint recognition technologies mostly support only a single language, and cannot meet the needs of a multilingual environment. The research on multilingual voiceprint recognition technology is still in its infancy and faces challenges including: (1) Lack of multilingual voiceprint recognition datasets, which limits model training and optimization and results in poor robustness and generalization ability of the model. (2) Differences between voiceprint features of different languages, such as differences caused by different language phonemes, differences in speech rhythm, and differences in pronunciation habits. Therefore, how to extract the global common features and local individual features of voiceprints of speakers of different languages is one of the important challenges of multilingual voiceprint recognition technology.

[0003] To this end, the present application proposes a voiceprint recognition method based on a pre-trained speech model. In terms of improving voiceprint recognition technology, the method uses a pre-trained model WavLM trained on nearly 100,000 hours of English dataset and fuses it with the traditional voiceprint recognition model ECAPA-TDNN. WavLM has a great advantage in extracting speech features, and inputting the features extracted by it into ECAPA-TDNN for voiceprint recognition can improve its recognition performance. In terms of multilingual recognition, to make the pre-trained model perform well on speech, we fine-tune it with a small amount of speech to make it learn the voiceprint features of the speaker more accurately. SUMMARY

[0004] The technical problem solved by the present application is to provide a voiceprint recognition method based on a pre-trained speech model, which can accurately recognize different pronunciations of speakers, thereby more effectively helping multilingual people to perform simple and efficient identity recognition, and promoting the convenience of life.

[0005] A multilingual voiceprint recognition method based on a pre-trained speech model, comprising the following steps:

[0006] Step (1), integration of an audio feature extractor;

[0007] Step (2), fusion architecture of WavLM and ECAPA-TDNN model;

[0008] Step (3), fine-tuning strategy on a multilingual dataset.

[0009] Further, in the step (1), the audio extractor comprises:

[0010] Audio feature representation of WavLM;

[0011] WavLM extracts deep features from input audio signals by combining a convolutional encoder CNN Encoder and a transformer encoder Transformer Encoder, forming an audio feature representation; the convolutional encoder uses multiple types of convolutional networks and is supplemented with regularization techniques to enhance the learning ability of the model; the transformer encoder is based on the Transformer architecture, and each encoding layer contains a multi-head attention mechanism and a feedforward neural network, wherein the self-attention mechanism enables the model to update the element representation of each position according to all other positions in the sequence, and the extracted features not only contain speech content but also incorporate context information;

[0012] The specific implementation of the convolutional encoder and the transformer encoder is as follows:

[0013] Convolutional encoder;

[0014] The convolutional encoder in WavLM functions similarly to a cochlear filter bank, processing 25ms of speech segments with a step size of 20ms, thereby extracting preliminary audio signal features; the convolutional encoder is composed of seven layers of time convolution layers, each followed by a normalization layer and a GELU activation layer; the time convolution layer uses 512 channels, with a step size of (5, 2, 2, 2, 2, 2) and a kernel width of (10, 3, 3, 2, 2, 2, 2); the time convolution layer uses one-dimensional or two-dimensional convolution to process the input signal, filtering the input data through a sliding window to extract local features; the output of each time convolution layer can be calculated by the following formula:

[0015]

[0016] where y[i] is the output feature map, x is the input signal, w is the convolution kernel, and b is the bias term.

[0017] The transformer encoder; the output of the convolutional encoder is taken as the input of the transformer encoder; the transformer encoder is composed of multiple Transformer layers, each Transformer block follows the design of HuBERT, with 16 attention heads and a hidden dimension of 768 or 1024; the transformer encoder also models local information using relative position bias, thereby improving the model's ability to capture short-term dependencies in speech signals; each layer of the transformer encoder processes the output of the previous layer, and each layer of the output is represented as H i = [h1, h2, …, h L L] where h L represents the output of the Lth layer; the core of this encoder is the self-attention mechanism, which allows the elements of each position to update their representation according to all other positions in the sequence, so that the extracted features contain speech content and contextual information, and its calculation is described as:

[0018]

[0019] where Q is the query matrix, K is the K key matrix, V is the value matrix, d k is the dimension of the key vector, is used to scale the dot product to prevent the problem of gradient disappearance, and Softmax is a normalization function that ensures the sum of the attention weights of the output is 1;

[0020] The multi-head attention (Multi-HeadAttention) in the self-attention mechanism is represented as:

[0021] MultiHead(Q, K, V) = Convat(head1, …, head h )W O (3)

[0022] where is the output of the i-th attention head, is a linear transformation matrix used to map the inputs Q, K, and V to different representation spaces, respectively, W O is a linear transformation matrix that finally combines the outputs of multiple heads back to the original dimension, is the attention score matrix, and Concat represents concatenating the outputs of multiple heads together.

[0023] By combining convolutional encoders and transformer encoders, the model's feature extraction capabilities are enhanced; convolutional encoders excel at capturing local speaker features, while transformer encoders help the model understand these features in a broader context; WavLM can effectively extract deep global common features from audio signals that contain rich speech and contextual information, including not only acoustic features, but also language features and speaker features.

[0024] Audio feature representation of ECAPA-TDNN;

[0025] Channel and context dependency pooling;

[0026] Channel and context-dependent pooling is a step to enhance speech feature representation by introducing an attention mechanism to improve statistical pooling. The attention mechanism is applied temporally and spatially as follows: for each time frame, an attention weight is calculated, reflecting the importance of that frame in the current context; the importance of different channels is calculated by weighted averaging and standard deviation for each time frame of each channel; pooling using a weighted matrix allows the model to focus more on important frames and channels, thereby extracting more accurate and robust speaker features; the attention mechanism is channel-dependent, as shown in the following formula:

[0027]

[0028] Where h t This represents the output at time step t, where parameter W ∈ R. R*C and b∈R R*1 The information from self-attention is projected into a smaller R-dimensional representation, which is shared across all C channels to reduce the risk of parameter counting and overfitting; after a nonlinear transformation f(·), it is weighted by v c ∈R R*1 The sum and deviation is k c The linear layer is converted into a channel-dependent self-attention score; then the softmax function is applied to normalize this self-attention score over all frames, as shown in the following formula:

[0029]

[0030] Among them, the self-attention score a t,c This represents the importance of each frame in a given channel, used to calculate the weighted statistics for channel c; for each speech segment, the weighted average vector... channel components The estimate is:

[0031]

[0032] Weighted standard deviation vector channel components is constructed as follows:

[0033]

[0034] The weighted mean and the weighted standard deviation are spliced to obtain the final output of the pooling layer;

[0035] SE-Res2Blocks;

[0036] SE-Res2Blocks combines the structural features of SE blocks and Res2Blocks;

[0037] SE block:

[0038] It includes Squeeze and Excitation operations, where the Squeeze operation compresses the features through global average pooling to generate a descriptor for each channel; the specific formula is:

[0039]

[0040] The Excitation operation uses z to calculate the weight of each channel, and the formula is as follows:

[0041] s = σ (W2f (W1z + b1) + b2) (9)

[0042] Where σ represents the sigmoid function, f represents the nonlinear activation function, W1∈R R×C and W2∈R C×R , C and R represent the number of input channels and the dimensionality reduction number respectively; the SE block applies the weight vector obtained by the Excitation operation to the original input through channel multiplication to obtain the final output feature map; this step actually recalibrates the original input feature map, so that the features of each channel are weighted according to their importance; the formula involved is as follows:

[0043]

[0044] Res2Blocks:

[0045] It processes multi-scale features by constructing hierarchical residual connections to expand the temporal context; the Res2Block module in SE-Res2Block has 8 scales, which expand the temporal context through kernel size and dilation interval;

[0046] SE-Res2Blocks:

[0047] The SE block is embedded into the Res2Blocks to form a residual block with stronger representation capability, i.e. SE-Res2Blocks; Specifically, SE-Res2Blocks contains extended convolution, front and back dense layers, and an SE block; The first dense layer is used to reduce the feature dimension, the second dense layer restores the feature quantity to the original dimension, and the SE module is used to expand each channel; The whole unit is covered by a jump connection, so as to retain the original input information and promote the flow of gradient;

[0048] Multi-layer feature aggregation:

[0049] For each frame, the ECAPA-TDNN model concatenates the output features of all SE-Res2Blocks; After this multi-layer feature aggregation, a dense layer processes the connected information to generate features for attention statistical pooling; A complementary method for using multi-layer information is to use the output of all previous SE-Res2Blocks and the initial convolutional layer as the input of each frame layer block; The ECAPA-TDNN model achieves this by defining the residual connection in each SE-Res2Block as the sum of the outputs of all previous blocks.

[0050] Further, in the step (2), a multi-layer perception MLP module is introduced between the WavLM and ECAPA-TDNN models:

[0051] A multi-layer perception MLP module is introduced between the pre-trained model WavLM and the speech recognition model ECAPA-TDNN, which is used to refine and convert the features extracted by WavLM, making it more suitable for the input requirements of the ECAPA-TDNN model; In the entire fusion model, WavLM serves as the primary audio feature extractor, extracting features containing rich speaker content information and audio features; The convolution operation, channel attention mechanism and multi-layer output fusion in the ECAPA-TDNN model can further refine and enhance the features extracted from the WavLM model, so that the model can more accurately identify and distinguish the identities of different speakers.

[0052] Further, the steps of step (3) are as follows:

[0053] The model is fine-tuned on multi-lingual datasets to adapt to the characteristics and challenges of different languages; During the fine-tuning process, the parameters of the pre-trained WavLM model are frozen to ensure that the knowledge learned by the pre-trained model is retained, while the parameters of the MLP module and the ECAPA-TDNN model are adjusted to adapt to the specific voiceprint recognition task; The specific implementation steps are as follows:

[0054] Freeze WavLM pre-trained model parameters:

[0055] Load the pre-trained model parameters and make the parameters of WavLM not updated during training, so that the knowledge learned by the pre-trained model is retained;

[0056] Adjust the parameters of the MLP module and the ECAPA-TDNN model:

[0057] By forward propagating the training data to obtain the speaker voiceprint representation vector, using the loss value and accuracy of the current batch of the loss function, and then back propagating to calculate the gradient, updating the parameters of the MLP module and the ECAPA-TDNN model; with the increase of the training round, the loss value is constantly reduced, the accuracy is constantly improved, and the model parameters are constantly optimized, and the model learns the ability of speaker recognition; a loss function based on angle, AAM-Softmax loss function, is used, which aims to maximize the inter-class distance while minimizing the intra-class distance.

[0058] Further, the specific implementation steps of the AAM-Softmax loss function are as follows:

[0059] Calculate the target output:

[0060]

[0061] where one_hot represents the label matrix of different categories, represents the cosine similarity between the feature vector x and the category center vector weight, represents the cosine value added with the additive angle margin m, to achieve the purpose of expanding the distance between the target category and the non-target category;

[0062] Scale the target output using the scale s:

[0063] output=output*s (12)

[0064] Calculate the final loss using the cross-entropy loss function:

[0065]

[0066] where output i and label i respectively represent the target output and the true target category label of sample i, and log represents the logarithmic function; the advantage of using the logarithmic function is that when the model predicts the probability of the target category is too small, it will output a very high loss, which will more strongly promote the model to further optimize.

[0067] Further, for the voiceprint recognition task, the evaluation mechanism steps are as follows:

[0068] Speaker voiceprint representation extraction:

[0069] Speaker embedding is extracted from the final full connection layer of the fusion model; speaker embedding represents the characteristics of the speaker, which is used for subsequent similarity comparison; specifically, for two new input audios x1, x2 to be judged, first process them into two groups of data, one group is the same audio x 11 21 as the original audio, and the other group is the splicing X 12 22 of 5 audio segments of fixed length randomly cut from the original audio; input x 11 21 12 22 into the fusion model to obtain speaker embedding 11 embedding 21 embedding 12 embedding 22 ;

[0070] Verification score calculation:

[0071] Calculate the score score1 of embedding 11 embedding 21 and the score score2 of embedding 12 embedding 22 using cosine similarity, and the formula is as follows:

[0072]

[0073] Speaker recognition result judgment:

[0074] Finally, take the average of score1 and score2 to obtain the final score score; according to the final score score and a trained threshold, judge whether the two audios come from the same person. BRIEF DESCRIPTION OF DRAWINGS

[0075] Figure 1 is the overall step flowchart of the present application

[0076] Figure 2 is the fusion model architecture diagram of the present application DETAILED DESCRIPTION

[0077] The present application will be described in detail below in combination with the drawings and examples.

[0078] ​​​​​The present application combines the powerful speech representation ability of the pre-trained speech large model WavLM and the perceptual ability of the ECAPA-TDNN model to capture speech features, aiming to propose an efficient and robust multilingual voiceprint recognition method. The technical route of this project is described as follows:

[0079] Integration of audio feature extractor

[0080] The present application proposes a method for integrating audio feature extractors for multilingual speakers. This method utilizes the deep speech representation ability of WavLM pre-trained on large-scale unsupervised datasets to extract deep features from audio signals. Through the WavLM model, we can obtain more rich and abstract audio feature representations, which not only contain the basic content of speech, but also contain the voiceprint features of speakers, providing strong support for voiceprint recognition. In addition, the present application also combines the feature extraction technology in ECAPA-TDNN, further enhancing the model's ability to capture features of speakers of different languages.

[0081] Architecture of fusion model

[0082] In the pre-trained model WavLM and the speech recognition model ECAPA-TDNN, the present application innovatively introduces a multi-layer perceptron (MLP) module to further refine and convert the features extracted by WavLM, making them more suitable for the input requirements of the ECAPA-TDNN model, and enhancing the model's ability to abstract and express features.

[0083] Fine-tuning strategy on multilingual dataset

[0084] The present application fine-tunes the model on a multilingual dataset to adapt to the characteristics and challenges of different languages. During the fine-tuning process, the parameters of the pre-trained WavLM model are frozen to ensure that the knowledge learned by the pre-trained model is retained, while the parameters of the MLP module and the ECAPA-TDNN model are adjusted to adapt to specific voiceprint recognition tasks. This strategy not only improves the generalization ability of the model, but also ensures the recognition performance on different languages.

[0085] Integration of audio feature extractor

[0086] The application provides a multi-lingual speaker audio feature extraction method, which utilizes the deep speech representation capability of WavLM pre-trained on a large-scale unsupervised dataset to perform deep feature extraction on audio signals. First, using WavLM, we can obtain more rich and abstract audio feature representations, which not only contain the basic content of speech, but also contain the speaker's voiceprint features, providing strong support for voiceprint recognition. In addition, the application also combines the feature extraction technology in ECAPA-TDNN to further enhance the model's ability to capture local individual features of speakers of different languages. The specific implementation is as follows:

[0087] Audio feature representation of WavLM

[0088] WavLM extracts deep features from input audio signals by combining a convolutional encoder (CNN Encoder) and a transformer encoder (Transformer Encoder) to form an audio feature representation. The convolutional encoder uses multiple types of convolutional networks and is supplemented with regularization techniques such as layer normalization and group normalization to enhance the learning ability of the model. The transformer encoder is based on the Transformer architecture, and each encoding layer contains a multi-head attention mechanism and a feedforward neural network. The self-attention mechanism in it allows the model to update the element representation of each position based on all other positions in the sequence, so that the extracted features not only contain speech content but also incorporate contextual information. The specific implementation of the convolutional encoder and the transformer encoder is as follows:

[0089] (1) Convolutional Encoder (CNN Encoder)

[0090] The role of the convolutional encoder in WavLM is similar to that of a cochlear filter bank, which processes 25ms of speech segments with a step size of 20ms to extract preliminary audio signal features. The encoder consists of seven temporal convolutional layers, each followed by a layer normalization layer and a GELU activation layer. These convolutional layers use 512 channels, a step size of (5, 2, 2, 2, 2, 2), and a kernel width of (10, 3, 3, 2, 2, 2, 2). The temporal convolutional layers use one-dimensional or two-dimensional convolution to process the input signal, filtering the input data through a sliding window to extract local features. The output of each temporal convolutional layer can be calculated by the following formula:

[0091]

[0092] where y[i] is the output feature map, x is the input signal, w is the convolution kernel, and b is the bias term.

[0093] (2) Transformer Encoder

[0094] The output of the convolutional encoder serves as the input of the transformer encoder. The transformer encoder consists of multiple Transformer layers, each of which follows the design of HuBERT with 16 attention heads and a hidden dimension of 768 (for the base model) or 1024 (for the large model). The transformer encoder also utilizes relative position bias (including gated relative position bias) to better model local information, thus improving the model's ability to capture short-term dependencies of the speech signal. Each layer of the transformer encoder processes the output of the previous layer, and each layer of the output is represented as H i = [hi, h2,..., h L L-1] where hi L represents the output of the Lth layer (the previous layer of the ith layer). The core of this encoder is the self-attention mechanism, which allows the elements of each position to update their representations according to all other positions in the sequence, so that the extracted features contain both speech content and contextual information, and its calculation can be described as:

[0095]

[0096] where Q is the Query matrix, K is the Key matrix, and V is the Value matrix, d k is the dimension of the key vector, is used to scale the dot product to prevent the problem of gradient disappearance, and Softmax is a normalization function that ensures the sum of the attention weights of the output is 1.

[0097] The multi-head attention in the self-attention mechanism can be represented as:

[0098] MultiHead(Q, K, V) = Concat(head1,..., head h ) W O (3)

[0099] where is the output of the ith attention head, is a linear transformation matrix used to map the inputs Q, K, v to different representation spaces, respectively, W O is a linear transformation matrix that finally combines the outputs of multiple heads back to the original dimension, is the attention score matrix, and Concat represents concatenating the outputs of multiple heads together.

[0100] By combining the convolutional encoder and the transformer encoder, the feature extraction capability of the model is enhanced. The convolutional encoder is good at capturing local voiceprint features such as phonemes and syllables, while the transformer encoder helps the model understand these features in a broader context. The convolutional encoder provides the transformer encoder with rich local information, which enables the model to make more accurate judgments in the global context. Therefore, WavLM can effectively extract deep global common features containing rich speech and context information from audio signals, including not only acoustic features such as phonemes and words, but also language features such as word meaning and semantics, and speaker features.

[0101] Audio feature representation of ECAPA-TDNN

[0102] (1) Channel and context-dependent pooling (Attentive Stat Pooling+BN)

[0103] Channel and context-dependent pooling is a step to enhance speech feature representation by introducing an attention mechanism to improve traditional statistical pooling. It not only considers the importance of each channel (or feature dimension), but also considers the importance of each frame in the time dimension (i.e. context). The core of this step is the use of attention mechanism in time and space. For each time frame, the model calculates an attention weight that reflects the importance of the frame in the current context. Then the model calculates the weighted average and standard deviation of each time frame for each channel to obtain the importance of different channels. Pooling in the weighted manner by the weight matrix enables the model to pay more attention to important frames and channels, thereby extracting more accurate and robust speaker features. Make the attention mechanism dependent on the channel, the formula is as follows:

[0104]

[0105] where h represents the output at time step t, parameter W ∈ R R*C and b ∈ R R*1 Project the information of self-attention into a smaller R-dimensional representation that is shared across all C channels to reduce the risk of overfitting and parameter count. After a nonlinear transformation f(·), this information is converted into channel-dependent self-attention scores by a linear layer with weights v c ∈ R R*1 and bias k c . Then apply the softmax function to normalize the scores across all frames, the formula is as follows:

[0106]

[0107] where the self-attention score a t,cimportance of each frame of a given channel, used to compute the weighted statistics of channel c. For each utterance, the weighted mean vector The channel components of the weighted standard deviation vector are estimated as:

[0108]

[0109] The channel components of the weighted standard deviation vector are estimated as:

[0110]

[0111] The vectors of weighted mean and weighted standard deviation are concatenated to obtain the final output of the pooling layer.

[0112] (2) SE-Res2Blocks SE-Res2Blocks combines the structural features of SE blocks and Res2Blocks. SE blocks aim to enhance the representation ability of the model by explicitly modeling the interdependence between channels and recalibrating channel feature responses. Res2Blocks is an improved residual block that processes multi-scale features by constructing internal hierarchical residual connections, improving performance and reducing model parameters. The specific implementation steps of SE-Res2Blocks are as follows:

[0113] 1) SE block: includes compression operation and activation operation, where the compression operation compresses the features by global average pooling to generate a descriptor for each channel, and the specific formula is:

[0114]

[0115] The activation operation uses z to calculate the weight of each channel, and the formula is as follows:

[0116] s = σ (W2f (W1z + b1) + b2) (9)

[0117] where σ represents the sigmoid function, f represents the nonlinear activation function, W1∈R R×C and W2∈R C×R , C and R represent the number of input channels and the dimensionality reduction number respectively. The SE block applies the weight vector obtained by the activation operation to the original input through channel multiplication to obtain the final output feature map. This step actually recalibrates the original input feature map, so that the features of each channel are weighted according to their importance. The formula involved is as follows:

[0118]

[0119] ​2) Res2Blocks: Process multi-scale features by constructing hierarchical residual connections, thus expanding the temporal context. Res2Block modules in SE-Res2Block have 8 scales, which expand the temporal context by kernel size and dilation stride.

[0120] 3) SE-Res2Blocks: Embed the SE block described above into Res2Blocks to form a residual block with stronger representation capability, i.e. SE-Res2Blocks. Specifically, SE-Res2Blocks contain an expanded convolution, a front and back dense layer, and an SE block. The first dense layer is used to reduce the feature dimension, the second dense layer restores the feature quantity to the original dimension, and the SE module is used to expand each channel. The whole unit is covered by a jump connection, so as to retain the original input information and promote the flow of gradient.

[0121] (3) Multi-layer feature aggregation

[0122] Due to the hierarchical nature of the ECAPA-TDNN model, deeper features are the most complex and should be closely related to the speaker identity. However, shallower feature maps also contribute to obtaining more accurate speaker information. For each frame, the ECAPA-TDNN model concatenates the output features of all SE-Res2Blocks. After this multi-layer feature aggregation, a dense layer (Conv1D+ReLU) processes the connected information to generate features for the attentive statistical pooling (Attentive Stat Pooling+BN). Another complementary method to utilize multi-layer information is to use the output of all previous SE-Res2Blocks and the initial convolutional layer (Conv1D+ReLU+BN) as the input of each frame layer block. The ECAPA-TDNN model achieves this by defining the residual connection in each SE-Res2Block as the sum of the outputs of all previous blocks.

[0123] Architecture of the fusion model and fine-tuning strategy

[0124] The present application proposes a method for extracting the voiceprint representation of a multilingual speaker. This method innovatively fuses the audio feature extraction capabilities of the pre-trained speech model WavLM and the traditional voiceprint recognition model ECAPA-TDNN to efficiently extract the voiceprint representation of a multilingual speaker and improve the accuracy of the voiceprint recognition task.

[0125] (1) Architecture of the fusion model

[0126] Specifically, by introducing the MLP module as an intermediary, the output of the WavLM model is received as the input of the ECAPA-TDNN model, and a fusion model is constructed. In the entire fusion model, WavLM serves as the primary audio feature extractor, extracting content information and audio features that contain rich speaker information; and the convolution operation, channel attention mechanism and multi-layer output fusion in the ECAPA-TDNN model can further refine and enhance the features extracted from the WavLM model, so that the model can more accurately identify and distinguish different speakers' identities.

[0127] (2) Fine-tuning strategy on multi-lingual dataset

[0128] The present application fine-tunes the model on a multi-lingual dataset to adapt to the characteristics and challenges of different languages. During the fine-tuning process, the parameters of the pre-trained WavLM model are frozen to ensure that the knowledge learned by the pre-trained model is retained, while the parameters of the MLP module and the ECAPA-TDNN model are adjusted to adapt to the specific voiceprint recognition task. This strategy not only improves the generalization ability of the model, but also ensures the recognition performance on different languages. The specific implementation steps are as follows:

[0129] 1) Freeze WavLM pre-trained model parameters: load the pre-trained model parameters and make the parameters of WavLM not updated with training, so that the knowledge learned by the pre-trained model is retained.

[0130] 2) Adjust the parameters of the MLP module and the ECAPA-TDNN model: pass the training data forward to obtain the speaker voiceprint representation vector, use the loss function to calculate the loss value and accuracy of the current batch, then perform back propagation to calculate the gradient and update the parameters of the MLP module and the ECAPA-TDNN model. As the number of training rounds increases, the loss value decreases and the accuracy improves, and the model parameters are continuously optimized, and the model learns the ability of speaker recognition. In particular, the present application uses an angle-based loss function, namely the AAM-Softmax loss function, which aims to maximize the inter-class distance while minimizing the intra-class distance. The following are the specific implementation steps of the AAM-Softmax loss function:

[0131] Calculate the target output:

[0132]

[0133] where one_hot represents the label matrix of different categories, represents the cosine similarity between the feature vector x and the class center vector weight, represents the cosine value with an additive angle margin m added to achieve a larger distance between the target class and non-target classes.

[0134] Scale the target output using scale s:

[0135] output = output * s (12)

[0136] Compute the final loss using the cross-entropy loss function:

[0137]

[0138] where output i and label i are the target output and the true target class label of sample i, respectively, and log denotes the logarithm function. The benefit of using the logarithm function is that when the model predicts a very small probability for the target class, it outputs a very high loss, which more strongly encourages the model to further optimize.

[0139] Judge the decision

[0140] For the task of voiceprint recognition, the main steps of the judgment decision taken by the present application are as follows:

[0141] Speaker voiceprint representation extraction: The speaker voiceprint representation is extracted from the final fully connected layer of the fusion model. These speaker voiceprint representations represent the characteristics of the speaker and are used for subsequent similarity comparison. Specifically, for two new input audio segments x1, x2 to be judged, first process them into two groups of data, one group being the same audio as the original audio x 11 , x 21 , and the other group being a concatenation of five audio segments of fixed length randomly cut from the original audio x 12 , x 22 . Input x 11 , x 21 , x 12 , x 22 into the fusion model to obtain the speaker voiceprint representation embedding 11 , embedding 21 , embedding 12 , embedding 22 for each audio segment, respectively.

[0142] Verification score calculation: Use the cosine similarity to calculate the score score1 of embedding 11 , embedding 21 and the score score2 of embedding 12 , embedding 22 , respectively, whose formula is as follows:

[0143]

[0144] Voiceprint recognition result judgment: Finally, score1 and score2 are averaged to obtain a final score score. According to the final score score and a trained threshold, it is judged whether the two audios are from one person.

Claims

1. A multi-lingual voiceprint recognition method based on a pre-trained speech model, characterized in that, The method comprises the following steps: Step (1), integration of an audio feature extractor; Step (2), a fusion architecture of WavLM and ECAPA-TDNN model; Step (3), a fine-tuning strategy on a multi-lingual dataset.

2. The method of claim 1, wherein the pre-trained speech model is a multilingual speech recognition model. In the step (1), the audio extractor comprises: An audio feature representation of WavLM; WavLM extracts deep features from input audio signals by combining a convolutional encoder CNN Encoder and a transformer encoder Transformer Encoder, forming an audio feature representation; the convolutional encoder uses multiple types of convolutional networks and is supplemented with regularization techniques to enhance the learning ability of the model; the transformer encoder is based on the Transformer architecture, and each encoding layer contains a multi-head attention mechanism and a feedforward neural network, in which the self-attention mechanism enables the model to update the element representation of each position according to all other positions in the sequence, and the extracted features contain not only speech content but also contextual information; The specific implementation of the convolutional encoder and the transformer encoder is as follows: Convolutional encoder; The role of the convolutional encoder in WavLM is similar to that of a cochlear filter bank, which processes a 25ms speech segment with a step size of 20ms, thereby extracting preliminary audio signal features; the convolutional encoder consists of seven time convolutional layers, each followed by a normalization layer and a GELU activation layer; the time convolutional layer uses 512 channels, a step size of (5, 2, 2, 2, 2, 2), and a kernel width of (10, 3, 3, 2, 2, 2, 2); the time convolutional layer uses one-dimensional or two-dimensional convolution to process the input signal, filters the input data through a sliding window, and extracts local features; the output of each time convolutional layer can be calculated by the following formula: Where y[i] is the output feature map, x is the input signal, w is the convolution kernel, and b is the bias term; The transformer encoder; the output of the convolutional encoder as the input of the transformer encoder; the transformer encoder consists of multiple Transformer layers, each Transformer block follows the design of HuBERT, with 16 attention heads and a hidden dimension of 768 or 1024; the transformer encoder also models local information using relative position bias, thereby improving the model's ability to capture short-term dependencies of the speech signal; each layer of the transformer encoder processes the output of the previous layer, and each layer of the output is represented as H i = [h1, h2, …, h L ] where h L represents the output of the Lth layer; the core of the encoder is the self-attention mechanism, which allows the elements of each position to update their representation according to all other positions in the sequence, so that the extracted features contain speech content and context information, and its calculation is described as: where Q is the query matrix, K is the key matrix, V is the value matrix, d k is the dimension of the key vector, for scaling the dot product to prevent the vanishing gradient problem, and Softmax is a normalization function that ensures the sum of the output attention weights is 1. The multi-head attention (Multi-HeadAttention) in the self-attention mechanism is represented as: MultiHead(Q, K, V) = Concat(head1,..., head h )W O (3) wherein is the output of the i-th attention head, is a linear transformation matrix for mapping the inputs Q, K, V into different representation spaces, W O is a linear transformation matrix for finally merging the outputs of multiple heads back into the original dimension, is an attention score matrix, and Concat denotes concatenating the outputs of multiple heads together; By combining the convolutional encoder and the transformer encoder, the feature extraction capability of the model is enhanced; the convolutional encoder is good at capturing local voiceprint features, and the transformer encoder helps the model understand these features in a broader context; WavLM can effectively extract deep global common features containing rich speech and contextual information from audio signals, and these features include not only acoustic features but also language features and speaker features; An audio feature representation of ECAPA-TDNN; Channel and context-dependent pooling; Channel and context-dependent pooling is a step to enhance the speech feature representation by introducing an attention mechanism to improve statistical pooling; the use of attention mechanism in time and space is as follows: for each time frame, an attention weight is calculated, which reflects the importance of the frame in the current context; the importance of different channels is calculated by weighted average and standard deviation of each time frame of each channel; the weight matrix weighted pooling makes the model pay more attention to important frames and channels, so as to extract more accurate and robust speaker features; the attention mechanism depends on the channel, and the formula is as follows: where h t denotes the output at time step t, parameters W ∈ R R*C and b ∈ R R*1 project the information of self-attention into a smaller R-dimensional representation that is shared across all C channels to reduce the parameter count and the risk of overfitting; after a non-linear transformation f(·), it is converted into channel-dependent self-attention scores by a linear layer with weights v c ∈ R R*1 and biases k c ; then a softmax function is applied to normalize the self-attention scores across all frames, which is formulated as follows: where the self-attention scores a t,c represent the importance of each frame of a given channel for computing the weighted statistics of the channel c; for each piece of speech, the weighted average vector the channel components of are estimated as: Weighted standard deviation vector Channel components of Are constructed as follows: concatenating the vector of weighted mean and weighted standard deviation to get the final output of the pooling layer; SE-Res2Blocks; SE-Res2Blocks combines the structural characteristics of SE blocks and Res2Blocks; SE block: including Squeeze and Excitation operations, wherein the Squeeze operation compresses the features by global average pooling to generate a descriptor for each channel; the specific formula is as follows: The Excitation operation uses z to calculate the weight of each channel, and the formula is as follows: s=σ(W2f(W1z+b1)+b2) (9) where σ denotes a sigmoid function, f denotes a nonlinear activation function, W1e R R×C and W2e R C×R , C and R denote the number of input channels and the dimension of the reduced dimension, respectively; the SE block applies the weight vector obtained by the Excitation operation to the original input through channel multiplication to obtain the final output feature map; this step actually recalibrates the original input feature map, so that the features of each channel are weighted according to their importance; the formula involved is as follows: Res2Blocks: The hierarchical residual connection is constructed to process multi-scale features, thereby expanding the temporal context; the Res2Block module in the SE-Res2Block has 8 scales, which expands the temporal context through kernel size and dilation interval; SE-Res2Blocks: The above SE block is embedded into the Res2Blocks to form a residual block with stronger representation capability, i.e. SE-Res2Blocks; specifically, SE-Res2Blocks contains extended convolution, front and back dense layers, and an SE block; the first dense layer is used to reduce the feature dimension, the second dense layer restores the feature quantity to the original dimension, and the SE module is used to expand each channel; the whole unit is covered by a jump connection, thereby retaining the original input information and promoting the flow of gradient; Multi-layer feature aggregation: For each frame, the ECAPA-TDNN model concatenates the output features of all SE-Res2Blocks; after this multi-layer feature aggregation, a dense layer processes the connected information to generate features for attention statistical pooling; the supplementary method of using multi-layer information is to use the output of all previous SE-Res2Blocks and the initial convolution layer as the input of each frame layer block; the ECAPA-TDNN model realizes this by defining the residual connection in each SE-Res2Block as the sum of the outputs of all previous blocks.

3. The method of claim 1, wherein the pre-trained speech model is a multilingual speech recognition model. In step (2), a multi-layer perceptron (MLP) module is introduced between the WavLM and ECAPA-TDNN models: A multi-layer perceptron (MLP) module is introduced between the pre-trained model WavLM and the speech recognition model ECAPA-TDNN to refine and convert the features extracted by WavLM, making them more suitable for the input requirements of the ECAPA-TDNN model. In the entire fusion model, WavLM serves as the primary audio feature extractor, extracting features containing rich speaker content information and audio features. The convolution operation, channel attention mechanism, and multi-layer output fusion in the ECAPA-TDNN model can further refine and enhance the features extracted from the WavLM model, enabling the model to more accurately identify and distinguish different speakers.

4. The method of claim 1, wherein the pre-trained speech model is a multilingual speech recognition model. The steps of step (3) are as follows: The model is fine-tuned on a multi-lingual dataset to adapt to the characteristics and challenges of different languages. During the fine-tuning process, the parameters of the pre-trained WavLM model are frozen to ensure that the knowledge learned by the pre-trained model is retained. At the same time, the parameters of the MLP module and the ECAPA-TDNN model are adjusted to adapt to the specific voiceprint recognition task. The specific implementation steps are as follows: Freeze WavLM pre-trained model parameters: Load the pre-trained model parameters and make the parameters of WavLM not updated during training, so that the knowledge learned by the pre-trained model is retained. Adjust the parameters of the MLP module and the ECAPA-TDNN model: By forward propagating the training data to obtain the speaker voiceprint representation vector, using the loss value and accuracy of the current batch of the loss function, and then performing backpropagation to calculate the gradient, the parameters of the MLP module and the ECAPA-TDNN model are updated. As the number of training rounds increases, the loss value decreases, the accuracy improves, and the model parameters are continuously optimized, and the model learns the ability of speaker recognition. An angle-based loss function, AAM-Softmax loss function, is used. This loss function aims to maximize the inter-class distance while minimizing the intra-class distance.

5. The method of claim 4, wherein the pre-trained speech model is a multilingual speech recognition model. Specific implementation steps of AAM-Softmax loss function: Calculate the target output: where one_hot represents a label matrix of different categories, represents the cosine similarity of the feature vector x and the category center vector weight, represents the cosine value added with the additive angle margin m to achieve the distance between the target category and the non-target category. Scale the target output using the scale s: output = output * s (12) Use the cross-entropy loss function to calculate the final loss: where output i and label i represent the target output and the true target class label for sample i, respectively, and log represents the logarithm function; the benefit of using the logarithm function is that when the model predicts a very small probability for the target class, it outputs a very high loss, more strongly encouraging the model to further optimize.

6. The method of claim 4, wherein the pre-trained voice model is a multilingual pre-trained voice model. For the voiceprint recognition task, the evaluation mechanism steps are as follows: Speaker voiceprint representation extraction: The speaker voiceprint representation is extracted from the final full connection layer of the fusion model; the speaker voiceprint representation represents the characteristics of the speaker and is used for subsequent similarity comparison; specifically, for two new input audio x1, x2 to be judged, first process them into two groups of data respectively, one group is the same audio x 11 , x 21 as the original audio, and the other group is a splicing x 12 , x 22 of five audio segments of fixed length randomly cut from the original audio; input x 11 , x 21 , x 12 , x 22 into the fusion model respectively to obtain the speaker voiceprint representation embedding 11 , embedding 21 , embedding 12 , embedding 22 of each audio segment; Verification score calculation: The cosine similarity is used to calculate scores score1 of embedding 11 , score2 of embedding 21 , score3 of embedding 12 , and score4 of embedding 22 , respectively, and the formulas are as follows: Voiceprint recognition result judgment: Finally, take the average of score1 and score2 to get the final score score. According to the final score score and a trained threshold, it is judged whether the two audios come from one person.

Citation Information

Cited By

  • Pluggable target speaker speech recognition method and system

    CN121963713A