Mask self-encoding voiceprint recognition method based on multi-level feature fusion

Through the multi-level feature fusion and self-supervised learning methods, the problems of data calibration time-consuming and privacy leakage in existing voiceprint recognition technology are solved, the accuracy and privacy protection of voiceprint recognition are improved, and a better hidden space representation is built.

CN120388568APending Publication Date: 2025-07-29THE FIRST AFFILIATED HOSPITAL OF WENZHOU MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411092231.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing voiceprint recognition technology has time-consuming, high cost and privacy leakage risks under the supervised learning paradigm, and the masked autoencoder does not fully utilize the different levels of the encoder, resulting in suboptimal representation of hidden space.

Method used

The masked self-encoding vocalprint recognition method is adopted with multi-level feature fusion, and the features of different layers in the encoder are dynamically fused, and the features are aligned and fused using linear projection layers and dynamic weighting strategies are aligned and fused, and the Mel spectrogram is reconstructed and fine-tuned in combination with self-supervised learning.

Benefits of technology

It improves the accuracy of voiceprint recognition tasks, reduces data labeling costs, enhances privacy protection, and builds better hidden space representations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388568A_ABST
    Figure CN120388568A_ABST
Patent Text Reader

Abstract

The invention discloses a mask self-encoding voiceprint recognition method based on multi-level feature fusion. The method comprises the following steps: converting original audio data into a Mel spectrogram through short-time Fourier transform and a Mel filter bank; blocking the Mel spectrogram, randomly masking the Mel spectrogram, and inputting the Mel spectrogram into an encoder; selecting a plurality of layers of middle features, performing semantic alignment on the middle features and the features of the last layer of the encoder by using a projection layer, and obtaining fusion features by using a dynamic weight fusion strategy; inputting the fusion features into a decoder, and completing pre-training by taking the minimum absolute value loss between the original Mel spectrogram and the reconstructed Mel spectrogram as an optimization target; in the fine tuning stage, a pre-trained encoder is used as an initial model, voiceprint classification is carried out by using a data set with a label, and fine tuning is carried out by using the probability of each category output by the model and the cross entropy loss between real labels as an optimization target. According to the scheme, the representation quality in the hidden space is enhanced, and the accuracy of the voiceprint recognition task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to voiceprint recognition, specifically a masked auto-encoding voiceprint recognition method based on multi-level feature fusion. Background Art

[0002] Language is a unique product of human society and the key for humans to express emotions and understand the world. Speakers have unique vocal organs and speaking styles, such as different vocal cord structures, oral shapes, accents, and speaking rhythms. Therefore, the voice signals of different speakers contain unique features, and the task of identifying the speaker's identity can be completed based on these different voice features. Thanks to the advantages of non-invasiveness and real-time of voiceprint recognition technology, it has a wide range of applications in fields such as financial payment, smart furniture, call service, and public security.

[0003] Most traditional voiceprint recognition technologies adopt the paradigm of supervised learning, which is divided into two stages: supervised training and inference testing. In the supervised training stage, a large-scale set of voice data and their corresponding identity labels need to be obtained, and then a network model is trained to establish the mapping relationship between the voices and identity labels in the training set. Researchers in related fields have successively proposed network models such as Gaussian mixture model, deep neural network, convolutional neural network, ECAPA-TDNN, and Transformer. In the inference testing stage, based on the speaker's voice, the identity ID corresponding to the voice with the highest similarity to the current voice is found from the backend database, and compared with the identity of the current speaker to complete the identity verification. However, the use of the supervised learning paradigm has problems such as time-consuming data calibration and high cost, especially for minority languages such as dialects. At the same time, the calibrated data has the risk of leakage and being hacked, which is not conducive to protecting personal privacy.

[0004] Self-supervised learning is a learning paradigm with lower cost and better performance, which is divided into three stages: pre-training, downstream fine-tuning, and inference testing. In the pre-training stage, the model is pre-trained on a large-scale unlabeled dataset to learn the general and essential features in the voice data, saving the time and cost of data annotation. In the fine-tuning stage, since the model has learned many common features from a large amount of data in the pre-training stage, a small amount of voice data and identity labels of the speaker to be identified can be given to fine-tune the parameters in the pre-trained model to make it adapt to different downstream scenarios.

[0005] The masked autoencoder is a mainstream self-supervised learning model originally proposed in the field of computer vision. Its core idea is to divide the mel-spectrogram obtained after Fourier transform of sound data into patches of equal size and randomly mask each patch according to a certain ratio. The unmasked patches are then encoded into high-dimensional feature vectors using a Transformer model as an encoder. A lightweight Transformer is then used as a decoder to map the high-dimensional feature vector output by the encoder and the learnable vector assigned to the masked patch position back to the original mel-spectrogram space. The absolute value loss between the reconstructed mel-spectrogram and the original mel-spectrogram is calculated as the model's loss function. After pre-training, only the encoder is retained for fine-tuning on specific downstream tasks.

[0006] Although some work has applied the architecture of masked autoencoders to the field of voiceprint recognition, there are still certain problems, mainly reflected in the following aspects: (1) Only the features of the last layer of the encoder are transmitted to the decoder to reconstruct the original Mel-spectrogram, which does not fully utilize the shallow features in the encoder. (2) The lack of full utilization of features at different levels results in the latent space representation obtained by the encoder being suboptimal, and there is still room for improvement in the downstream voiceprint recognition task. Based on the above analysis, for the actual application scenarios of voiceprint recognition such as financial payment, smart furniture, and call services, there is an urgent need for a voiceprint recognition technology based on masked autoencoders that can focus on difficult samples. Summary of the Invention

[0007] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a masked self-encoding voiceprint recognition method based on multi-level feature fusion, dynamically fuse the features of different layers in the encoder, and input the fused features into the decoder for reconstruction, thereby further enhancing the representation quality in the latent space, solving the problem of suboptimal representation caused by only using the last layer of encoder features for reconstruction in the existing method, and improving the accuracy of voiceprint recognition tasks.

[0008] To achieve the above purpose, the present invention provides the following technical solution: a masked self-encoding voiceprint recognition method based on multi-level feature fusion, comprising the following steps:

[0009] S1. Obtain a large amount of unlabeled voiceprint data as a pre-training dataset; obtain a smaller amount of labeled voiceprint data as a downstream fine-tuning dataset;

[0010] S2, using short-time Fourier transform and Mel frequency filter bank to convert the original voiceprint data in the dataset into a Mel frequency spectrogram describing the speech spectrum characteristics;

[0011] S3: The Mel-spectrogram is divided into blocks and randomly masked. The visible blocks after masking are embedded using a projection layer. The features of the visible blocks are then input into an encoder composed of Transformer layers. The multi-head self-attention mechanism is used to calculate the dependencies between different blocks to obtain the output of each layer of the encoder model.

[0012] S4. Select several layers of features in the middle of the encoder and use a linear projection layer to semantically align the features of different layers with the features of the last layer of the encoder. Use a dynamic weighting strategy to fuse the aligned features of the middle layers with the features of the last layer, and input the fused features into the decoder.

[0013] S5. Assign a learnable feature vector to the mask position. Use a decoder consisting of a Transformer layer and a projection layer to reconstruct the multi-level fused encoding features and the learnable feature vector of the mask position back into the original Mel-spectrogram space. Calculate the absolute value loss between the original Mel-spectrogram at the mask position and the reconstructed Mel-spectrogram, and optimize the model based on this loss.

[0014] S6. Use the labeled fine-tuning dataset to fine-tune the encoder to complete the voiceprint recognition task.

[0015] As a further improvement of the present invention, S2 includes

[0016] S2-1. Assume that the input speech signal is x(t); perform frame operation on x(t) to obtain a locally stable speech signal. Assume that the length of each frame is T and the displacement between adjacent frames is Δt; the nth frame signal x n (t) is expressed as x(t+n·Δt); in each frame signal x n (t) is applied with the Hamming window function w(t) to reduce the edge effect after the speech signal is framed, and the windowed speech signal x is obtained. n (t)·w(t);

[0017] S2-2, perform N-point discrete Fourier transform on each frame of windowed speech signal to obtain the short-time spectrum X n (k);

[0018] S2-3, convert the original linear frequency k into Mel frequency mf; assuming that the minimum value of the original linear frequency is f min , the maximum value is f max , in this interval, M uniform Mel frequencies are divided as the center frequency of the Mel filter, and the center frequency of the mth filter is f(m); M triangular filters are used to filter X in the Mel frequency domain n (k) Perform weighted summation to obtain the Mel spectrum S n(m), where m is an integer with a value range of [1, M].

[0019] As a further improvement of the present invention, S3 includes

[0020] S3-1. Denote the Mel spectrogram of the original speech as where f represents the number of filters and t represents the number of frames; divide the Mel spectrogram into N blocks of size p×p, and send each block into the block projection layer to be encoded into a one-dimensional vector, denoted as where d represents the dimension of the projection, and the block projection layer is a linear projection layer; use a combination of sine and cosine functions to encode the position information, expressing the position information of each block in the entire sequence, denoted as where d represents the dimension of the position encoding; add x pos and x proj to obtain the representation x of the mixed speech features and position features of each tile;

[0021] S3-2. Perform a random masking operation on x according to the masking rate p to obtain the visible part x vis ∈ and the masked part Only send the visible part x vis into the encoder composed of Transformer Blocks; Transformer Blocks consists of a multi-head attention calculation module and a feed-forward neural network layer; the multi-head attention mechanism module is composed of multiple self-attention mechanism modules spliced together, and each self-attention mechanism module models the input sequence from different perspectives, learning different weight allocation methods, so as to more comprehensively capture the information in the input sequence; denote the input of each layer of Block as x in , and the output as x out , and the input of the first layer is x vis ;

[0022] S3-3. The output of each layer of Transformer Blocks is used as the input of the next layer of Blocks until the result of the Lth layer is calculated; thus, the output of each layer of the encoder is obtained, denoted as {e1, e2…e L}}.

[0023] As a further improvement of the present invention, S4 includes

[0024] S4-1. Assume that M intermediate features and the last layer feature are selected for fusion, and denote the indices of the selected M layers as {s1, s2…s M}}; the semantic information of different layers in the encoder is different. In order to align the feature spaces between different layers, use M linear projection layers {P1, P2…P MAlign the features of the selected layer with the features of the last layer to avoid the impact of semantic space differences on training the decoder; after the above alignment operation, obtain the multi-level features of M+1 layers for fusion.

[0025] S4-2. Perform dynamic weight fusion on the aligned features of different layers obtained; let {W1, W2... w M , W M+1} represent the allocated dynamic weights, which can be dynamically adjusted according to the size of the reconstruction loss during the learning process, and their sum is always 1; use {W1, W2... w M , W M+1} to weight the selected M+1 layer features to obtain the fused feature O; the feature dimension of O is the same as the feature vector dimension of each layer of the encoder and is input into the encoder.

[0026] As a further improvement of the present invention, S5 includes

[0027] S5-1. Use the feature vector generated by the encoder through the visible tiles and the position information of the masked tiles to restore the original spectral information of the masked tiles; assign a same learnable feature vector to the masked tiles, denoted as whose feature dimension is the same as the dimension of the fused feature O of the visible part;

[0028] Input O and mtoken into the decoder composed of Transformer Blocks as well; the internal of Transformer Blocks also uses the multi-head attention mechanism for calculation, and the principle is the same as that of the encoder in S3-2;

[0029] Finally, obtain the decoded feature vectors of the visible part and the masked part at the same time, denoted as The decoded feature vector of the masked part, denoted as

[0030] S5-2. Use the linear projection layer to project the decoded feature vector back to the original mel-spectrogram space, that is, predict the original mel frequency of each masked tile pixel by pixel, denoted as Let the frequency at the masked position in the original mel-spectrogram be Calculate the per-pixel absolute value loss between the reconstructed spectrogram and the original spectrogram, denoted as Loss rec ; reduce the reconstruction loss through the training of the model.

[0031] As a further improvement of the present invention, S6 includes

[0032] S6-1. Use the fully connected layer and softmax function to calculate the probability of the voiceprint data belonging to each category based on its feature vector; in represents the parameters in the fully connected layer, c represents the number of voiceprint categories;

[0033] S6-2. Assume that the label of the voiceprint data is It is in the form of a one-hot code; using cross entropy loss L cross Perform model optimization, the formula is After several rounds of training, the cross entropy loss converges and the fine-tuning phase ends.

[0034] Beneficial effects of the present invention:

[0035] 1) The idea of multi-level feature fusion is introduced to fuse the low-level semantic information contained in the shallow features of the encoder and the high-level semantic information contained in the deep features to better reconstruct the original Mel-spectrogram.

[0036] 2) A linear projection layer is used to align the semantic differences between the shallow features of the encoder and the last layer features to prevent the influence of semantic differences on the optimization process.

[0037] 3) Use a dynamic weight fusion strategy to assign a learnable weight to each layer of features to be fused, and adaptively adjust it during the training process to autonomously weigh the importance of each layer of features.

[0038] 4) Using self-supervision methods to complete voiceprint recognition tasks has stronger generalization and can reduce the cost of data annotation.

[0039] 5) Fuse the multi-level features of the encoder and use the semantic alignment layer and dynamic fusion strategy to improve the quality of the features used for reconstruction, thereby building a latent space with better voiceprint representation and better reconstructing the original Mel-spectrogram.

[0040] 6) Thanks to the better latent space representation, higher voiceprint recognition accuracy can be achieved when fine-tuning the downstream voiceprint recognition task, further enhancing the reliability of the voiceprint system. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a principle flow chart of the present invention.

[0042] Figure 2 Schematic diagram of a flow chart of an embodiment of the present invention. DETAILED DESCRIPTION

[0043] The present invention will be further described below with reference to the embodiments shown in the accompanying drawings.

[0044] The present invention provides a masked auto-encoding voiceprint recognition method based on multi-level feature fusion, as Figure 1 and Figure 2 shown. First, the original audio data is converted into a Mel spectrogram through short-time Fourier transform and Mel filter bank; the Mel spectrogram is segmented and randomly masked, and then input into an encoder composed of Transformer blocks; several intermediate features are selected, and a projection layer is used to semantically align them with the features of the last layer of the encoder, and then a dynamic weight fusion strategy is used to obtain the fused features; the fused features are input into the decoder and reconstructed back to the original Mel spectrogram space, and the absolute value loss between the original Mel spectrogram and the reconstructed Mel spectrogram is minimized as the optimization objective to complete the pre-training; in the fine-tuning stage, the pre-trained encoder is used as the initial model, and a small amount of labeled dataset is used for voiceprint classification, and the cross-entropy loss between the probability of each category output by the model and the true label is used as the optimization objective for fine-tuning.

[0045] Specifically, after obtaining the pre-training and fine-tuning voiceprint datasets, the present invention performs the following operations in sequence:

[0046] Step (1). Use short-time Fourier transform and Mel frequency filter bank to convert the original voiceprint data in the dataset into a Mel spectrogram describing the speech spectral features.

[0047] Step (2). Segment and randomly mask the Mel spectrogram, embed the visible blocks after masking using a projection layer, and then input the features of the visible blocks into an encoder composed of Transfomer layers, and use the multi-head self-attention mechanism to calculate the dependencies between different segments to obtain the output of each layer of the encoder model.

[0048] Step (3). Select several intermediate layer features of the encoder, and use a linear projection layer to semantically align the features of different layers with the features of the last layer of the encoder. Use the dynamic weight strategy to fuse the aligned features of several intermediate layers with the features of the last layer, and input the fused features into the decoder.

[0049] Step (4). Assign a learnable feature vector to the masked positions, and use a decoder composed of Transformer layers and projection layers to reconstruct the multi-level fused encoded features and the learnable feature vectors at the masked positions back to the original Mel spectrogram space. Calculate the absolute value loss between the original Mel spectrogram at the masked positions and the reconstructed Mel spectrogram, and optimize the model according to this loss.

[0050] Step (5). After the pre-training phase of steps (2) to (4), the encoder has acquired the semantic perception capability of voiceprint data. In the fine-tuning phase, the encoder is fine-tuned using a labeled fine-tuning dataset to complete the voiceprint recognition task.

[0051] Step (1) is specifically:

[0052] Assume that the input speech signal is x(t). In order to obtain a locally stable speech signal, x(t) is divided into frames, and the length of each frame is T, and the displacement between adjacent frames is Δt. The nth frame signal x n (t) can be expressed as x(t+n·Δt). In order to reduce the edge effect after the speech signal is framed, n Apply the Hamming window function w(t) to (t) to obtain the windowed speech signal x n (t)·w(t).

[0053] Perform N-point discrete Fourier transform on each frame of windowed speech signal to obtain the short-time spectrum X n (k). The specific formula for this operation is The value range of k is an integer in [O, N-1], and j represents the imaginary unit.

[0054] Since the human ear perceives frequency nonlinearly, the original linear frequency k is converted into the Mel frequency mf that is more in line with the characteristics of the human ear. The specific formula for this operation is: Assume that the minimum value of the original linear frequency is f min , the maximum value is f max , in this interval, M uniform Mel frequencies are divided as the center frequency of the Mel filter. The calculation formula for the center frequency f(m) of the mth filter is Use M triangular filters to filter X in the Mel frequency domain n (k) Perform weighted summation to obtain the Mel spectrum S n (m), the specific formula is The value range of m is an integer in [1, M]. m (k) is the weight function of the m-th filter bank, defined as follows:

[0055]

[0056] Step 2 specifically is:

[0057] (2-1) After step 1, the Mel-spectrogram of the original speech is obtained, which is recorded as Where f represents the number of filters and t represents the number of frames. The Mel spectrum map is divided into N blocks of size p×p, and each block is sent to the block projection layer and encoded into a one-dimensional vector, which is recorded as where d represents the dimension of the projection, and the block projection layer is a linear projection layer. To represent the position information of each block in the entire sequence, a combined encoding of sine and cosine functions is used to encode the position information, denoted as where d represents the dimension of the position encoding. Add x pos and x proj to obtain the representation x of the mixed speech features and position features of each tile.

[0058] (2-2) Perform a random masking operation on x according to the masking rate p to obtain the visible part and the masked part Only send the visible part x vis into the encoder composed of Transformer Blocks. Transformer Blocks consists of a multi-head attention calculation module and a feed-forward neural network layer. The multi-head attention mechanism module is composed of multiple self-attention mechanism modules spliced together. Each self-attention mechanism module models the input sequence from different perspectives and learns different weight assignment methods, so as to capture the information in the input sequence more comprehensively. Denote the input of each layer of Block as x in , and the output as x out , and the input of the first layer is x vis . The specific calculation process is as follows. First, multiply the input sequence x in on the right by n groups of different mapping transformation matrices and to project the original input into the query, key, and value spaces of different attention mechanisms. The calculation formula is Calculate the inner product of each group of Q i , K i T to obtain the strength of the correlation between each patch and other patches. Multiply the attention weight matrix by V i to obtain the final output of the i-th head. The calculation formula is where d k represents the size of the dimension of Q i and K i . Concatenate the outputs obtained by each head through the self-attention mechanism, multiply on the right by a linear transformation matrix W2 to restore the dimension to the size at the input. The calculation formula is X middle = Concate(Head1, Head2…Head n )·W2. Perform a residual connection between x in and x middle . After the concatenated data is layer-normalized again, it is sent into a multi-layer perceptron. The MLP changes the dimension of each block from p2 Mapped to 8p 2 and then mapped back to p 2 。

[0059] (2 - 3) The output of each layer of Transformer Blocks is used as the input to the next layer of Blocks until the result of the L-th layer is calculated. Thus, the output of each layer of the encoder is obtained, denoted as {e1, e2…e L}.

[0060] Step 3 is specifically as follows:

[0061] (3 - 1) Assume that M intermediate features and the last layer feature are selected for fusion, and the indices of the selected M layers are denoted as {s1, s2…s M}. Since the semantic information of different layers in the encoder is different, in order to align the feature spaces between different layers, M linear projection layers {P1, P2…P M} are used to align the features of the selected layers with the last layer feature, avoiding the influence of semantic space differences on training the decoder. After the above alignment operation, M + 1 layers of multi-level features for fusion are obtained

[0062] (3 - 2) Perform dynamic weight fusion on the aligned features of different layers obtained. Let {W1, W2…w M , W M+1} represent the dynamic weights assigned to each layer, which can be dynamically adjusted according to the size of the reconstruction loss during the learning process, and their sum is always 1. Use {W1, W2…w M , W M+1} to weight the selected M + 1 layer features to obtain the fused feature O. The specific calculation formula is After two steps of alignment and fusion, O contains richer information from different levels and is aligned through the projection layer. The feature dimension of O is the same as the feature vector dimension of each layer of the encoder, both being d.

[0063] Step 4 is specifically as follows:

[0064] (4 - 1) After the above three steps, the encoder obtains the fused feature vector O of the visible patches. Use the feature vector generated by the encoder through the visible patches and the position information of the masked patches to restore the original spectral information of the masked patches. First, assign the same learnable feature vector to the masked patches, denoted as Its feature dimension is the same as that of the fused feature O of the visible part. Input O and mtoken into the decoder which is also composed of TransformerBlocks. The internal of Transformer Blocks also uses the multi-head attention mechanism for calculation, and the principle is the same as that of the encoder in step (2-2). Finally, the decoded feature vectors of the visible part and the masked part are obtained simultaneously, denoted as Since the training objective of the model is to reconstruct the original information of the masked patches, only the decoded feature vectors of the masked part are needed, denoted as

[0065] (4-2) Use a linear projection layer to project the decoded feature vectors back to the original Mel spectrogram space, that is, predict the original Mel frequency of each masked patch pixel by pixel, denoted as Let the frequency at the masked position in the original Mel spectrogram be Calculate the per-pixel absolute value loss between the reconstructed spectrogram and the original spectrogram, denoted as Loss rec . The calculation formula is The training objective of the model is to make the reconstruction loss as small as possible.

[0066] Step 5 is specifically as follows:

[0067] (5-1) After the above pre-training process, the parameters of the model encoder have been optimized to a relatively good position. Use a small amount of labeled downstream dataset for fine-tuning. Only the encoder obtained in the pre-training stage is retained during the fine-tuning stage. Use a fully connected layer and the softmax function to calculate the probability of each category that the voiceprint data belongs to according to the feature vectors of the voiceprint data. Where represents the parameters in the fully connected layer, and c represents the number of voiceprint categories.

[0068] (5-2) Assume that the label of the voiceprint data is which is in the form of a one-hot code. Use the cross-entropy loss L cross to optimize the model, and the formula is After several rounds of training, the cross-entropy loss converges and the fine-tuning stage ends.

[0069] The above is only the preferred implementation mode of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions within the idea of the present invention belong to the protection scope of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and retouches should also be regarded as the protection scope of the present invention.

Claims

1. A masked auto-encoding voiceprint recognition method based on multi-level feature fusion, characterized in that, It includes the following steps S1. Obtain unlabeled voiceprint data with a large scale as the dataset for pre-training; obtain labeled voiceprint data with a small scale as the dataset for downstream fine-tuning; S2. Use the short-time Fourier transform and the Mel frequency filter bank to convert the original voiceprint data in the dataset into a Mel spectrogram that describes the speech spectrum features; S3. After dividing the Mel spectrogram into blocks, randomly mask them. Embed the visible blocks after masking using a projection layer, and then input the features of the visible blocks into an encoder composed of Transformer layers. Use the multi-head self-attention mechanism to calculate the dependencies between different blocks to obtain the output of each layer of the encoder model; S4. Select the features of several intermediate layers of the encoder, and use a linear projection layer to semantically align the features of different layers with the features of the last layer of the encoder; Use a dynamic weight strategy to fuse the aligned features of several intermediate layers with the features of the last layer, and input the fused features into the decoder; S5. Assign a learnable feature vector to the masked positions. Use a decoder composed of Transformer layers and a projection layer to reconstruct the multi-level fused encoded features and the learnable feature vectors at the masked positions back to the original Mel spectrogram space; calculate the absolute value loss between the original Mel spectrogram at the masked positions and the reconstructed Mel spectrogram, and optimize the model according to this loss; S6. Fine-tune the encoder using the labeled fine-tuning dataset to complete the voiceprint recognition task.

2. The masked auto-encoding voiceprint recognition method based on multi-level feature fusion according to claim 1, wherein S2 includes S2-1. Assume that the input speech signal is x(t); perform a framing operation on x(t) to obtain a locally stable speech signal. Let the length of each frame be T and the displacement between adjacent frames be Δt; the nth frame signal x n (t) is expressed as x(t + n·Δt); Apply the Hamming window function w(t) to each frame of the signal x n (t) to reduce the edge effect generated after the speech signal is framed, and obtain the windowed speech signal x n (t)·w(t); S2-2. For each frame of the windowed speech signal, perform an N-point discrete Fourier transform to obtain the short-time spectrum X n (k); S2-3. Convert the original linear frequency \(k\) to the Mel frequency \(mf\); assume that the minimum value of the original linear frequency is \(f\) min , and the maximum value is \(f\) max . Divide \(M\) evenly spaced Mel frequencies within this range as the center frequencies of the Mel filters, and the center frequency of the \(m\)-th filter is \(f(m)\); use \(M\) triangular filters to perform weighted summation on \(X\) n (k) in the Mel frequency domain to obtain the Mel spectrum \(S\) n (m), where \(m\) ranges from 1 to \(M\) as an integer.

3. The masked auto-encoding voiceprint recognition method based on multi-level feature fusion according to claim 2, wherein S3 includes S3-1. Denote the Mel spectrogram of the original speech as where f represents the number of filters and t represents the number of frames; divide the Mel spectrogram into N blocks of size p×p, and send each block into the block projection layer to be encoded as a one-dimensional vector, denoted as where d represents the dimension of the projection, and the block projection layer is a linear projection layer; use a combination of sine and cosine functions to encode the position information, expressing the position information of each block in the entire sequence, denoted as where d represents the dimension of the position encoding; add x pos and x proj to obtain the representation x of the mixed speech features and position features of each tile; S3-2. Randomly mask x according to the masking rate p to obtain the visible part and the masked part Only send the visible part x vis into the encoder composed of Transformer Blocks; Transformer Blocks are composed of a multi-head attention calculation module and a feed-forward neural network layer; the multi-head attention mechanism module is composed of multiple self-attention mechanism modules spliced together. Each self-attention mechanism module models the input sequence from different perspectives, learns different weight allocation methods, so as to more comprehensively capture the information in the input sequence; denote the input of each layer of Block as x in and the output as x out The input of the first layer is x vis ; S3-3. The output of each layer of Transformer Blocks is used as the input of the next layer of Blocks until the result of the L-th layer is calculated; thus, the output of each layer of the encoder is obtained, denoted as {e1, e2…e L}.

4. The masked auto-encoding voiceprint recognition method based on multi-level feature fusion according to claim 3, characterized in that, S4 includes S4-1. Assume that the middle features of M layers and the features of the last layer are selected for fusion, and the indices of the selected M layers are denoted as {s1, s2... s M}; Since the semantic information of different layers in the encoder is different, in order to align the feature spaces between different layers, M linear projection layers {P1, P2... P M} are used to align the features of the selected layers with the features of the last layer, avoiding the influence of semantic space differences on training the decoder; After the above alignment operation, multi-level features for fusion of M + 1 layers are obtained S4-2. Dynamically fuse the alignment features of different layers obtained; let {W1, W2…w M , W M+1} represent the dynamic weights assigned to each layer, which can be dynamically adjusted according to the size of the reconstruction loss during the learning process, and their sum is always 1; use {W1, W2…w M , W M+1} to weight the selected M + 1 layer features to obtain the fused feature O; the feature dimension of O is the same as the feature vector dimension of each layer of the encoder and is input into the encoder.

5. The masked auto - encoding voiceprint recognition method based on multi - level feature fusion according to claim 4, wherein, S5 includes S5-1. Use the feature vectors generated by the encoder from the visible patches and the position information of the masked patches to recover the original spectral information of the masked patches; assign a same learnable feature vector to the masked patches, denoted as whose feature dimension is the same as that of the fused feature O of the visible part; input O and mtoken into the decoder also composed of Transformer Blocks; the internal calculation of Transformer Blocks also adopts the multi-head attention mechanism, and the principle is the same as that of the encoder in S3-2; Finally, the decoded feature vectors of the visible part and the mask part are obtained simultaneously, denoted as The decoded feature vector of the mask part, denoted as S5-2. Use a linear projection layer to project the decoded feature vector back to the original mel-spectrogram space, i.e., predict the original mel-frequency of each masked patch pixel-wise, denoted as Let the frequency at the masked position in the original mel-spectrogram be Calculate the pixel-wise absolute value loss between the reconstructed spectrogram and the original spectrogram, denoted as Loss rec ; Minimize the reconstruction loss through model training.

6. The masked auto-encoding voiceprint recognition method based on multi-level feature fusion according to claim 5, wherein, S6 includes S6-1. Calculate the probability that the feature vector of the voiceprint data belongs to each category using the fully connected layer and the softmax function; where represents the parameters in the fully connected layer, and c represents the number of voiceprint categories; S6-2. Assume that the label of the voiceprint data is in the form of one-hot code; use the cross-entropy loss L cross to optimize the model, and the formula is After several rounds of training, the cross-entropy loss converges and the fine-tuning stage ends.