A voiceprint recognition method based on random mask training and a computer device
By employing a voiceprint recognition method trained with random masking and utilizing the Wav2Vec2 model to construct lossy speech feature vectors, the robustness problem of voiceprint recognition in noisy environments is solved, and accurate recognition and feature extraction of lossy speech are achieved.
Patent Information
- Application Number
- CN202211193071.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-28
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2042-09-28
AI Technical Summary
Existing voiceprint recognition technology has poor robustness in noisy environments, making it difficult to accurately recognize lossy speech. Furthermore, the time-varying nature of speech features and noise interference affect the recognition results.
A feature extraction model is constructed using a random masking training method. The Wav2Vec2 model is used to encode the random masking features and construct a lossy speech feature vector. The model is then trained by masking and loss iteration to obtain the feature extraction model. Finally, cosine similarity calculation is used for voiceprint verification.
It achieves accurate recognition of lossy speech in noisy environments, improves the robustness of voiceprint recognition, and can accurately extract the speaker's voiceprint features in various environments.
Smart Images

Figure CN115691510B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech recognition technology, and in particular to a voiceprint recognition method and computer device based on random masking training. Background Technology
[0002] Voiceprint recognition, also known as speaker identification, extracts voiceprint features from a speaker's speech to build a model for identifying the speaker. Currently, voiceprint recognition research faces three main challenges. First, the semantic information and individual vocal characteristics of a speaker's speech are complexly mixed, making it difficult to separate individual speaker features from speech features, thus posing a challenge to voiceprint recognition. Second, a speaker's voice features are not static or fixed but rather time-varying physical characteristics, often closely related to external factors such as the speaker's environment, emotions, health status, time, and age. Even the same person speaking the same words can produce different voice signals. Third, the performance of speaker recognition systems in applications needs improvement. Sound transmission through communication lines is inevitably affected by line noise, and different transmission lines introduce different types of noise. Therefore, the challenge of speaker recognition in noisy environments is how to effectively remove various additive and multiplicative noise interferences and improve the robustness of the system under conditions of high noise or even lossy speech. Summary of the Invention
[0003] Based on the above analysis, the present invention aims to provide a voiceprint recognition method and computer device based on random masking training; it solves the problem that existing voiceprint recognition methods cannot accurately recognize lossy speech and have poor robustness.
[0004] The objective of this invention is mainly achieved through the following technical solutions:
[0005] On one hand, this invention discloses a voiceprint recognition method based on random masking training, the method comprising the following steps:
[0006] A user voice registration library is obtained by registering multiple user voices through a pre-trained feature extraction model; the feature extraction model is obtained by constructing lossy voice feature vectors using a random masking method and then training them.
[0007] The speech to be recognized is acquired, and the feature extraction model is used to extract features from the speech to be recognized to obtain the feature vector of the speech to be recognized.
[0008] The feature vector of the speech to be identified is compared with the cosine similarity value of all registered speech in the user speech registration database; the user to whom the speech to be identified belongs is determined based on the cosine similarity value.
[0009] Furthermore, a lossy speech feature vector is constructed using a random masking method and trained to obtain the feature extraction model, including:
[0010] Based on the Wav2Vec2 model, a random masking feature encoding model is constructed.
[0011] The speech sample data is input into the random masking feature coding model for feature extraction, random masking and feature coding to obtain a lossy speech feature vector.
[0012] The lossy speech feature vectors are masked and the masking vector is predicted. The feature extraction model is obtained by iteratively updating the loss and training.
[0013] Furthermore, the random masking feature encoding model performs the following steps to obtain the lossy speech feature vector:
[0014] Feature extraction is performed on the input speech sample data to obtain a fixed-dimensional speech feature vector;
[0015] The speech feature vectors are randomly masked according to a uniform distribution; the positions of the masked vectors and the remaining vectors are recorded, and the position information is embedded into the remaining vectors to obtain the lossy speech data after random masking.
[0016] The lossy speech data after random masking is subjected to inter-frame attention calculation to obtain a lossy speech feature vector that has temporal and contextual relationships.
[0017] Furthermore, the feature extraction model is trained through the following steps:
[0018] Construct a voiceprint recognition training dataset, which includes voice sample data and labels representing the person to whom the voice data belongs;
[0019] The lossy speech feature vector is constructed based on the speech sample data using the random masking feature coding model.
[0020] The lossy speech feature vector is padded with vectors and embedded with positional information to obtain the mask-padded feature vector.
[0021] Based on the feature vector after mask filling, the speech feature vector corresponding to the randomly masked part is predicted to obtain the complete speech feature vector;
[0022] The complete speech feature vector is subjected to average pooling and softmax classification to obtain the label category distribution corresponding to the complete speech feature vector;
[0023] The labels obtained from the voiceprint recognition training dataset and the classifier are iteratively updated using a loss function to obtain a trained random masking feature encoding model; based on the trained random masking feature encoding model, the feature extraction model is obtained by disabling the random masking function.
[0024] Furthermore, the step of performing vector padding and positional information embedding on the lossy speech feature vector includes: assigning a shared learning vector or a fixed non-zero vector to the masked vector; and embedding corresponding positional information into the masked vector and the remaining vector to obtain the masked feature vector.
[0025] Furthermore, the mask-padded feature vector is input into the decoder, and the complete speech feature vector is obtained using the following formula:
[0026]
[0027] in, This represents the speech feature vector after adding the padding vector. Represents location encoding information, f bn f represents the normalization operation. transformer This represents the transformer decoder.
[0028] Furthermore, the complete speech feature vector is subjected to average pooling and softmax classification to obtain the prediction vector.
[0029]
[0030] The loss function formula is:
[0031]
[0032] Where x is the input speech sequence and y is the speaker category label.
[0033] Furthermore, the step of registering multiple user voices using a pre-trained feature extraction model to obtain a user voice registration library includes: inputting the voice of the user to be registered into the feature extraction model to obtain the voice feature vector of the user to be registered; labeling the voice feature vector of the user to be registered based on the user ID; saving the feature vector of the voice to be registered and the corresponding label to obtain the user voice registration library.
[0034] Furthermore, the shielding rate of the random shielding is greater than 50%.
[0035] On the other hand, a computer device is also disclosed, including at least one processor and at least one memory communicatively connected to said processor;
[0036] The memory stores instructions that can be executed by the processor to implement the aforementioned voiceprint recognition method.
[0037] The present invention can achieve at least the following beneficial effects:
[0038] 1. This invention uses a powerful voiceprint recognition feature extraction model trained by a random masking method to extract speech features, and combines it with a cosine similarity calculation method for voiceprint verification, thereby achieving accurate recognition of lossy speech in different environments.
[0039] 2. The random masking feature extraction model used in this invention is trained through random masking operations with a high masking rate. It can accurately extract features from discontinuous or even lossy speech to be identified, and obtain voiceprint features that can represent the speaker.
[0040] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0041] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.
[0042] Figure 1 This is a flowchart of the voiceprint recognition method according to an embodiment of the present invention. Detailed Implementation
[0043] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.
[0044] This embodiment discloses a voiceprint recognition method based on random masking training, such as... Figure 1 As shown, it includes the following steps:
[0045] Step S1: Register multiple user voices using a pre-trained feature extraction model to obtain a user voice registration library;
[0046] Specifically, multiple user voices are registered using a pre-trained feature extraction model to obtain a user voice registration library. This includes: inputting the voice of the user to be registered into the feature extraction model to obtain the voice feature vector of the user to be registered; labeling the voice feature vector of the user to be registered based on the user ID; and saving the feature vector of the voice to be registered and the corresponding label to obtain the user voice registration library.
[0047] The feature extraction model is obtained by constructing a lossy speech feature vector using a random masking method and then training it.
[0048] Specifically, a lossy speech feature vector is constructed using a random masking method and trained to obtain a feature extraction model, including:
[0049] Based on the Wav2Vec2 model, a random masking feature encoding model is constructed.
[0050] The speech sample data is input into the random masking feature coding model for feature extraction, random masking and feature coding to obtain a lossy speech feature vector.
[0051] The lossy speech feature vectors are masked and the masking vector is predicted. The model is then trained by iterative updating of the loss function to obtain the feature extraction model.
[0052] As a specific example, the feature extraction model can be trained through the following steps:
[0053] Step S101: Construct a voiceprint recognition training dataset, which includes voice sample data and labels representing the person to whom the voice data belongs;
[0054] The speech data was obtained by preprocessing the speech of unconstrained speakers in the Chinese speech dataset. Labels were obtained by classifying and annotating the speech of different speakers; for example, each speaker's speech was labeled with a number representing the label category corresponding to each speech.
[0055] Specifically, constructing the voiceprint recognition training dataset includes:
[0056] Obtain the dataset; the dataset includes speech sample data from multiple speakers;
[0057] Construct voice data labels, where each voice data label is the speaker ID corresponding to the voice sample data;
[0058] By preprocessing the speech sample data through duct splitting, speech segmentation, and standardized formatting, monophonic, fixed-length speech data is obtained, forming a voiceprint recognition training dataset.
[0059] Preferably, the dataset used in this embodiment is the Chinese speech datasets cn-celeb1 and cn-celeb2 for unconstrained speaker recognition. cn-celeb1 contains recordings or interviews of 997 people, and cn-celeb2 contains recordings or interviews of 1996 people. Speech segments of speakers with less than fifty segments are deleted, and the remaining data in the cn-celeb1 and cn-celeb2 datasets are labeled according to the speaker ID.
[0060] Further, the data is preprocessed, and the steps are as follows:
[0061] Channel splitting: All speech with more than 1 channel is split into mono speech, and silence removal is performed on the data of each channel separately: The WebRTC speech endpoint detection method is used to detect speech segments. First, the input speech is divided into several segments in 20ms increments. Then, the segment is checked to see if it is silent. If it is, the segment is deleted; otherwise, it is retained.
[0062] Speech segmentation: Speech is segmented into segments of fixed length based on min_token and max_token; segments shorter than min_token are discarded, and segments longer than max_token are truncated. In this embodiment, min_token is 56000, representing a segment of speech no less than 5 seconds, and max_token is 480000, representing a segment of speech no more than 30 seconds.
[0063] Standardized format: The segmented audio is uniformly converted into a format with a sampling rate of 16000 and a sampling precision of 16 bits.
[0064] The preprocessed speech data and corresponding labels constitute a voiceprint recognition dataset, and the training data in it is used as a voiceprint recognition training dataset.
[0065] Step S102: Construct a lossy speech feature vector based on speech sample data using a random masking feature coding model;
[0066] Preferably, the feature encoding model in this embodiment uses the wav2vec2 model as the initial model structure without modifying the original model structure. Its input is preprocessed speech sampling data from the training dataset. The original speech data, after sampling, is used as the model input. It passes through the CNN feature extraction layer in wav2vec2 to obtain speech feature vector representations. Then, a random masking operation is performed, masking a portion of the obtained speech feature vectors while remembering the positions of the masked vectors and the remaining vectors. This embodiment uses a high masking ratio, with a masking probability greater than 50%. The remaining speech vectors are then sequentially arranged, and the embedded position vector information is input into the transformer structure of wav2vec2 to obtain the lossy speech vector representation output by the feature encoder.
[0067] More specifically, based on the wav2vec2 model structure, the feature encoder obtains X = (x0, x1, ..., x) after sampling the original speech data. lAs input, the signal is processed through a 7-layer convolutional network, with the output of each layer serving as the input to the next. The stride of each layer is (5, 2, 2, 2, 2, 2, 2), and the kernel width is (10, 3, 3, 3, 3, 2, 2). After feature encoding, a fixed-dimensional 512 speech feature vector is generated, resulting in the hidden layer feature C = (c0, c1, ..., c2). L The vector has dimensions (1, L, 512), where L is equal to l / 320. The 512-dimensional vector is then mapped to a 768-dimensional space to obtain the speech feature vector representation S. Next, the feature vector S is randomly masked, and the remaining vectors after masking are embedded with positional information. This information is then fed into a 12-layer block for inter-frame attention calculation. Each block is a transformer structure with 768 hidden units and a self-attention network, which fully utilizes the transformer's powerful ability to handle temporal relationships. This allows for sufficient contextual encoding of the features of each speech frame, resulting in speech feature vectors that have temporal and contextual relationships. The output features of the speech encoded by the 12-layer transformer are: H = (h0, h1, ..., h...). l ), with dimensions [1, 1, 768]; where vectors S and H can be represented as:
[0068] S = f bn (Wf cnn (X)+b);
[0069]
[0070] Among them, f cnn Represents a CNN feature extractor, W∈R p*q This represents a mapping from 512 dimensions to 768 dimensions, where p and q represent the dimensions, which are 512 and 768 respectively. This represents the remaining part of S after being masked, P emb represent Location information, f bn f represents the normalization operation. transformer This represents the transformer encoder in wav2vec2.
[0071] It should be noted that for the feature encoder, random masking is only performed during training; in practical applications, this function needs to be disabled. Since the masked portions of the same speech are random, a speech can be divided into several segments with different textual information, serving as data augmentation. Furthermore, due to the high masking ratio, the left and right segments of each masked portion vector are highly likely to be masked. The model can learn the speaker's voiceprint features through discontinuous speech features, making the trained feature encoder highly robust. Finally, the decoder reconstructs the masked portion vectors, and through iterative training updates, this model can recognize all speech features using only a subset of them. This places higher demands on the encoder's extraction capabilities, allowing for a deeper understanding of speech features and resulting in a powerful feature extraction model.
[0072] For random masking, random sampling masking is performed according to a uniform distribution. Using a high masking rate (i.e., greater than 50%) avoids the task of easily predicting by inferring from unmasked neighboring vectors. A uniform distribution prevents potential center shifts (i.e., higher masking probability near the center of a speech segment). Furthermore, the highly sparse input creates conditions for training efficient feature encoders; very powerful encoders can be trained with minimal computation and memory, improving training efficiency by 2-3 times.
[0073] Step S103: Perform vector padding and position information embedding on the lossy speech feature vector to obtain the masked feature vector; including: assigning a shared learning vector or a fixed non-zero vector to the masked vector and the remaining vector; embedding the corresponding position information to obtain the masked feature vector.
[0074] Specifically, for the vector H output by the feature encoder, a unified, learnable vector I is added to the masked part based on the recorded position information, and position vector encoding is added to both the masked and unmasked parts to obtain the feature representation Z.
[0075] Step S104: Based on the feature vector after mask filling, predict the speech feature vector corresponding to the randomly masked part to obtain the complete speech feature vector;
[0076] Specifically, the vector Z obtained after masking is input into the transformer decoder for decoding to obtain the final vector representation Z. out :
[0077]
[0078]
[0079] in, This represents the complete speech feature vector after adding vector I to the vector H output by the encoder. Represents location encoding information, f bn f represents the normalization operation. transformer This represents the transformer decoder.
[0080] Preferably, in this embodiment, when using the wav2vec2 model as the initial model for training, the encoder and decoder of the wav2vec2 model can be designed with an asymmetric structure. Specifically, in this embodiment, the main function of the decoder is to restore the filled feature vector, making it as close as possible to the information before it was masked. Its structure is the same as the transformer encoder. The input is a vector composed of the unmasked feature vector, the padding vector, and the positional encoding; where the positional encoding can represent the positional information of the feature vector in the complete speech. Since the transformer decoder is only used during training to perform the speech reconstruction task, and the final recognition task is completed by the wav2vec encoder, the decoder architecture can be designed independently of the encoder design. Experiments can be conducted with a very small decoder, which is narrower and shallower than the encoder. For example, the encoder uses a 12-layer, 768-dimensional transformer structure, and the decoder can use 2 or 3 layers, with a size of 384 dimensions. Compared with the encoder, the computational cost per token in the decoder is less than 10%. Using this asymmetric design, all tokens are processed through a lightweight decoder, which greatly reduces the training time.
[0081] Step S105: Perform average pooling and softmax classification on the complete speech feature vector to obtain the label category distribution corresponding to the complete speech feature vector;
[0082] The output layer of this invention uses an average pooling operation followed by a softmax layer to predict the speaker label category distribution. The output vector dimension is the number of label categories in the training dataset, and the predicted vector is obtained after passing through the softmax layer.
[0083]
[0084] It should be noted that in the original Wav2vec2 model pre-training process, the model loss consists of both adversarial loss and diversity loss. However, in this embodiment, due to the high masking rate operation, predicting the masked information is too difficult for the model, preventing convergence. Furthermore, the ultimate goal of this invention is to obtain the speaker's speech feature vector representation, a text-independent task. This embodiment, based on the decoded complete speech feature vector, uses average pooling and a fully connected layer network to convert the complete speech feature vector into the label dimension corresponding to the task. Probability normalization is then performed using an activation function, and finally, the category information of the speech data is obtained by selecting the category corresponding to the highest confidence value.
[0085] Step S106: Based on the labels in the voiceprint recognition training dataset and the labels obtained by the classifier, the training random masking feature encoding model is iteratively updated through the loss function to obtain the trained random masking feature encoding model; based on the trained random masking feature encoding model, the random masking function is turned off to obtain the feature extraction model.
[0086] Specifically, this embodiment uses the following loss function for iterative loss updates to obtain the trained feature encoding model:
[0087]
[0088] Where x is the input speech sequence and y is the speaker category label.
[0089] In practical applications, the random masking function of the trained feature encoding model is turned off, thus obtaining the feature extraction model. The speech data to be recognized is input into the feature extraction model, and after CNN feature extraction and transformer encoding, a feature vector with the speaker's voiceprint features is obtained. Based on the obtained feature vector, the speaker's identity is identified and confirmed.
[0090] Step S2: Obtain the speech to be recognized, and extract features from the speech to be recognized using the feature extraction model to obtain the feature vector of the speech to be recognized.
[0091] Specifically, this embodiment uses existing speech recognition methods for speech recognition. The recognized speech is preprocessed and then input into a trained feature extraction model for feature extraction. The feature extraction model trained using a random masking training method has powerful feature extraction capabilities, enabling it to extract and encode features from the input speech to obtain a feature vector containing the user's voiceprint information. This vector is then used for subsequent cosine similarity calculation to verify the user's identity.
[0092] Step S3: Calculate the cosine similarity value between the feature vector of the speech to be identified and all registered speech in the user speech registration database; determine the user to which the speech to be identified belongs based on the cosine similarity value.
[0093] Specifically, the similarity value between the speech feature vector of the speech to be recognized and the speech feature vector in the speech registration database is calculated, and the user to which the registered speech feature vector with the highest similarity value and greater than the threshold belongs is selected as the user to which the speech to be recognized belongs.
[0094] In this embodiment, the speech to be identified should be the speech of a user in the speech registration database. For unregistered speakers, after feature extraction by the feature extraction model, the similarity score between the speech feature vector and the speech feature vector in the speech registration database will be less than the threshold, and then a message will be displayed indicating that there is no matching user.
[0095] Another embodiment of the present invention also discloses a computer device, the computer device including at least one processor and at least one memory communicatively connected to the processor;
[0096] The memory stores instructions that can be executed by a processor to implement the aforementioned voiceprint recognition method.
[0097] In summary, this invention employs a voiceprint recognition method trained using a random masking approach, enabling accurate recognition of discontinuous or lossy speech in various noisy environments. During training, the input speech undergoes high-density random masking and feature extraction to construct a lossy speech feature vector for training. The randomly masked portion of the speech features is reconstructed during training, resulting in a content-independent voiceprint recognition feature extraction model and method. This improves the robustness of voiceprint recognition, making it applicable to various voiceprint recognition scenarios and achieving accurate recognition of lossy speech.
[0098] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0099] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A voiceprint recognition method based on random masking training, characterized in that, Includes the following steps: A user voice registration library is obtained by registering multiple user voices through a pre-trained feature extraction model. The feature extraction model is obtained by constructing lossy speech feature vectors using a random masking method and then training them, including: Construct a voiceprint recognition training dataset, which includes voice sample data and labels representing the person to whom the voice data belongs; Lossy speech feature vectors are constructed using a random masking feature coding model as follows: feature extraction is performed on the input speech sample data to obtain a fixed-dimensional speech feature vector; the speech feature vector is randomly masked according to a uniform distribution, and the masking rate is greater than 50%; the positions of the masked vectors and the remaining vectors are recorded, and position information is embedded in the remaining vectors to obtain randomly masked lossy speech data; inter-frame attention is calculated on the randomly masked lossy speech data to obtain a lossy speech feature vector with temporal and contextual relationships. The lossy speech feature vector is padded with vectors and embedded with positional information to obtain a masked feature vector; the speech feature vector corresponding to the randomly masked part is predicted based on the masked feature vector to obtain a complete speech feature vector; average pooling and softmax classification are performed on the complete speech feature vector to obtain the label category distribution corresponding to the complete speech feature vector. The labels obtained from the voiceprint recognition training dataset and the classifier are iteratively updated using a loss function to obtain a trained random masking feature encoding model; based on the trained random masking feature encoding model, the feature extraction model is obtained by disabling the random masking function. The speech feature vector and the lossy speech feature vector are represented as follows: S=f bn (Wf cnn (X)+b); Where S is a fixed-dimensional speech feature vector, H is a lossy speech feature vector, and f cnn Denotes a CNN feature extractor, W∈R p*q A mapping from dimension p to dimension q. P represents the portion remaining after S is randomly masked. emb represent Location information, f bn f represents the normalization operation. transformer This represents the transformer encoder in wav2vec2; The step of performing vector padding and position information embedding on the lossy speech feature vector includes: assigning a shared learning vector or a fixed non-zero vector to the masked vector; embedding corresponding position information into the masked vector and the remaining vector to obtain the masked feature vector. The speech to be recognized is acquired, and the feature extraction model is used to extract features from the speech to be recognized to obtain the feature vector of the speech to be recognized. The feature vector of the speech to be identified is compared with the cosine similarity value of all registered speech in the user speech registration database; the user to whom the speech to be identified belongs is determined based on the cosine similarity value.
2. The voiceprint recognition method according to claim 1, characterized in that, The mask-padded feature vector is input into the decoder, and the complete speech feature vector is obtained using the following formula: in, This represents the speech feature vector after adding the padding vector. Represents location encoding information, f bn f represents the normalization operation. transformer This represents the transformer decoder.
3. The voiceprint recognition method according to claim 1, characterized in that, The predicted vector is obtained by performing average pooling and softmax classification on the complete speech feature vector. The loss function formula is: Where x is the input speech sequence and y is the speaker category label.
4. The voiceprint recognition method according to claim 1, characterized in that, The step of registering multiple user voices using a pre-trained feature extraction model to obtain a user voice registration library includes: inputting the voice of the user to be registered into the feature extraction model to obtain the voice feature vector of the user to be registered; labeling the voice feature vector of the user to be registered based on the user ID; and saving the voice feature vector of the user to be registered and the corresponding label to obtain the user voice registration library.
5. A computer device, characterized in that, It includes at least one processor and at least one memory communicatively connected to the processor; The memory stores instructions that can be executed by the processor to implement the voiceprint recognition method according to any one of claims 1-4.
Citation Information
Patent Citations
Noise-containing speech emotion recognition method based on deep learning
CN115035916A
Voiceprint detection model training method and voiceprint recognition method
CN115101077A
Training method of voiceprint recognition feature extraction model and voiceprint recognition system
CN115547344A