A momentum contrast learning based voiceprint recognition method and device
Patent Information
- Application Number
- CN202311463284.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-11-06
AI Technical Summary
即使将目标用户的训练数据大幅提升,也很难符合 GMM 的充分训练要求
[0037]本发明提供一种基于动量对比学习的声纹识别方法和设备,应用MOCO主干网络来提取声音的特征,使用了队列的数据结构思想来减轻训练网络时对电脑储存的依赖,并且使用动量更新的方法来保证数据样本的一致性,从而保证了训练的有效性;并且结合了文本嵌入的方式,实现了音频和文本的多模态学习;训练MLP网络来更好地完成说话人识别的下游任务;使用了无标签数据的对比学习在降低了对训练数据依赖同时还增强模型的泛化能力,可以在复杂的环境下实现更准确的声纹识别。
Smart Images

Figure CN117457005B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voiceprint recognition, and more specifically to a voiceprint recognition method and device based on momentum contrast learning. Background Technology
[0002] With the advent of the digital age, identity recognition and authentication have become an indispensable technology in daily life. Commonly used technologies include key authentication, face authentication, iris authentication, and fingerprint authentication. Voiceprint authentication has always been a relatively immature technology, and there is currently no mature open-source or closed-source voiceprint recognition authentication technology on the market.
[0003] The background of voiceprint recognition technology can be traced back to the 1960s, when researchers began to try to identify speakers using voice features. The key to this technology is that human voice features are unique; even for the same sentence, different people will pronounce it differently.
[0004] Theoretically, there are two main methods for voiceprint recognition and authentication. The first is feature extraction-based methods: these methods first extract features from the sound signal, such as the sound spectrum, vocal tract length, and formants. Then, these features are combined to construct a variable with unique characteristics similar to human voice timbre for comparison, thereby identifying and distinguishing the voiceprints of different individuals. The second method is based on neural networks: in recent years, with the development of deep learning technology, neural networks have been applied to various aspects of NLP, CV, etc., and many speaker recognition models based on neural networks have emerged. These models typically input sound signals into the neural network and learn to extract high-level features for recognition.
[0005] In practical applications, voiceprint recognition technology has been widely adopted. In telephone banking and customer service, voiceprint recognition can be used to verify customer identity and improve user experience. In the field of intelligent voice assistants, voiceprint recognition can differentiate between different users, thereby providing personalized services. Different users can enjoy customized suggestions, reminders, and other features specific to them.
[0006] In terms of technical methods, the mainstream voiceprint recognition technologies currently employed mainly include the following: MFCC (MelFrequency Cepstral Coefficients) is a commonly used feature extraction method for voiceprint recognition. It simulates the human ear's perception mechanism of sound, converting sound signals into a series of frequency coefficients to extract useful features; GMM-UBM (Gaussian Mixture Model - Universal Background Model) is a commonly used statistical modeling method that models sound features as a Gaussian mixture model and compares them with a universal background model to achieve speaker recognition; i-vector is a method for representing speaker features. It projects acoustic features into a low-dimensional space, encoding the speaker's acoustic features into a vector for constructing a speaker recognition system; with the development of deep learning technology, deep learning models such as Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN) have achieved significant results in voiceprint recognition for extracting sound features; Siamese networks are used to learn speaker similarity, achieving good results in tasks such as speaker verification by comparing the similarity between two sound signals to determine whether they belong to the same speaker.
[0007] However, when extracting a speaker's voiceprint features, noise often interferes with the process, blurring acoustic details and affecting model performance and recognition rate. Furthermore, Gaussian Mixture Models (GMMs) using Gaussian components also have limitations. As the GMM size increases, the model's representational ability improves. However, the number of parameters also increases, meaning that Gaussian components require a large amount of data for training to simulate the distribution. More data is needed to drive the GMM training; otherwise, its performance will deteriorate. Even significantly increasing the training data for the target user is insufficient to meet the requirements for adequate GMM training. When the amount of data is limited, Gaussian Mixture Models are prone to overfitting and cannot be applied to multiple scenarios. Summary of the Invention
[0008] To address the technical problems existing in the prior art, this invention provides a voiceprint recognition method and device based on momentum contrastive learning. This method can obtain a large amount of voiceprint sample data from different environments. During training, the recognition model captures more effective features, enhancing generalization and effectively preventing interference from noise factors. It can perform voiceprint recognition tasks well in complex environments, thereby improving recognition efficiency.
[0009] The first objective of this invention is to provide a voiceprint recognition method based on momentum contrast learning.
[0010] A second objective of this invention is to provide a computer device.
[0011] The first objective of this invention can be achieved by adopting the following technical solution:
[0012] A voiceprint recognition method based on momentum contrast learning, the method comprising:
[0013] S1. Obtain training voiceprint samples and test voiceprint samples;
[0014] S2. Collect text data related to training voiceprint samples and test voiceprint samples, construct a synchronous sound-text correspondence dataset, and generate diverse training voiceprint samples based on the synchronous sound-text correspondence dataset using data augmentation methods.
[0015] S3. Use training voiceprint samples to train the MOCO model backbone network to obtain the trained MOCO model backbone network, and use test voiceprint samples to verify the training effect of the MOCO model backbone network.
[0016] S4. Randomly select two MOCO model backbone networks trained with task audio input, and output the feature vectors of the two audios respectively; map the text data to a high-dimensional space to obtain word vectors, and multiply the word vectors with the audio vectors to obtain the composite vector of the two audios;
[0017] S5. Input the synthesized vectors of the two audios into the MLP network for advanced feature extraction, and then use the MLP network to determine the similarity of the synthesized vectors of the two audios.
[0018] S6. Obtain labeled voiceprint sample data, fine-tune the parameters of the MLP network using the labeled voiceprint sample data to obtain the fine-tuned MLP network, and use the fine-tuned MLP network and the trained MOCO backbone network as the voiceprint recognition model.
[0019] S7. Input the two audio files into the voiceprint recognition model to determine whether the two audio files come from the same person, and output the voiceprint recognition result.
[0020] In a preferred embodiment, the preprocessing of the speech dataset includes:
[0021] The audio data is denoised by using Fourier transform to remove some types of noise and DenoisingAutoencoders to denoise the audio.
[0022] Standardize the sampling rate, bit depth, and format of audio data;
[0023] Segment the speech data.
[0024] In a preferred embodiment, the step of collecting text data related to training and testing voiceprint samples to construct a synchronized sound-text mapping dataset includes:
[0025] Collect text data related to training voiceprint samples and test voiceprint samples, clean and standardize the text data to obtain text representations of training voiceprint samples and test voiceprint samples, match the sound representation and text representation of the training voiceprint samples of the same speaker to obtain a synchronized sound-text correspondence dataset.
[0026] In the preferred technical solution, the step of training the MOCO model backbone network using training voiceprint samples to obtain the trained MOCO model backbone network includes:
[0027] The MOCO model backbone network is trained using training voiceprint samples. An appropriate mini-batch size and agent task are selected. Positive and negative sample pairs of training voiceprint samples are generated using any training voiceprint sample as the original audio data sample. A dynamic data dictionary is constructed based on the positive and negative sample pairs of training voiceprint samples.
[0028] Multiple audio samples are randomly selected from the training voiceprint samples of the dynamic data dictionary. For each audio sample, the feature vector of the audio signal is extracted using the MFCC method, and then propagated forward through the encoder of the ResNet-18 network to generate a query vector. For other audio samples, the feature vectors of their audio signals are extracted using the MFCC method, and then propagated forward through the momentum encoder of the ResNet-18 network to output a key vector set, which is then added to a queue.
[0029] The InfoNCE loss function is used for feature learning, and the encoder and momentum encoder parameters are updated to obtain the trained MOCO model backbone network.
[0030] In the preferred technical solution, the step of using the InfoNCE loss function for feature learning and updating the encoder parameters and momentum encoder parameters includes:
[0031] Obtain key vectors generated from positive samples from the dynamic data dictionary, and randomly select a number of key vectors generated from negative samples. Perform a query-key lookup on the key vectors generated from positive samples, the key vectors generated from negative samples, and a corresponding query vector, and input them into the loss function to calculate the loss. Then perform backpropagation to directly update the encoder parameter and perform momentum-based update of the encoder parameter.
[0032] In a preferred embodiment, the step of inputting the synthesized vectors of the two audio files into an MLP network for advanced feature extraction and determining the similarity of the synthesized vectors of the two audio files includes:
[0033] The synthesized vectors of the two audio segments are input into an MLP network. The linear layer of the MLP network performs high-level feature extraction on the synthesized vectors. The cosine similarity function is used to calculate the cosine similarity between the two synthesized vectors. The cosine similarity is compared with a pre-set threshold. If the cosine similarity is greater than or equal to the threshold, it is determined that the two audio segments are from the same person; otherwise, it is determined that the two audio segments are not from the same person.
[0034] The second objective of this invention can be achieved by adopting the following technical solution:
[0035] A computer device includes a processor and a memory for storing a processor-executable program, wherein when the processor executes the program stored in the memory, it implements the aforementioned voiceprint recognition method based on momentum contrast learning.
[0036] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0037] This invention provides a voiceprint recognition method and device based on momentum contrastive learning. It employs a MOCO backbone network to extract sound features, utilizes a queue data structure to reduce reliance on computer storage during network training, and uses momentum updates to ensure data sample consistency, thereby guaranteeing training effectiveness. Furthermore, it combines text embedding to achieve multimodal learning of audio and text. An MLP network is trained to better perform downstream tasks such as speaker recognition. The use of contrastive learning with unlabeled data reduces dependence on training data while enhancing the model's generalization ability, enabling more accurate voiceprint recognition in complex environments. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0039] Figure 1 This is a flowchart of a voiceprint recognition method based on momentum contrast learning in an embodiment of the present invention;
[0040] Figure 2 This is a schematic diagram of the framework for training the MOCO backbone network in an embodiment of the present invention. Detailed Implementation
[0041] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments, and the implementation of the present invention is not limited thereto. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0042] Example 1:
[0043] This invention discloses a voiceprint recognition method based on momentum contrastive learning. It employs a MOCO network as the backbone to extract sound features, utilizes a queue data structure to reduce reliance on computer storage during network training, and uses momentum updates to ensure data sample consistency, thereby guaranteeing training effectiveness. Furthermore, it combines text embedding to achieve multimodal learning of audio and text. While freezing the backbone network, an MLP network is trained to better perform downstream tasks such as speaker recognition. The use of contrastive learning with unlabeled data reduces dependence on training data while enhancing the model's generalization ability, enabling more accurate voiceprint recognition in complex environments.
[0044] like Figure 1 As shown, the present invention provides a voiceprint recognition method based on momentum contrast learning, which includes the following steps:
[0045] Step S1: Obtain training voiceprint samples and test voiceprint samples
[0046] S11. Obtain multiple speech datasets from the deep learning community.
[0047] Specifically, acquiring training voiceprint samples requires obtaining some mature and reliable speech datasets from the deep learning community, such as THCHS30 and the open-source version of AISHELL-1. For example, THCHS30 is an open-source Chinese speech database published by the Center for Speech and Language Technology (CSLT) at Tsinghua University, containing over 10,000 speech files, approximately 40 hours of Chinese speech data, mainly consisting of texts and poems, all in female voices. AISHELL-1, an open-source Chinese speech dataset released by Beijing Hill Technology Co., Ltd., contains approximately 178 hours of open-source data. This dataset contains the voices of 400 people from different regions of China with different accents. Recordings were made in a quiet indoor environment using high-fidelity microphones, with sampling reduced to 16kHz. Through professional speech annotation and rigorous quality checks, the manual transcription accuracy reached over 95%. These voice data represent diverse pronunciation styles, speech rates, and sound qualities.
[0048] S12. Preprocess the speech dataset and divide the preprocessed speech dataset into training voiceprint samples and test voiceprint samples according to a predetermined ratio.
[0049] Specifically, preprocessing the speech dataset includes:
[0050] To denoise the speech data, Fourier transform is first used to remove some types of noise, especially periodic noise, and then Denoising Autoencoders are used to denoise the sound to remove environmental noise.
[0051] Standardize the sampling rate, bit depth, and format of the speech data to ensure data consistency. The sampling rate is set to 16kHz, which is suitable for speech sampling. The bit depth is 8 bits. The format is unified by converting all MP3 files from the open-source dataset to WAV format, and the audio files are split into 120-second segments.
[0052] Speech data is segmented for better analysis and processing.
[0053] Step S2: Collect text data related to training voiceprint samples and test voiceprint samples, construct a synchronized sound-text correspondence dataset, and generate diverse training voiceprint samples based on the synchronized sound-text correspondence dataset using data augmentation methods.
[0054] S21. Collect text data related to the training and testing voiceprint samples. Clean and standardize the text data to obtain the text representations of the training and testing voiceprint samples. Match the voice representations and text representations of the training voiceprint samples of the same speaker to obtain a synchronized voice-text mapping dataset. The synchronized voice-text mapping dataset includes the voice representation and corresponding text representation of the same speaker.
[0055] Specifically, text data related to voiceprint samples is collected. This text data includes information corresponding to the voice data, such as the speaker's gender, intonation, and vocal quality. To improve the quality of the text representation, the text data is cleaned and standardized to ensure consistency and comparability. Then, words matching the descriptions of each voice data point are extracted from the text data and paired. This synchronous representation construction allows for capturing the strong correlation between voice and text, contributing to improved accuracy and robustness of voiceprint recognition. The data in the synchronous voice-text mapping dataset is unlabeled. Contrastive learning using unlabeled data reduces dependence on training data while enhancing the model's generalization ability, enabling more accurate voiceprint recognition in complex environments.
[0056] S22. Based on the synchronous sound-text correspondence dataset, data augmentation methods are used to generate diverse training voiceprint samples to ensure the diversity of the dataset and to ensure coverage of multiple different speakers, languages, and contexts. The balance of the dataset is maintained to avoid speaker imbalance and to ensure a relatively balanced number of positive pairs for each speaker. To improve the robustness of the voiceprint recognition system, multiple data augmentation methods are applied to the synchronous sound-text correspondence dataset to create more diverse and rich training samples. Preferably, the multiple data augmentation methods include speech rate variation, noise addition, and text perturbation, wherein:
[0057] Speech rate variation is achieved by adjusting the playback speed of the audio samples to create multiple audio samples with different speaking speeds. This helps our model adapt to speakers with different speaking speeds, improving the adaptability of the recognition.
[0058] Noise addition involves adding environmental noise to a clean human voice dataset. For example, the sound of wind and waves at the beach can be added to a maritime law enforcement system.
[0059] Text perturbation involves altering the text information to a certain extent, such as deleting or replacing certain characters, to increase the model's robustness to text incompleteness.
[0060] S3. Train the MOCO model backbone network using training voiceprint samples to obtain the trained MOCO model backbone network, and use test voiceprint samples to verify the training effect of the MOCO model backbone network.
[0061] S31. Train the MOCO model backbone network using training voiceprint samples. Select an appropriate mini-batch size and agent task. Use any training voiceprint sample as the original audio data sample to generate positive and negative sample pairs of the training voiceprint sample. Construct a dynamic data dictionary based on the positive and negative sample pairs of the training voiceprint samples. The dynamic data dictionary includes positive and negative sample pairs and is used for feature learning.
[0062] Specifically, training the MOCO model backbone network only requires an audio dataset, and the mini-batch size is determined by the available GPU memory. The proxy task used is the widely used instance discrimination. A mini-batch of training audio data serves as the original sample for generating query vectors. For any original audio data sample of the mini-batch size, data augmentation is applied to the original audio data to obtain the corresponding positive sample. For example, data augmentation results in audio 1, audio 2, audio 3, etc., all from the same person. All other audio data in the mini-batch size, excluding the original audio data itself, are mixed with the augmented data to form the negative sample of the original audio data. For example, all other audio data in the mini-batch size, excluding the original audio data itself, are mixed with the augmented data to obtain audio 1, audio 2, audio 3, etc., which are not from the same person. Audio 1 and audio 2, audio 3, etc., are unrelated. Positive and negative samples are used together as positive and negative sample pairs to construct the data dictionary.
[0063] like Figure 2 The diagram illustrates the framework for training the MOCO backbone network. MOCO (Momentum Contrast) is a self-supervised learning method used to train neural networks to learn useful feature representations. MOCO training does not require manually labeled data; instead, it learns from unlabeled data, making it perform exceptionally well on large-scale datasets. This is invaluable for avoiding the time and cost of labeled data. The MOCO model is trained through contrastive learning, which encourages the model to learn useful feature representations by maximizing the similarity metric between positive sample pairs while minimizing the similarity metric between negative sample pairs. This helps the model map similar samples to similar feature spaces.
[0064] S32. Randomly extract multiple audio samples from the training voiceprint samples of the dynamic data dictionary. For each audio sample, use the MFCC method to extract the feature vector of the audio signal, and then propagate it forward through the encoder of the ResNet-18 network to generate a query vector. For other audio samples, use the MFCC method to extract the feature vector of their audio signals, and then propagate it forward through the momentum encoder of the ResNet-18 network to output a key vector set. Add the key vector set to the queue.
[0065] like Figure 2As shown, audio samples are randomly selected from the training voiceprint samples of the dynamic data dictionary, such as audio 1, audio 2, audio 3, audio 4, etc. Audio 1 first has its feature vector extracted using the MFCC method, and then it is forward-propagated through the encoder of the ResNet-18 network to generate a query vector. Audio 2, audio 3, audio 4, etc. have their feature vectors extracted using the MFCC method, and then forward-propagated through the momentum encoder of the ResNet-18 network to output a key vector set, which is then queued.
[0066] The audio signal first passes through the MFCC function to extract its feature vector. The feature vector then passes through a CNN convolutional layer to adjust its dimensions to match the ResNet-18 input. It then undergoes forward propagation, which passes through a series of convolutional layers, pooling layers, and fully connected layers in sequence, ultimately generating an output vector.
[0067] MFCC (Mel-Frequency Cepstral Coefficients) is a method for feature extraction from audio signals, commonly used in speech recognition and audio processing tasks. The MFCC method applies Discrete Cosine Transform (DCT) to convert the audio signal from the frequency domain into cepstral coefficients, retaining only the first few (usually 13) cepstral coefficients as MFCC features. These cepstral coefficients are used to represent the characteristics of the audio signal.
[0068] ResNet-18 (Residual Network-18) is a deep convolutional neural network used for image classification and computer vision tasks. It is a member of the ResNet family of models, with a relatively shallow depth, yet still achieving excellent performance. ResNet-18 consists of multiple basic blocks, each containing several convolutional layers and an identity shortcut. The overall structure of the ResNet-18 network includes: an input layer, which accepts the input image, typically a 224x224 pixel color image; convolutional and pooling layers, starting with a 7x7 convolutional layer followed by max-pooling layers to reduce the resolution. These layers help extract low-level features of the image; and basic blocks, each consisting of four basic blocks, with each block containing multiple convolutional layers. Each basic block's structure includes: Convolutional Layer 1, a 3x3 convolutional layer for feature extraction; Convolutional Layer 2, a 3x3 convolutional layer for further feature extraction; Identity Shortcut: To ensure smooth information flow, each basic block has a skip connection that adds the input to the convolutional layer's output; Global Average Pooling: At the top of the network, global average pooling is performed on the output of the last basic block, reducing the feature map size to 1x1; Fully Connected Layer: Finally, a fully connected layer maps the feature vector to a class probability distribution. Typically, the fully connected layers in ResNet-18 output a probability distribution equal to the number of classes in the task.
[0069] S33. Use the InfoNCE loss function to learn features, update the encoder parameters and momentum encoder parameters, and obtain the trained MOCO model backbone network.
[0070] Feature learning is performed using the InfoNCE loss function to update the encoder and momentum encoder parameters. The InfoNCE loss function is as follows:
[0071] infoNCE(q, k) = - log( exp(q · k / τ) / Σ exp(q · ki / τ) )
[0072] Where q is the query vector, k is the set of positive samples, which usually includes samples related to the query vector, τ is a temperature parameter, which is usually used to adjust the similarity scale, exp(x) represents the exponential function of x, Σ represents the summation operation, which is performed on all positive samples ki in the set of positive samples k, and q · k represents the dot product or inner product between q and k.
[0073] The system retrieves key vectors generated from positive samples from a dynamic data dictionary and randomly selects key vectors generated from a certain number (typically 4096) of negative samples. These two vectors are then used to perform a query-key lookup with a corresponding query vector, and the result is fed into the loss function to calculate the loss. Backpropagation is then performed to directly update the encoder parameters, while momentum-based updates are applied to the encoder parameters.
[0074] θ' = m * θ' + (1 - m) * θ
[0075] Where θ' represents the parameters of the target encoder, θ represents the parameters of the current encoder, and m is the momentum parameter, which is usually a value less than 1 and is used to specify the update speed.
[0076] To ensure consistency in the dynamic data dictionary, the value of m should be set very close to 1, with 0.999 showing good training performance. This process is then repeated until the accuracy on the test set no longer improves, resulting in the trained MOCO model backbone network. Afterward, two sets of training audio are input to optimize the feature representation before proceeding to the next steps.
[0077] S34. To test the training effect of the trained MOCO model backbone network using test voiceprint samples, the test voiceprint samples are input into the model, and the model output is compared with the sample labels to calculate the model's accuracy.
[0078] S4. Randomly select two MOCO model backbone networks trained with task audio input, and output the feature vectors of the two audios respectively; map the text data to a high-dimensional space to obtain word vectors, and multiply the word vectors with the audio vectors to obtain the composite vector of the two audios.
[0079] Multimodal learning is achieved by amplifying the differences between audio vectors from different speakers or reducing the differences between audio vectors from the same speaker through speech embedding and the introduction of text embedding. Methods such as Word2Vec are used to map text data to a high-dimensional space to obtain word vectors while preserving semantic connections. MFCC is first used to extract features from two audio segments, which are then processed by a backbone network to obtain feature vectors. The word vectors are multiplied by the audio vectors to create new vector representations, thus expanding or reducing the differences in the high-dimensional space.
[0080] like Figure 2As shown, in this example, the Word2Ve text encoder is used to embed the text dataset to obtain the corresponding word vectors T1, T2, T3, ..., Tn, etc. The text dataset is the text dataset related to the voiceprint samples trained in step 2, a typeof voice (type: high, low, man, woman, etc.). These texts are mapped to a high-dimensional space while still preserving the semantic relationships between the texts. The two audio WAV files are first processed by MFCC to extract features, and then processed by the backbone network to obtain the feature vectors of the two audio segments. The feature vectors of audio 1 are v1, v2, v3, ..., vn, etc., and the feature vectors of audio 2 are v1`, v2`, v3`, ..., vn`, etc. The obtained word vectors are multiplied by the feature vectors of the two audio segments (i.e., the subscript parts of the T vector and the v vector are multiplied), resulting in two new vectors: composite vector 1 and composite vector 2. By adding multimodal features, the difference between the two audio vectors of different people in the high-dimensional space is better amplified, or the difference between the two audio vectors of the same person in the high-dimensional space is reduced.
[0081] S5. Input the synthesized vectors of the two audios into the MLP network for advanced feature extraction and determine the similarity of the synthesized vectors of the two audios.
[0082] Two audio data segments are preprocessed and their features extracted by the backbone network to transform them into vectors. These vectors capture important information about the audio, such as sound frequencies and temporal features. Further, the synthesized vectors from these two audio segments are input into an MLP (Multilayer Perception) network. The MLP is one of the most fundamental neural network models in deep learning. An MLP consists of multiple linear layers and activation functions (typically ReLU). High-level feature extraction is performed through the linear layers of the MLP network. These linear layers map the input vectors to a high-dimensional space, and the non-linearity introduced by the activation functions helps extract higher-level features and representations to better capture the complexity of speech.
[0083] The cosine similarity function is used to measure the similarity between two synthesized vectors. The results are compared with a preset threshold, and the cosine similarity is used to determine whether the two audio segments belong to the same person.
[0084] Cosine Similarity = dot(A, B) / (||A|| * ||B||)
[0085] Where C represents cosine similarity, A represents the synthesized vector of one audio, and B represents the synthesized vector of another audio.
[0086] This will produce a value between -1 and 1, representing the similarity between the two vectors. The closer the value is to 1, the higher the similarity; the closer the value is to -1, the lower the similarity. This measures the probability that two audio clips are from the same person. Further, a cosine similarity function is used to measure the similarity between the two vectors and compared to a pre-set threshold. If the cosine similarity is greater than or equal to the threshold, we can determine that the two audio clips are from the same person; otherwise, they are considered not to be from the same person.
[0087] S6. Obtain labeled voiceprint sample data, fine-tune the parameters of the MLP network using the labeled voiceprint sample data to obtain the fine-tuned MLP network, and use the fine-tuned MLP network and the trained MOCO backbone network as the voiceprint recognition model.
[0088] Specifically, labeled voiceprint samples from the synchronized sound-text mapping dataset are associated with the speaker's voiceprint labels to obtain labeled voiceprint sample data. Voiceprint labels for training voiceprint samples are determined by the source of the sound data or its text representation. The sound representation and text representation of each voiceprint sample are associated with the speaker's voiceprint label. These labels can be used to guide the model in learning the connection between sound and voiceprint identity during training. Voiceprint labels typically include the speaker's identity information, ensuring the accuracy of recognition. Labeled voiceprint sample data can be manually annotated, specifying the person corresponding to the specific audio and the corresponding text data, such as labels like "man," "woman," "high-pitched," and "low-pitched." For example, two input audio clips may be labeled to indicate whether they are from the same person; if they are, the label value is 1; if they are not, the label value is 0.
[0089] Labeled voiceprint sample data can be used to fine-tune the parameters of the obtained MoCo feature representation model and downstream subdivided voiceprint recognition models, enabling the entire model to learn the ability to match sound representations with specific speaker identities. The multilayer perceptron (MLP) is fine-tuned using labeled voiceprint sample data to train the MLP network parameters, obtaining the best-performing MLP network parameters and the previously obtained MoCo backbone network as the final model. A small amount of labeled audio data is used to adjust the MLP network parameters to improve voiceprint recognition performance.
[0090] The MOCO backbone network is used to extract sound features, while the MLP network is used for classification and discrimination. After the sound passes through the MOCO backbone network to extract sound features, the obtained features are... Figure 2 The audio feature vector and the corresponding text vector are multiplied to obtain the final synthesized vector. The synthesized vector is then passed through an MLP network for classification and discrimination, thus completing the voiceprint recognition task.
[0091] Speaker recognition is treated as a downstream task. This is achieved by segmenting an audio file and inputting it into the model, freezing the backbone network, and training only the MLP network. This results in an MLP network specifically designed for this downstream task. The parameters of the best-performing MLP network and the previously obtained MOCO backbone network are then used as the speaker recognition model.
[0092] S7. Input the two audio segments into the voiceprint recognition model to determine whether the two audio segments come from the same person, and output the voiceprint recognition result.
[0093] The voiceprint recognition model includes a MOCO backbone network, a multiplication function of sound and text vectors, and an MLP network. After the model is processed, two audio clips are input into the voiceprint recognition model. The audio clips first pass through the MOCO backbone network to extract features, then the resulting text vectors are multiplied, and finally the synthesized vector is input into the MLP network for processing. The result is then determined using 0 and 1, where 1 indicates that the two audio clips come from the same person, and 0 indicates that the two audio clips do not come from the same person.
[0094] Specifically, both audio files are converted to WAV format, and then features are extracted using a MOCO network. A text encoder encodes the text data based on preset parameters, combining this with the audio feature vectors extracted by the MOCO network. The MLP network outputs the similarity result and compares it to a set threshold to determine if the two audio clips belong to the same person. Since the model's training data is primarily unlabeled and very large, the model is well-trained and does not suffer from overfitting. Therefore, it can accurately extract audio features even in complex environments, leading to accurate identification. Alternatively, an audio clip can be segmented, and a simple loop can input each segment in pairs into a voiceprint recognition model. The loop then identifies which segments belong to the same person, thus identifying how many people are speaking in an audio clip and what each person is saying. As a downstream task of voiceprint recognition, this approach is accurate and can extract the main speaker's voice features and perform voiceprint recognition even in complex environments (such as noisy background sounds like wind, rain, and other people's noise).
[0095] Example 2:
[0096] This embodiment provides a computer device, which may be a server, computer, etc., including a processor, memory, input device, display, and network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. When the processor executes the computer programs stored in the memory, it implements a voiceprint recognition method based on momentum contrast learning as described in Embodiment 1 above, including:
[0097] S1. Obtain training voiceprint samples and test voiceprint samples;
[0098] S2. Collect text data related to training voiceprint samples and test voiceprint samples, construct a synchronous sound-text correspondence dataset, and generate diverse training voiceprint samples based on the synchronous sound-text correspondence dataset using data augmentation methods.
[0099] S3. Use training voiceprint samples to train the MOCO model backbone network to obtain the trained MOCO model backbone network, and use test voiceprint samples to verify the training effect of the MOCO model backbone network.
[0100] S4. Randomly select two MOCO model backbone networks trained with task audio input, and output the feature vectors of the two audios respectively; map the text data to a high-dimensional space to obtain word vectors, and multiply the word vectors with the audio vectors to obtain the composite vector of the two audios;
[0101] S5. Input the synthesized vectors of the two audios into the MLP network for advanced feature extraction, and then use the MLP network to determine the similarity of the synthesized vectors of the two audios.
[0102] S6. Obtain labeled voiceprint sample data, fine-tune the parameters of the MLP network using the labeled voiceprint sample data to obtain the fine-tuned MLP network, and use the fine-tuned MLP network and the trained MOCO backbone network as the voiceprint recognition model.
[0103] S7. Input the two audio files into the voiceprint recognition model to determine whether the two audio files come from the same person, and output the voiceprint recognition result.
[0104] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.
Claims
1. A voiceprint recognition method based on momentum contrast learning, characterized in that, Includes the following steps S1. Obtain training voiceprint samples and test voiceprint samples; S2. Collect text data related to training voiceprint samples and test voiceprint samples, construct a synchronous sound-text correspondence dataset, and generate diverse training voiceprint samples based on the synchronous sound-text correspondence dataset using data augmentation methods. S3. Use training voiceprint samples to train the MOCO model backbone network to obtain the trained MOCO model backbone network, and use test voiceprint samples to verify the training effect of the MOCO model backbone network. The MOCO model backbone network is trained using training voiceprint samples. The mini-batch size and agent task are selected. Positive and negative sample pairs of training voiceprint samples are generated using any training voiceprint sample as the original audio data sample. A dynamic data dictionary is constructed based on the positive and negative sample pairs of training voiceprint samples. Multiple audio samples are randomly selected from the training voiceprint samples of the dynamic data dictionary. For each audio sample, the feature vector of the audio signal is extracted using the MFCC method and propagated forward through the encoder of the ResNet-18 network to generate a query vector. For other audio samples, the feature vectors of their audio signals are extracted using the MFCC method and propagated forward through the momentum encoder of the ResNet-18 network to output a key vector set, which is then queued. The InfoNCE loss function is used for feature learning, and the encoder and momentum encoder parameters are updated to obtain the trained MOCO model backbone network. The InfoNCE loss function is as follows: infoNCE(q, k) = - log( exp(q · k / τ) / Σ exp(q · ki / τ) ) Where q is the query vector, k is the set of positive samples, τ is the temperature parameter used to adjust the similarity scale, exp(x) represents the exponential function of x, Σ represents the summation operation, which is performed on all positive samples ki in the set of positive samples k, and q · k represents the dot product or inner product between q and k; The step of using the InfoNCE loss function for feature learning and updating the encoder and momentum encoder parameters includes: Obtain key vectors generated from positive samples from the dynamic data dictionary, and randomly select a number of key vectors generated from negative samples. Perform a query-key lookup on the key vectors generated from positive samples, the key vectors generated from negative samples, and a corresponding query vector, and input them into the loss function to calculate the loss. Then perform backpropagation to directly update the encoder parameter and perform momentum-based update of the encoder parameter. S4. Randomly select two task audios as inputs to the trained MOCO model backbone network, and output the feature vectors of the two audios respectively; map the text data to a high-dimensional space to obtain word vectors, and multiply the word vectors with the audio vectors to obtain the composite vector of the two audios; S5. Input the synthesized vectors of the two audios into the MLP network for advanced feature extraction, and then use the MLP network to determine the similarity of the synthesized vectors of the two audios. S6. Obtain labeled voiceprint sample data, fine-tune the parameters of the MLP network using the labeled voiceprint sample data to obtain the fine-tuned MLP network, and use the fine-tuned MLP network and the trained MOCO backbone network as the voiceprint recognition model. S7. Input the two audio files into the voiceprint recognition model to determine whether the two audio files come from the same person, and output the voiceprint recognition result.
2. The voiceprint recognition method based on momentum contrast learning according to claim 1, characterized in that, The acquisition of training voiceprint samples and test voiceprint samples includes: Obtain multiple speech datasets from the deep learning community; The speech dataset is preprocessed, and the preprocessed dataset is divided into training voiceprint samples and test voiceprint samples according to a predetermined ratio.
3. The voiceprint recognition method based on momentum contrast learning according to claim 2, characterized in that, The preprocessing of the speech dataset includes The speech data is denoised by using Fourier transform to remove some types of noise and DenoisingAutoencoders to denoise the speech data. Standardize the sampling rate, bit depth, and format of speech data; Segment the speech data.
4. The voiceprint recognition method based on momentum contrast learning according to claim 1, characterized in that, The text data related to the collected and trained voiceprint samples and test voiceprint samples are used to construct a synchronized sound-text mapping dataset, including: Collect text data related to training voiceprint samples and test voiceprint samples, clean and standardize the text data to obtain text representations of training voiceprint samples and test voiceprint samples, match the sound representation and text representation of the training voiceprint samples of the same speaker to obtain a synchronized sound-text correspondence dataset.
5. The voiceprint recognition method based on momentum contrast learning according to claim 1, characterized in that, The step of inputting the synthesized vectors of the two audio files into an MLP network for high-level feature extraction and determining the similarity of the synthesized vectors of the two audio files includes: The synthesized vectors of the two audio segments are input into an MLP network. The linear layer of the MLP network performs high-level feature extraction on the synthesized vectors. The cosine similarity function is used to calculate the cosine similarity between the two synthesized vectors. The cosine similarity is compared with a pre-set threshold. If the cosine similarity is greater than or equal to the threshold, it is determined that the two audio segments are from the same person; otherwise, it is determined that the two audio segments are not from the same person.
6. The voiceprint recognition method based on momentum contrast learning according to claim 5, characterized in that, The cosine similarity function is: C = dot(A, B) / (||A|| * ||B||) Where C represents cosine similarity, A is the synthesized vector of one audio, and B is the synthesized vector of another audio.
7. A computer device comprising a processor and a memory for storing a processor-executable program, characterized in that, When the processor executes the program stored in the memory, it implements the voiceprint recognition method based on momentum contrast learning as described in any one of claims 1-6.
Citation Information
Patent Citations
Voiceprint segmentation method, apparatus and device, and readable storage medium
CN112201275A
Voiceprint recognition model training method, voiceprint recognition method, voiceprint recognition device and voiceprint recognition equipment
CN115424621A