Speaker recognition method based on emotion transfer learning
By introducing emotional transfer learning and attention mechanisms in speaker recognition, the problem of degradation of recognition performance in traditional methods in the case of emotional domain mismatch is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510128716.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-04
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional speaker recognition methods perform poorly in the case of emotional domain mismatch, resulting in a decrease in recognition accuracy and robustness.
The speaker recognition method based on emotion transfer learning is adopted to transform emotionally related speaker embeddings into emotionally independent embeddings through emotional domain adaptation strategies, combining attention mechanisms and deep neural networks to improve the accuracy and robustness of recognition.
It significantly reduces the impact of emotional changes on speaker identification, improves speaker identification accuracy and robustness under emotional changes, and provides more reliable technical support for identity verification and security monitoring.
Smart Images

Figure CN119993213A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech processing technology, and in particular to a speaker recognition method based on emotion transfer learning, which is used to solve the problem of emotion domain mismatch in speaker recognition. Background Art
[0002] Speaker Recognition (SR) is a technology that identifies the identity of a speaker through a voice signal. As an important part of biometric technology, speaker recognition technology is widely used in identity authentication, security monitoring and other fields. Traditional speaker recognition methods mainly rely on the acoustic features of speech signals, such as Mel-frequency cepstral coefficients (MFCC) and linear prediction cepstral coefficients (LPCC). However, in real scenarios, the speaker's speech signal is affected by many factors, such as age, gender, environmental noise, etc., among which the change in emotional state is particularly significant. The emotional domain mismatch problem refers to the inconsistency of the speaker's emotional state in the registration voice and the test voice, which leads to a decrease in the performance of the speaker recognition model.
[0003] In order to overcome this challenge, researchers have proposed a variety of improvement methods, such as feature extraction based on deep learning, speaker model adaptation, etc. However, the recognition performance of these methods in emotional speech environments still needs to be improved. Therefore, the present invention proposes a speaker recognition method based on emotion transfer learning, which aims to improve the accuracy of speaker recognition under emotional changes by combining emotion transfer learning and attention mechanism. Summary of the invention
[0004] In speaker recognition tasks, the problem of emotion domain mismatch is a long-standing challenge. Specifically, when the speaker's emotional state in the registration speech and the test speech is inconsistent, traditional speaker recognition models often perform poorly. This emotion domain mismatch will cause the distribution of speaker embeddings to change, thereby reducing the recognition accuracy and robustness of the model. The present invention proposes a speaker recognition method based on emotion transfer learning, which aims to improve the accuracy and robustness of speaker recognition under conditions of emotion changes by combining emotion transfer learning and attention mechanism. This method needs to overcome the shortcomings of traditional speaker recognition methods in emotional speech environments, improve the accuracy and robustness of recognition, and provide more reliable technical support for identity authentication, security monitoring and other fields.
[0005] 1. The purpose of the present invention is to provide a speaker recognition method based on emotion transfer learning, which converts emotion-related speaker embedding into emotion-independent embedding through an emotion domain adaptation strategy, thereby solving the emotion domain mismatch problem. The specific technical solution includes
[0006] S1. Input speech signal S containing different emotional states, where S = {s 1 ,s 2 ,...,s N} is a speech dataset containing N samples, and each sample s i With speaker label and domain labels The domain label is used to distinguish the emotional domain category of the speech data, where the source domain consists of neutral speech, while the target domain contains a variety of emotional states, such as anger, excitement, panic, and sadness. The speech features are extracted using the X-Vector structure of the Kaldi toolkit to obtain the embedded representation X i , where X i =f(s i ), f is the feature extraction function;
[0007] S2, for X i Perform feature transformation to enhance speaker features: Input the embedding representation X obtained in step S1 i , transformed through the fully connected layer FC to get the speaker embedding
[0008] S3, deep speaker feature extraction: input the speaker embedding obtained in step S2 Use a 3-layer deep neural network DNN to extract deep speaker features and get
[0009] S4. Calculate the embedding cross-mapping: Input the speaker embedding obtained in step S2 Calculate the internal relationship matrix M i And apply nonlinear transformation to get the embedded cross-mapping feature The specific calculation is as follows:
[0010] S5. Combine deep features and cross-map features to get the final speaker embedding: Input the deep speaker features obtained in step S3 and the embedded cross-mapping features obtained in step S4 After the following calculation: Get the final speaker embedding
[0011] S6: Perform sentiment transfer learning to remove sentiment information: Input the final speaker embedding obtained in step S5 Use the gradient reversal layer (GRL) for sentiment transfer learning and process it through the fully connected layer to get the embedding without sentiment influence. where γ is the scaling factor, is the gradient of the domain classification loss with respect to the embedding;
[0012] Calculate the emotion-related speaker embedding: Input the embedding obtained in step S6 without emotion influence After the following calculation: Get sentiment-related speaker embedding Then we pass it to the domain classifier cls domain Input and according to Get domain classification results
[0013] Calculate the total loss and perform gradient backpropagation: Finally optimize the objective function L total L is calculated by the following formula: total =αL sp +βL domain , where L sp is the speaker classification loss, L domain is the domain classification loss, α and β are parameters that control the influence of the two losses;
[0014] S9, system test: input test voice signal S test According to steps S1-S6, the emotion-related speaker embedding of the speech to be tested is obtained and remove emotional embedding according to: Get the final speaker recognition result
[0015] 2. The method according to claim 1, characterized in that the speaker classification loss L in step S8 sp , the input is the speaker recognition result of the i-th sample and the speaker label of the i-th sample The specific calculation method is:
[0016] 3. The method according to claim 1, characterized in that the domain classification loss L in step S8 domain , the input is the domain classification result of the i-th sample and the domain label of the i-th sample The specific calculation method is:
[0017] The beneficial effects of the present invention are as follows: by introducing an emotion transfer framework, the present invention achieves the following beneficial effects: through an emotion domain classifier and a gradient reversal layer (GRL), the present invention can convert emotion-related speaker embeddings into emotion-independent embeddings, significantly reducing the impact of emotion changes on speaker recognition; by embedding a cross-mapping module and an attention mechanism compensation module, the present invention retains sufficient speaker identity information during the emotion domain adaptation process, avoiding the problem of information loss; the framework of the present invention is mainly composed of a fully connected layer and simple matrix operations, with low computational complexity, and is suitable for promotion in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is the overall flow chart of the present invention; DETAILED DESCRIPTION
[0019] The specific implementation of the present invention is divided into three main modules: speaker feature integration module, emotion transfer learning module and attention speaker identity compensation module. The implementation details of each module are described in detail below.
[0020] Part 1: Speaker Feature Integration Module
[0021] In the speaker feature integration module, the first step is to process the input speech signal S, where S = {s 1 ,s 2 ,...,s N} is a speech dataset containing N samples, and each sample s i With speaker label and domain labels The domain label is used to distinguish the emotional domain category of the speech data, where the source domain consists of neutral speech, while the target domain contains a variety of emotional states, such as anger, excitement, panic, and sadness. By extracting speech features using the X-Vector structure of the Kaldi toolkit, the embedded representation X i , where X i =f(s i ), f is the X-Vector feature extraction function. These embedded representations capture the key features of speech and lay the foundation for further processing.
[0022] Next, the extracted embedding representation is feature transformed to enhance the speaker characteristics. This step is achieved through a fully connected layer (FC) to transform the original embedding representation X i Convert to more representative speaker embeddings This conversion process aims to highlight the unique characteristics of the speaker and reduce interference from factors such as emotions:
[0023] Deep speaker features are further extracted through a deep neural network (DNN). Input the speaker embedding after the full connection layer conversion After three layers of DNN processing, deep speaker features are obtained
[0024] By computing the embedding cross-map, we capture the internal relations between speaker embeddings. This step is done by computing the internal relation matrix M i , and apply nonlinear transformation to get the embedded cross-mapping feature The specific calculation is as follows: These features can reflect the complex relationships between embeddings of different speakers.
[0025] Finally, the deep speaker features Cross-mapping features with embedding Combine to get the path speaker embedding This combination is achieved through a fully connected layer, which aims to combine the advantages of different features to generate a more representative speaker embedding:
[0026] Part 2: Emotion Transfer Learning Module
[0027] The sentiment transfer learning module enables speaker recognition across emotional states by converting emotion-dependent embeddings into emotion-independent embeddings while retaining speaker identity information. This allows the system to recognize speakers even when trained only with neutral speech data, thereby improving robustness to emotional changes. To achieve this goal, a sentiment classifier with a gradient reversal layer (GRL) is introduced to ensure emotion-invariant feature learning. GRL reverses the gradient during backpropagation, prompting the model to generate emotion-independent embeddings. Unlike traditional domain adaptation methods, this method is able to generalize to multiple emotional states and remove path speaker embeddings. The emotional information in where γ is the scaling factor, is the gradient of the domain classification loss with respect to the embedding. This process effectively reduces the impact of emotion on speaker recognition and improves the recognition accuracy.
[0028] Embedding to remove emotional influence Further processing is performed with fully connected layers to generate sentiment-related speaker embeddings Get a more precise representation of the speaker identity:
[0029] Finally, the emotion-related speaker is embedded Speaker Embedding with Path Combined, through a classifier cls consisting of two fully connected layers and an output layer domain Classify the embedding into discrete emotion categories and get the domain classification result
[0030] Part III: Attention Speaker Identity Compensation Module
[0031] The attention speaker identity compensation module aims to adapt to the emotional changes in speech while preserving the speaker identity information. To achieve speaker information compensation, Passed as additional input to the speaker classifier cls sp The speaker classifier consists of three fully connected layers, the last of which is combined with a softmax layer as the output layer. The classification layer input is Calculated and fed into the fully connected layer for speaker recognition.
[0032] Finally, the total loss is calculated and gradient backpropagation is performed to optimize the entire system. total The speaker classification loss L sp and domain classification loss L domain Composition, L sp The input is the speaker recognition result of the i-th sample and the speaker label of the i-th sample L domain The input is the domain classification result of the i-th sample and the domain label of the i-th sample L total =αL sp +βLdomain Among them, α and β are parameters that control the influence of the two losses. By minimizing the total loss, the system can better adapt to speaker recognition tasks under different emotional states and improve the robustness and accuracy of the system.
[0033] System test: Input test voice signal S test Through the above three modules, the emotion-related speaker embedding of the speech to be tested is obtained and remove emotional embedding And input speaker classifier cls sp ,according to: Get the final speaker recognition result Experimental design
[0034] Experimental Dataset: The IEMOCAP dataset used in the experiment is a multimodal database widely used in emotion recognition research, released by the University of Southern California in 2008. The dataset contains about 12 hours of audio, video, and text data, recording the emotional expressions of 10 professional actors (5 men and 5 women) in two-way conversations. The data is divided into 5 sessions, each of which contains the interaction of two actors, covering a variety of emotional states such as anger, happiness, sadness, neutrality, excitement, frustration, and surprise. IEMOCAP not only provides high-quality emotion annotations, but also contains detailed motion capture data of facial expressions, gestures, and voice, supporting multimodal emotion analysis. This dataset has important application value in the fields of emotion computing, speech emotion recognition, natural language processing, etc., and provides researchers with rich resources to develop and evaluate emotion recognition models.
[0035] Experimental strategy: It is divided into three stages. In the migration / pre-training stage, 15,000 samples containing multiple emotions (Anger, Excited, Neutral, Sad, Panic) are used for model pre-training; in the registration stage, the first 70% of the segments (1,192 samples) from the neutral emotional speech of 10 speakers are selected to register the speaker identity information; in the testing stage, the last 30% of the segments (1,684 samples) of all emotional speech of the same batch of speakers are used for testing, covering four emotions: Anger, Excited, Neutral, and Sad, to verify the recognition ability of the model in cross-emotional scenarios.
[0036] Performance indicators: The present invention uses three key indicators, namely, equal error rate (EER), detection cost function (DCF) and precision, to evaluate the performance of the speaker recognition system. The equal error rate (EER) represents the error rate when the false acceptance rate (FAR) and the false rejection rate (FRR) are equal. The lower the EER, the better the system performance. FAR refers to the proportion of non-target speakers that the system mistakenly identifies as target speakers. FRR refers to the proportion of target speakers that the system mistakenly rejects. EER is determined by adjusting the threshold so that FAR and FRR are equal.
[0037] The detection cost function (DCF) combines FAR and FRR and evaluates the cost weight in actual applications. The lower the DCF, the better the system performance. The calculation formula of DCF is: DCF=C miss ×P target ×FRR+C fa ×(1-P target )×FAR Among them C miss is the cost of omission, C fa is the false alarm cost, P target is the prior probability of the target speaker.
[0038] Precision reflects the system's ability to correctly identify the speaker. The higher the precision, the higher the recognition accuracy of the system. The calculation formula for precision is: Where True Positives (TP) are the number of target speakers correctly identified by the system, and False Positives (FP) are the number of non-target speakers incorrectly identified by the system.
[0039] The present invention includes two hyperparameters: the parameter α represents the weight of the speaker recognition loss, which is fixed to 1 in the experiment; the parameter β controls the influence of the transfer learning loss, and its value is adjusted from 0.6 to 1.1 with a step size of 0.1. In the speaker classifier, the number of units in the three fully connected layers is set to 500, 500 and 50+n, where 50 represents the number of speakers included in the data set in the pre-training stage, and n represents the number of speakers in the registration and test sets. All attention fully connected layers in the attention module have the same number of units as the corresponding fully connected layers in the emotion and classifier as additional layers. In the training parameter setting, both pre-training and registration training include 1000 training rounds (Epochs), and the learning rate is 0.0001. The present invention adopts the Adam optimizer, and the batch size is 512. Experimental Results
[0040] Experimental results under different classification models: This paper comprehensively compares the proposed speaker recognition method based on emotion transfer learning with several baseline methods and state-of-the-art speaker recognition (SR) methods, covering three categories of methods: CNN-based methods, GAN-based methods, and embedding-based methods. CNN-based methods, including ResNet18, VGG16, and DenseNet121, directly use spectrograms as input for speaker recognition without utilizing pre-trained models or emotional speech data during the training process. GAN-based methods, such as StarGAN and IREC-StarGAN, adopt data augmentation strategies by generating emotional speech from neutral speech, and use the distribution of emotional data learned in the pre-training stage to enrich the training process, thereby improving the speaker recognition model. Embedding-based methods, including X-Vector, ECAPA-TDNN, RawNet3, HubertSR, and ETP, first fine-tune the emotional speech data in the pre-training stage to capture the emotional representation. These fine-tuned models are then used to extract embeddings for enrollment and test data, and processed by deep neural networks to complete the speaker recognition task.
[0041] The results in Table 1 show that the present invention is significantly superior to existing methods, which is specifically manifested in the significant reduction of equal error rate and detection cost function and the significant improvement of precision. These results prove that the present invention effectively solves the limitations of GAN-based and embedding-based methods, and establishes its superiority in the task of emotional speaker recognition by robustly preserving speaker identity information in different emotional scenes. Table 1 Performance comparison of different SR methods on IEMOCAP Table 1 Performance comparative results on IEMOCAP with different SRmethods.
Claims
1. A speaker recognition method based on emotion transfer learning, comprising the following steps: S1, input speech signal S containing different emotional states, where S = {s1, s2, ..., s N } is a speech dataset containing N samples, and each sample s i With speaker label and domain labels The domain label is used to distinguish the emotional domain category of the speech data. The source domain consists of neutral speech, while the target domain contains a variety of emotional states, such as anger, excitement, panic, and sadness. The X-Vector structure of the Kaldi toolkit is used to extract speech features and obtain the embedded representation X i , where X i =f(s i ), f is the feature extraction function; S2, for X i Perform feature transformation to enhance speaker features: Input the embedding representation X obtained in step S1 i , transformed through the fully connected layer FC to get the speaker embedding S3, deep speaker feature extraction: input the speaker embedding obtained in step S2 Use a 3-layer deep neural network DNN to extract deep speaker features and get S4. Calculate the embedding cross-mapping: Input the speaker embedding obtained in step S2 Calculate the internal relationship matrix M i And apply nonlinear transformation to get the embedded cross-mapping feature The specific calculation is as follows: S5: Combine the deep features and cross-map features to get the final speaker embedding, and input the deep speaker features obtained in step S3 and the embedded cross-mapping features obtained in step S4 After the following calculation: Get the final speaker embedding S6: Perform sentiment transfer learning to remove sentiment information: Input the final speaker embedding obtained in step S5 Use the gradient reversal layer (GRL) for sentiment transfer learning and process it through the fully connected layer to get the embedding without sentiment influence. where γ is the scaling factor, is the gradient of the domain classification loss with respect to the embedding; S7, calculate the emotion-related speaker embedding: input the embedding obtained in step S6 without the emotion influence After the following calculation: Get sentiment-related speaker embedding Then we pass it to the domain classifier cls domain Input and according to Get domain classification results S8. Calculate the total loss and perform gradient backpropagation: Finally optimize the objective function L total L is calculated by the following formula: total =αL sp +βL domain , where L sp is the speaker classification loss, L domain is the domain classification loss, α and β are parameters that control the influence of the two losses; S9, system test: input test voice signal S test According to steps S1-S6, the emotion-related speaker embedding of the speech to be tested is obtained and remove emotional embedding And input speaker classifier cls sp ,according to: Get the final speaker recognition result 2. The method according to claim 1, characterized in that The speaker classification loss L in step S8 sp , the input is the speaker recognition result of the i-th sample and the speaker label of the i-th sample The specific calculation method is:
3. The method according to claim 1, characterized in that The domain classification loss L in step S8 domain , the input is the domain classification result of the i-th sample and the domain label of the i-th sample The specific calculation method is: