Speech emotion recognition method and recognition system based on double adversarial learning

By adopting a dual adversarial learning method in speech emotion recognition, the speaker information and content information in the voice signal are removed, and the problem of low recognition accuracy in the prior art is solved, and a higher accuracy rate of speech emotion recognition is achieved.

CN120015014AActive Publication Date: 2025-05-16QIZHITU TECHNOLOGY (GUANGZHOU) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510143057.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-16
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The existing speech emotion recognition methods are difficult to effectively remove the speaker information and content information in the voice signal, resulting in a decrease in recognition accuracy.

Method used

Using a dual adversarial learning method, by training the adversarial speaker classifier and the adversarial phoneme classifier, the speaker information and content information in the voice signal are removed, thereby extracting features containing only emotional information for emotional classification.

Benefits of technology

It effectively improves the accuracy of speech and emotional recognition and reduces the interference of speaker information and content information on emotional recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015014A_ABST
    Figure CN120015014A_ABST
Patent Text Reader

Abstract

The invention discloses a voice emotion recognition method and recognition system based on double adversarial learning, and relates to the technical field of voice signal processing, and the method comprises the steps: obtaining a voice signal, carrying out the preprocessing of the voice signal, and extracting a WavLM feature from the preprocessed voice signal through a WavLM pre-training model in an emotion classifier; and respectively sending the extracted WavLM features into an emotion encoder, an adversarial phoneme classifier and an adversarial speaker classifier, removing speaker information and content information in the to-be-recognized voice signal through double adversarial learning, and obtaining an emotion category of the to-be-recognized voice signal through the emotion classifier. According to the method, adversarial learning is carried out on the speaker classifier and the phoneme classifier respectively, speaker information and content information in the voice signals are removed, so that features only containing emotion information are extracted for voice emotion recognition, and the accuracy of voice emotion recognition is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech signal processing, and more specifically, to a speech emotion recognition method and recognition system based on dual adversarial learning. Background Art

[0002] In recent years, with the rapid development of speech signal processing technology, speech emotion recognition has received extensive attention as an important research direction in the field of human-computer interaction. However, speech signals usually contain a large amount of speaker information and content information, which will interfere with the emotion recognition task and reduce the recognition ability of the model. Therefore, an effective technical means is needed to extract speech features that only contain emotional information, so as to improve the accuracy of speech emotion recognition.

[0003] Existing speech emotion recognition methods usually directly use the features of speech signals for emotion classification, lacking an effective mechanism to remove interfering information related to the speaker and content. Therefore, how to remove speaker information and content information to achieve higher-precision speech emotion recognition is an urgent problem to be solved. Summary of the invention

[0004] In order to solve the above technical problems, the present invention proposes a speech emotion recognition method and recognition system based on dual adversarial learning. By training an adversarial speaker classifier and an adversarial phoneme classifier, the speaker information and content information in the speech signal are removed, thereby extracting features containing only emotion information for emotion classification.

[0005] The present invention provides a speech emotion recognition method based on dual adversarial learning, comprising the following steps:

[0006] Acquire speech signals and preprocess them, and use the WavLM pre-trained model in the sentiment classifier to extract WavLM features from the preprocessed speech signals;

[0007] The extracted WavLM features are sent to the emotion encoder, adversarial phoneme classifier and adversarial speaker classifier respectively, and the cross entropy loss of the emotion classifier, adversarial phoneme classifier and adversarial speaker classifier is calculated respectively;

[0008] The three calculated cross entropy losses are added together to obtain a total loss function, and the total loss function is used to simultaneously train the sentiment classifier, the adversarial phoneme classifier, and the adversarial speaker classifier;

[0009] The speech signal to be recognized is respectively imported into the trained emotion classifier, adversarial phoneme classifier and adversarial speaker classifier. The speaker information and content information in the speech signal to be recognized are removed through double adversarial learning, and the emotion category of the speech signal to be recognized is obtained through the emotion classifier.

[0010] In this solution, the speech signal is obtained and preprocessed, and the WavLM pre-trained model in the sentiment classifier is used to extract WavLM features from the preprocessed speech signal. Specifically:

[0011] Obtain a large amount of speech signals with phoneme annotations, emotion annotations, and speaker annotations, perform frequency domain analysis on the speech signals, obtain the frequency band component distribution corresponding to the speech signals, determine the corresponding frequency segment according to the frequency band component distribution, and configure bandpass filtering according to the frequency segment to remove signals that do not meet the frequency requirements;

[0012] Down-sample the speech signal after bandpass filtering, perform digital filtering using a biorthogonal wavelet basis, obtain the denoised speech signal, calculate the wavelet entropy of the denoised speech signal, and obtain the interval between the maximum wavelet entropy and the minimum wavelet entropy to generate a threshold interval;

[0013] Using the threshold interval to perform fuzzy speech discrimination on the de-noised speech signal, when the wavelet entropy of the speech signal is not within the threshold interval, it is eliminated, and all speech signals are traversed to obtain the pre-processed speech signal;

[0014] Constructing a WavLM pre-training model, in the training of the WavLM pre-training model, using a convolutional encoder and a Transformer encoder speech signal for feature encoding, randomly transforming the input speech signal, randomly covering a preset proportion of the speech signal, and predicting the label corresponding to the covered position;

[0015] After the training is completed, the WavLM pre-training model is used to extract the probability distribution of the label sequence corresponding to the preprocessed speech signal, and the probability distribution is used as the WavLM feature. The WavLM feature includes emotion information, phoneme information and speaker information.

[0016] In this scheme, the sentiment classifier consists of a WavLM pre-trained model, a sentiment encoder, a fully connected layer, and a softmax classification layer;

[0017] The adversarial phoneme classifier consists of a gradient inversion, a phoneme encoder, a fully connected layer and a softmax classification layer;

[0018] The adversarial speaker classifier is composed of gradient inversion, speaker encoder, fully connected layer and softmax classification layer;

[0019] The obtained WavLM features are used as the input of the sentiment encoder, adversarial phoneme classifier and adversarial speaker classifier, and the cross entropy losses of the sentiment classifier, adversarial factor classifier and adversarial speaker classifier are calculated respectively.

[0020] In this solution, the adversarial phoneme classifier is specifically:

[0021] The acquired WavLM features are imported into the phoneme encoder, where an initial convolution is performed through one convolution layer, followed by downsampling using two convolution layers to reduce the feature size. The downsampled features are then extracted using three identical residual modules, and a multi-head self-attention mechanism is introduced in feature extraction to obtain phoneme encoding.

[0022] The obtained phoneme code is imported into the discriminator, and the negative coefficient is multiplied by the error control back propagation through the gradient reversal layer, so that the network learning objectives before and after the gradient reversal layer are opposite, realizing the adversarial learning of phoneme features;

[0023] Use the fully connected layer and Softmax activation function to classify and predict the phoneme information in the WavLM feature.

[0024] In this solution, the parameters of the sentiment classifier and the adversarial speaker classifier are configured by sharing features based on the adversarial phoneme classifier, and the WavLM features corresponding to the annotated speech signal are used to simultaneously perform supervised training on the adversarial phoneme classifier, the sentiment classifier and the adversarial speaker classifier;

[0025] Import the acquired WavLM features into the configured adversarial speaker classifier to classify and predict the speaker information in the WavLM features;

[0026] The content information and speaker information are marked according to the phoneme information label and speaker information label in the WavLM feature, and the marked content information and speaker information are removed.

[0027] In this scheme, the cross entropy loss of the sentiment classifier, adversarial factor classifier, and adversarial speaker classifier is calculated separately, specifically:

[0028] The WavLM features corresponding to the annotated speech signals are divided into training sets and test sets according to the ratio, the framework parameters and learning rates of the emotion classifier, adversarial factor classifier and adversarial speaker classifier are initialized, and the training samples in the training set are input into the three classifiers for training;

[0029] In the training process of the three classifiers, the sentiment cross entropy loss between the output of the sentiment classifier and the sentiment annotation, the phoneme cross entropy loss between the output of the adversarial phoneme classifier and the phoneme annotation, and the speaker cross entropy loss between the output of the adversarial speaker classifier and the speaker annotation are calculated based on the subordinate relationship between the training samples and the label information;

[0030] The emotion cross entropy loss, phoneme cross entropy loss and speaker cross entropy loss are added together to construct a total loss function to supervise the training of the three classifiers, the network parameters of the three classifiers are iteratively updated according to the total loss in the forward propagation, and the classification performance is tested using the test set. When the performance test results meet the preset standards, the training of the three classifiers is completed.

[0031] In this solution, a pre-processed speech signal to be recognized is obtained, and the speech signal to be recognized is respectively introduced into a trained emotion classifier, an adversarial phoneme classifier, and an adversarial speaker classifier;

[0032] Through double adversarial learning, the speaker information and content information in the speech signal to be recognized are removed, and the emotion label probability distribution corresponding to the speech signal to be recognized is obtained through the fully connected layer and Softmax function in the emotion classifier, and the emotion category of the speech signal to be recognized is output according to the probability distribution.

[0033] The second aspect of the present invention provides a speech emotion recognition system based on dual adversarial learning, the system comprising: a speech signal input module, an emotion classifier module, an adversarial phoneme classifier module, an adversarial speaker classifier module, a classifier training module and a speech emotion output module;

[0034] The voice signal input module is responsible for acquiring the voice signal to be recognized and preprocessing the voice signal to be recognized;

[0035] The emotion classifier module is responsible for extracting the WavLM features of the speech signal to be recognized, and obtaining the emotion category of the speech signal to be recognized according to the WavLM features;

[0036] The adversarial phoneme classifier module is responsible for removing content information of the speech signal to be recognized;

[0037] The adversarial speaker classifier module is responsible for removing speaker information of the speech signal to be recognized;

[0038] The classifier training module is responsible for training the emotion classifier, the adversarial phoneme classifier and the adversarial speaker classifier using the annotated speech signal;

[0039] The speech output module is responsible for outputting the emotion category information corresponding to the speech signal to be recognized without the content information and the speaker information, and displaying it in a preset manner.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] The present invention performs adversarial learning on the speaker classifier and the phoneme classifier respectively, removes the speaker information and content information in the speech signal, thereby extracting features containing only emotion information for speech emotion recognition, effectively improving the accuracy of speech emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments or exemplary embodiments of the present invention, the drawings required for use in the embodiments or exemplary descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained according to the drawings without paying creative work.

[0043] Figure 1 A flow chart of a speech emotion recognition method based on dual adversarial learning is shown;

[0044] Figure 2 The process and structural diagram of speech emotion recognition are shown;

[0045] Figure 3 A schematic diagram of the process of classifier training is shown;

[0046] Figure 4 A block diagram of a speech emotion recognition system based on dual adversarial learning is shown. DETAILED DESCRIPTION

[0047] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

[0048] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.

[0049] Figure 1 A flowchart of a speech emotion recognition method based on dual adversarial learning is shown.

[0050] like Figure 1 As shown, in a first embodiment of the present invention, a speech emotion recognition method based on dual adversarial learning is provided, comprising:

[0051] S102, acquiring and preprocessing a speech signal, and extracting WavLM features from the preprocessed speech signal using a WavLM pre-trained model in a sentiment classifier;

[0052] S104, sending the extracted WavLM features to the emotion encoder, the adversarial phoneme classifier and the adversarial speaker classifier respectively, and calculating the cross entropy loss of the emotion classifier, the adversarial phoneme classifier and the adversarial speaker classifier respectively;

[0053] S106, adding the three calculated cross entropy losses to obtain a total loss function, and using the total loss function to simultaneously train the sentiment classifier, the adversarial phoneme classifier, and the adversarial speaker classifier;

[0054] S108, importing the speech signal to be recognized into the trained emotion classifier, adversarial phoneme classifier and adversarial speaker classifier respectively, removing the speaker information and content information in the speech signal to be recognized through double adversarial learning, and obtaining the emotion category of the speech signal to be recognized through the emotion classifier.

[0055] It should be noted that a large number of speech signals with phoneme annotations, emotion annotations and speaker annotations are obtained, and the speech signals are analyzed in the frequency domain to obtain the frequency band component distribution corresponding to the speech signal, and the corresponding frequency segment is determined by the frequency band component distribution. According to the frequency segment, a bandpass filter is configured to remove signals that do not meet the frequency requirements; in order to reduce the amount of calculation of the model, the speech signal after the bandpass filter is downsampled, and a biorthogonal wavelet basis is used for digital filtering. The translation and scaling of the wavelet function are controlled by setting the translation amount and scale of the wavelet transform to filter out high-frequency noise and DC noise. Due to environmental factors and speaker factors, fuzzy speech may be generated. The wavelet entropy can reflect the time and energy spectrum of the signal. The wavelet entropy is used to distinguish which are fuzzy speech signals, so as to eliminate them and enhance the quality of the speech signal. The denoised speech signal is obtained, the wavelet entropy of the denoised speech signal is calculated, and the interval between the maximum wavelet entropy and the minimum wavelet entropy is obtained to generate a threshold interval. The calculation formula of the wavelet entropy S(a) is:

[0056] S(a)=-∫P(a,b)log(P(a,b))db

[0057] Among them, P(a,b) represents the wavelet energy probability distribution, a represents the total time variable, and b represents the effective time variable; the threshold interval is used to perform fuzzy speech discrimination on the noise-reduced speech signal, and when the wavelet entropy of the speech signal is not in the threshold interval, it is eliminated, and the preprocessed speech signal is obtained after traversing all speech signals.

[0058] Construct the WavLM pre-trained model, which includes a convolutional encoder and a Transformer encoder. The convolutional encoder has 7 layers, each of which includes a time domain convolution layer, a layer normalization layer, and a GELU activation function layer. In the Transformer encoder, the relative position is introduced into the calculation of the attention network using gated relative position encoding to better model local information. The WavLM pre-trained model converts continuous signals into discrete labels through the Kmeans method and models the discrete labels as targets.

[0059] In the training of the WavLM pre-training model, the convolution encoder and the Transformer encoder speech signal are used for feature encoding, the input speech signal is randomly transformed, and then a preset proportion of the speech signal is randomly covered, and the label corresponding to the covered position is predicted, such as mixing the two input speech signals, or adding background noise. After that, about 50% of the audio signal is randomly covered, and the label corresponding to the covered position is predicted at the output. After the training is completed, the WavLM pre-training model is used to extract the probability distribution of the label sequence corresponding to the preprocessed speech signal, and the probability distribution is used as the WavLM feature, and the WavLM feature contains emotional information, phoneme information and speaker information.

[0060] It should be noted that if Figure 2 As shown in the figure, the sentiment classifier is composed of a WavLM pre-trained model, a sentiment encoder, a fully connected layer and a softmax classification layer; the adversarial phoneme classifier is composed of a gradient inversion, a phoneme encoder, a fully connected layer and a softmax classification layer; the adversarial speaker classifier is composed of a gradient inversion, a speaker encoder, a fully connected layer and a softmax classification layer; the acquired WavLM features are used as inputs of the sentiment encoder, the adversarial phoneme classifier and the adversarial speaker classifier, respectively, and the cross entropy losses of the sentiment classifier, the adversarial factor classifier and the adversarial speaker classifier are calculated respectively.

[0061] The obtained WavLM features are imported into the phoneme encoder, where an initial convolution is performed through one convolution layer, followed by downsampling using two convolution layers to reduce the feature size. In the downsampling part, several convolution layers are used to gradually reduce the size of the input vector, and the phoneme features are initially extracted. Three identical residual modules are used to extract the phoneme features from the downsampled features. Each residual module has two convolution layers and a direct connection structure. The residual blocks are connected in series to further extract the phoneme features. A multi-head self-attention mechanism is introduced in feature extraction. The purpose of introducing multi-head attention is to model the relative dependency between elements at different positions in the WavLM feature sequence, extract more significant features, and obtain phoneme encoding based on the output of the multi-head attention structure; the obtained phoneme encoding is imported into the discriminator, and the gradient reversal layer uses a negative coefficient to multiply the error to control the back propagation, so that the network learning objectives before and after the gradient reversal layer are opposite, realizing the adversarial learning of phoneme features. The gradient reversal layer is equivalent to an identity transformation function during forward propagation. During back propagation, the gradient is multiplied by a negative coefficient to control back propagation. The features of adversarial learning are input into the fully connected layer, and the output layer is activated using the Softmax activation function to classify and predict the phoneme information in the WavLM features.

[0062] Based on the adversarial phoneme classifier, the parameters of the emotion classifier and the adversarial speaker classifier are configured by sharing features, and the WavLM features corresponding to the annotated speech signal are used to perform supervised training on the adversarial phoneme classifier, the emotion classifier and the adversarial speaker classifier simultaneously; the acquired WavLM features are imported into the configured adversarial speaker classifier, and the speaker information in the WavLM features is classified and predicted; the content information and the speaker information are marked according to the phoneme information labels and the speaker information labels in the WavLM features, and the marked content information and the speaker information are eliminated.

[0063] like Figure 3 As shown, the cross entropy losses of the sentiment classifier, adversarial factor classifier, and adversarial speaker classifier are calculated respectively. The WavLM features corresponding to the annotated speech signal are divided into a training set and a test set according to the ratio, the framework parameters and learning rates of the sentiment classifier, adversarial factor classifier, and adversarial speaker classifier are initialized, and the training samples in the training set are input into the three classifiers for training; during the training process of the three classifiers, the sentiment cross entropy loss between the output of the sentiment classifier and the sentiment annotation, the phoneme cross entropy loss between the output of the adversarial phoneme classifier and the phoneme annotation, and the speaker cross entropy loss between the output of the adversarial speaker classifier and the speaker annotation are calculated based on the subordinate relationship between the training samples and the label information. The cross entropy loss L e The formal expression is:

[0064]

[0065] Among them, n represents the total number of training samples, x i represents the i-th training sample, y i Represents the label information corresponding to the training sample, θ k represents the encoder shared feature configuration parameters, θ f represents the parameters of the corresponding classifier e, e is the sentiment classifier, adversarial phoneme classifier and adversarial speaker classifier, and P represents the probability distribution representation;

[0066] The emotion cross entropy loss, phoneme cross entropy loss and speaker cross entropy loss are added together to construct a total loss function to supervise the training of the three classifiers, the network parameters of the three classifiers are iteratively updated according to the total loss in the forward propagation, and the classification performance is tested using the test set. When the performance test results meet the preset standards, the training of the three classifiers is completed.

[0067] According to an embodiment of the present invention, a preprocessed speech signal to be recognized is obtained, and the speech signal to be recognized is respectively introduced into a trained emotion classifier, an adversarial phoneme classifier and an adversarial speaker classifier; speaker information and content information in the speech signal to be recognized are removed through double adversarial learning, and the emotion label probability distribution corresponding to the speech signal to be recognized is obtained through the fully connected layer and the Softmax function in the emotion classifier, and the emotion category of the speech signal to be recognized is output according to the probability distribution.

[0068] Figure 4 A block diagram of a speech emotion recognition system based on dual adversarial learning is shown.

[0069] The second embodiment of the present invention provides a speech emotion recognition system based on dual adversarial learning, the system comprising: a speech signal input module 401, an emotion classifier module 402, an adversarial phoneme classifier module 403, an adversarial speaker classifier module 404, a classifier training module 405 and a speech emotion output module 406;

[0070] The voice signal input module is responsible for acquiring the voice signal to be recognized and preprocessing the voice signal to be recognized;

[0071] The emotion classifier module is responsible for extracting the WavLM features of the speech signal to be recognized, and obtaining the emotion category of the speech signal to be recognized according to the WavLM features;

[0072] The adversarial phoneme classifier module is responsible for removing content information of the speech signal to be recognized;

[0073] The adversarial speaker classifier module is responsible for removing speaker information of the speech signal to be recognized;

[0074] The classifier training module is responsible for training the emotion classifier, the adversarial phoneme classifier and the adversarial speaker classifier using the annotated speech signal;

[0075] The speech output module is responsible for outputting the emotion category information corresponding to the speech signal to be recognized without the content information and the speaker information, and displaying it in a preset manner.

[0076] In the several embodiments provided in the present application, it should be understood that the disclosed method can be implemented in other ways. The system embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as: multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms. In addition, the functional modules in the various embodiments of the present invention can be all integrated into one processing module, or each module can be a separate module, or two or more modules can be integrated into one module; the above-mentioned integrated modules can be implemented in the form of hardware or in the form of hardware plus software functional modules.

[0077] If the above-mentioned integrated module of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0078] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A speech emotion recognition method based on dual adversarial learning, characterized in that: The following steps are involved: Acquire speech signals and preprocess them, and use the WavLM pre-trained model in the sentiment classifier to extract WavLM features from the preprocessed speech signals; The extracted WavLM features are sent to the emotion encoder, adversarial phoneme classifier and adversarial speaker classifier respectively, and the cross entropy loss of the emotion classifier, adversarial phoneme classifier and adversarial speaker classifier is calculated respectively; The three calculated cross entropy losses are added together to obtain a total loss function, and the total loss function is used to simultaneously train the sentiment classifier, the adversarial phoneme classifier, and the adversarial speaker classifier; The speech signal to be recognized is respectively imported into the trained emotion classifier, adversarial phoneme classifier and adversarial speaker classifier. The speaker information and content information in the speech signal to be recognized are removed through double adversarial learning, and the emotion category of the speech signal to be recognized is obtained through the emotion classifier.

2. The speech emotion recognition method based on dual adversarial learning according to claim 1, characterized in that: The speech signal is obtained and preprocessed, and the WavLM pre-trained model in the sentiment classifier is used to extract WavLM features from the preprocessed speech signal. Specifically: Obtain a large amount of speech signals with phoneme annotations, emotion annotations, and speaker annotations, perform frequency domain analysis on the speech signals, obtain the frequency band component distribution corresponding to the speech signals, determine the corresponding frequency segment according to the frequency band component distribution, and configure bandpass filtering according to the frequency segment to remove signals that do not meet the frequency requirements; Down-sample the speech signal after bandpass filtering, perform digital filtering using a biorthogonal wavelet basis, obtain the denoised speech signal, calculate the wavelet entropy of the denoised speech signal, and obtain the interval between the maximum wavelet entropy and the minimum wavelet entropy to generate a threshold interval; Using the threshold interval to perform fuzzy speech discrimination on the de-noised speech signal, when the wavelet entropy of the speech signal is not within the threshold interval, it is eliminated, and all speech signals are traversed to obtain the pre-processed speech signal; Constructing a WavLM pre-training model, in the training of the WavLM pre-training model, using a convolutional encoder and a Transformer encoder speech signal for feature encoding, randomly transforming the input speech signal, randomly covering a preset proportion of the speech signal, and predicting the label corresponding to the covered position; After the training is completed, the WavLM pre-training model is used to extract the probability distribution of the label sequence corresponding to the preprocessed speech signal, and the probability distribution is used as the WavLM feature. The WavLM feature includes emotion information, phoneme information and speaker information.

3. The speech emotion recognition method based on dual adversarial learning according to claim 1 is characterized in that: The sentiment classifier consists of a WavLM pre-trained model, a sentiment encoder, a fully connected layer and a softmax classification layer; The adversarial phoneme classifier consists of a gradient inversion, a phoneme encoder, a fully connected layer and a softmax classification layer; The adversarial speaker classifier is composed of gradient inversion, speaker encoder, fully connected layer and softmax classification layer; The obtained WavLM features are used as the input of the sentiment encoder, adversarial phoneme classifier and adversarial speaker classifier, and the cross entropy losses of the sentiment classifier, adversarial factor classifier and adversarial speaker classifier are calculated respectively.

4. The speech emotion recognition method based on dual adversarial learning according to claim 3 is characterized in that: The adversarial phoneme classifier is specifically: The acquired WavLM features are imported into the phoneme encoder, where an initial convolution is performed through one convolution layer, followed by downsampling using two convolution layers to reduce the feature size. The downsampled features are then extracted using three identical residual modules, and a multi-head self-attention mechanism is introduced in feature extraction to obtain phoneme encoding. The obtained phoneme code is imported into the discriminator, and the negative coefficient is multiplied by the error control back propagation through the gradient reversal layer, so that the network learning objectives before and after the gradient reversal layer are opposite, realizing the adversarial learning of phoneme features; Use the fully connected layer and Softmax activation function to classify and predict the phoneme information in the WavLM feature.

5. The speech emotion recognition method based on dual adversarial learning according to claim 3 is characterized in that: Based on the adversarial phoneme classifier, parameters of the sentiment classifier and the adversarial speaker classifier are configured by sharing features, and the adversarial phoneme classifier, the sentiment classifier and the adversarial speaker classifier are simultaneously supervised trained using WavLM features corresponding to the annotated speech signal; Import the acquired WavLM features into the configured adversarial speaker classifier to classify and predict the speaker information in the WavLM features; The content information and speaker information are marked according to the phoneme information label and speaker information label in the WavLM feature, and the marked content information and speaker information are removed.

6. The speech emotion recognition method based on dual adversarial learning according to claim 3 is characterized in that: The cross entropy losses of the sentiment classifier, adversarial factor classifier, and adversarial speaker classifier are calculated separately, specifically: The WavLM features corresponding to the annotated speech signals are divided into training sets and test sets according to the ratio, the framework parameters and learning rates of the emotion classifier, adversarial factor classifier and adversarial speaker classifier are initialized, and the training samples in the training set are input into the three classifiers for training; In the training process of the three classifiers, the sentiment cross entropy loss between the output of the sentiment classifier and the sentiment annotation, the phoneme cross entropy loss between the output of the adversarial phoneme classifier and the phoneme annotation, and the speaker cross entropy loss between the output of the adversarial speaker classifier and the speaker annotation are calculated based on the subordinate relationship between the training samples and the label information; The emotion cross entropy loss, phoneme cross entropy loss and speaker cross entropy loss are added together to construct a total loss function to supervise the training of the three classifiers, the network parameters of the three classifiers are iteratively updated according to the total loss in the forward propagation, and the classification performance is tested using the test set. When the performance test results meet the preset standards, the training of the three classifiers is completed.

7. The speech emotion recognition method based on dual adversarial learning according to claim 1, characterized in that: Obtaining the preprocessed speech signal to be recognized, and importing the speech signal to be recognized into the trained emotion classifier, adversarial phoneme classifier and adversarial speaker classifier respectively; Through double adversarial learning, the speaker information and content information in the speech signal to be recognized are removed, and the emotion label probability distribution corresponding to the speech signal to be recognized is obtained through the fully connected layer and Softmax function in the emotion classifier, and the emotion category of the speech signal to be recognized is output according to the probability distribution.

8. A speech emotion recognition system based on dual adversarial learning, characterized in that: Implementing the speech emotion recognition method based on dual adversarial learning as described in any one of claims 1 to 7, the system comprises: a speech signal input module, an emotion classifier module, an adversarial phoneme classifier module, an adversarial speaker classifier module, a classifier training module and a speech emotion output module; The voice signal input module is responsible for acquiring the voice signal to be recognized and preprocessing the voice signal to be recognized; The emotion classifier module is responsible for extracting the WavLM features of the speech signal to be recognized, and obtaining the emotion category of the speech signal to be recognized according to the WavLM features; The adversarial phoneme classifier module is responsible for removing content information of the speech signal to be recognized; The adversarial speaker classifier module is responsible for removing speaker information of the speech signal to be recognized; The classifier training module is responsible for training the emotion classifier, the adversarial phoneme classifier and the adversarial speaker classifier using the annotated speech signal; The speech output module is responsible for outputting the emotion category information corresponding to the speech signal to be recognized without the content information and the speaker information, and displaying it in a preset manner.

Citation Information

Patent Citations

  • Voice emotion recognition method based on adversarial semantic erasure

    CN111128240A

  • Multi-modal emotion recognition method and device based on pre-training model

    CN116778967A

  • Multi-language speech emotion recognition system based on domain adversarial learning

    CN117831566A

  • Method and system for fair speech emotion recognition

    US20240404511A1

  • Network model training method and device for speaker recognition and storage medium

    WO2023103375A1