Speech Emotion Recognition System and Method Based on Bi-MGAN and ResTCN-FDA Networks

The speech emotion recognition system using Bi-MGAN and ResTCN-FDA networks achieves bidirectional mapping and feature fusion of acoustic and phonetic features, solving the problem of insufficient feature evaluation in traditional models and improving the accuracy and applicability of speech emotion recognition.

CN115631769BActive Publication Date: 2026-01-30TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211094361.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-08
Publication Date
2026-01-30
Estimated Expiration
2042-09-08

AI Technical Summary

Technical Problem

In existing technologies, traditional acoustic and phonetic conversion models lack a feedback mechanism for evaluating mapped features, and classification and recognition algorithms ignore the dependencies between elements within features and the differences in emotional information contained in the channel dimensions of different types of features, resulting in insufficient accuracy in speech emotion recognition.

Method used

A speech emotion recognition system based on Bi-MGAN and ResTCN-FDA networks is adopted. The system realizes bidirectional mapping of acoustic and phonetic features through forward and backward generators, discriminators and loss function calculation modules. The system performs feature fusion and attention processing through ResTCN-FDA network and improves the accuracy of mapped features and emotion recognition rate by using a bounded mapping loss function and feature attention mechanism.

Benefits of technology

It improves the accuracy of speech emotion recognition, enhances the applicability of the model, maximizes the use of emotional information from different feature and dimension channels, and improves the emotion recognition rate, especially showing outstanding performance in the study of the correlation between multimodal signals and emotional states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115631769B_ABST
    Figure CN115631769B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of speech signal processing, specifically relating to a speech emotion recognition method and system based on Bi-MGAN and ResTCN-FDA networks. It includes a Bi-MGAN network, a ResTCN-FDA network, and a softmax module. The Bi-MGAN network comprises a forward generator, a backward generator, an articulatory discriminator, and an acoustic discriminator. The forward generator maps acoustic features to articulatory features, the backward generator maps articulatory features to acoustic features, the articulatory discriminator compares the actual articulatory features with the mapped articulatory features, and uses a loss function to retrieve the weight parameters of the forward generator. The acoustic discriminator compares the actual acoustic features with the mapped acoustic features and uses a loss function to retrieve the weight parameters of the backward generator. The ResTCN-FDA network includes a ResTCN network, an FA module, and a DA module. The softmax module calculates the corresponding emotion classification based on the output of the DA module. This invention can improve the speech emotion recognition rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of speech signal processing, specifically relating to a speech emotion recognition system and method based on Bi-MGAN and ResTCN-FDA networks. Background Technology

[0002] Emotion recognition (ER) is an important interface for human-computer interaction today. [1] The goal is to enable computers to understand, recognize, and generate emotional expressions. Integrating acoustic and phonetic features for emotion recognition is an important branch of emotion research, involving issues such as building emotion databases, data preprocessing, feature extraction, feature transformation algorithms, and classification algorithms. Among these, a multimodal database rich in emotional information, high-precision feature transformation algorithms, and effective classification algorithms are crucial for improving system performance.

[0003] Research in the field of emotion requires substantial data support. Researchers have constructed databases suitable for different research directions based on the diversity of information expressed by participants, such as CHEAVD. [2] NNIME [3] and IEMOCAP [4] In addition to the databases mentioned above, scholars have created a wide variety of other databases, but each database has its limitations. Only by choosing a database that aligns with one's research direction can one achieve twice the result with half the effort.

[0004] In their research on the human vocal mechanism, scholars have discovered a strong correlation between speech and the vocal organs; that is, some of the acoustic signals emitted by the human body are generated by the unique movement trajectories of each vocal organ. [5] To explore the relationship between acoustic signals and vocal organs, scholars have conducted extensive research, with the most in-depth studies focusing on forward mapping. [6] and reverse mapping [7] Forward mapping refers to converting the phonetic features of the articulatory organs (tongue, lips, and jaw, etc.) into acoustic features; the corresponding conversion of acoustic features into phonetic features is called inverse mapping. Currently, researchers have proposed using deep learning to explore forward and inverse mappings and applying them to various research fields. [8] The joint distribution relationship between phonetics and acoustic features was explored using a hidden markov model (HMM), and positive mapping was applied to speech synthesis. [9]Mel-frequency cepstral coefficients (MFCCs) were extracted, and the correlation between acoustic and phonetic features was explored using a Gaussian mixture model (GMM). Backpropagation was then applied to speaker recognition. Although these methods achieved good results in acoustic-phonetic conversion, they all relied solely on backpropagation to train the model, neglecting the evaluation and analysis of the conversion results.

[0005] In ER research, different emotion features and classification algorithms have been proposed to improve system performance. In feature research, Guo...

[10] Phase features were extracted, and their application in speech emotion recognition was explored. In research on recognition algorithms, Long Short-Term Memory (LSTM) units have been proposed.

[11] Convolutional neural networks (CNNs)

[12] Deep recurrent neural networks (RNNs)

[13] And Deep Neural Networks (DNN)

[14] Algorithms such as these establish a model of the connection between the speaker and emotion, but they do not analyze the dependencies between elements within features, nor do they consider the differences in emotional information contained in different dimensional channels of different feature classes. [9] .

[0006] References:

[0007] [1] Sun Ying, Hu Yanxiang, Zhang Xueying, et al. PAD prediction of emotion dimension for emotion speech recognition[J]. Journal of Zhejiang University (Engineering Science), 2019, 53(10):2041-2048.

[0008] [2]Li Y, Tao J, Chao L, et al. CHEAVD: a Chinese natural emotional audio–visual database [J]. Journal of Ambient Intelligence and Humanized Computing, 2017, 8(6): 913-924.

[0009] [3]Chou H C,Lin W C,Chang L C,et al.Nnime:The nthu-ntuachineseinteractive multimodal emotion corpus[C] / / 2017Seventh InternationalConference on Affective Computing and Intelligent Interaction.Texas:ACII,2017:292-298.

[0010] [4]Busso C,Bulut M,Lee C C,et al.IEMOCAP:Interactive emotional dyadicmotion capture database[J].Language resources and evaluation,2008,42(4):335-359.

[0011] [5]Qin C,Carreira M A.An empirical investigation of the nonuniquenessin the acoustic-to-articulatory mapping[C] / / Eighth Annual Conference of theInternational Speech Communication Association.Belgium:INTERSPEECH,2007:27-31.

[0012] [6]Ren G,Fu J,Shao G,et al.Articulatory-to-Acoustic Conversion ofMandarin Emotional Speech Based on PSO-LSSVM[J].Complexity,2021,29(3):696-706.

[0013] [7]Hogden J, Lofqvist A, Gracco V, et al. Accurate recovery ofarticulator positions from acoustics: New conclusions based on human data[J]. The Journal of the Acoustical Society of America, 1996, 100(3): 1819-1834.

[0014] [8]Ling ZH, Richmond K, Yamagishi J, et al. Integrating articulatory features into HMM-based parametric speech synthesis [J]. IEEE Transactions onAudio, Speech, and Language Processing, 2009, 17(6): 1171-1185.

[0015] [9]Li M,Kim J,Lammert A,et al.Speaker verification based on the fusion of speech acoustics and inverted articulatory signals[J].Computerspeech&language,2016,36:196-211.

[0016]

[11] Chen Q, Huang GA novel dual attention-based BLSTM with hybridfeatures in speech emotion recognition[J]. Engineering Applications of Artificial Intelligence, 2021,102:104277.

[0017]

[12] Zhang Jing, Zhang Xueying, Chen Guijun, et al. EEG emotion recognition combining 3D-CNN and frequency-space attention mechanism [J]. Journal of Xidian University, 2022, 49(03):191-198+205.

[0018]

[13] Kumaran U, Radha RS, Nagarajan SM, et al. Fusion of mel and gammatone frequency cepstral coefficients for speech emotion recognition using deep C-RNN[J]. International Journal of Speech Technology, 2021, 24(2): 303-314.

[0019]

[14] Lieskovská E,Jakubec M,Jarina R,et al.A review on speech emotionrecognition using deep learning and attention mechanism[J].Electronics,2021,10(10):1163.

[0020] In summary, in acoustic and phonetic feature conversion tasks, traditional regression models can only update weight parameters through backpropagation, lacking a feedback mechanism for evaluating mapped features. Furthermore, in ER tasks, current classification and recognition algorithms mainly model the speaker's features and emotional state, ignoring the dependencies between elements within features and the differences in emotional information contained in the channel dimensions of different feature classes. Moreover, in the field of emotion research, the impact of acoustic and phonetic feature conversion on emotion recognition has not yet been explored. Summary of the Invention

[0021] To address the limitations of existing traditional acoustic-phonetics conversion models and the resulting loss of accuracy, this invention overcomes these shortcomings by combining the advantages of acoustic signals and phonetics kinematic signals in emotion recognition. It provides a speech emotion recognition method and system based on Bi-MGAN and ResTCN-FDA networks to achieve accurate speech emotion recognition.

[0022] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: a speech emotion recognition system based on Bi-MGAN and ResTCN-FDA networks, including a Bi-MGAN network, a ResTCN-FDA network and a softmax module, wherein the Bi-MGAN network includes a forward generator, a backward generator, an articulatory discriminator, an acoustic discriminator and a loss function calculation module.

[0023] The forward generator is used to map the acoustic features of the phonetic features, the reverse generator is used to map the phonetic features of the acoustic features, the acoustic discriminator is used to compare the real acoustic features and the mapped acoustic features to obtain a first adversarial loss function, and the phonetic discriminator is used to compare the real phonetic features and the mapped phonetic features to obtain a second adversarial loss function.

[0024] The loss calculation module is used to calculate the positive overall loss function based on the original phonetic features, the mapped phonetic features, and the first adversarial loss function, and to calculate the negative overall loss function based on the original acoustic features, the mapped acoustic features, and the second adversarial overall loss function; the positive overall loss function is used to call back the weight parameters of the positive generator; the negative overall loss function is used to call back the weight parameters of the negative generator.

[0025] The ResTCN-FDA network includes a ResTCN network, an FA module, and a DA module. The acoustic features mapped by the forward generator and the corresponding phonetic features are fused together and then output to the ResTCN-FDA network. The phonetic features mapped by the reverse generator and the corresponding acoustic features are fused together and then output to the ResTCN-FDA network. The ResTCN network is used to process the input fused features. The FA module is used to perform feature attention processing on the processed fused features to obtain feature attention weight coefficients, and then applies the feature attention weights to the fused features output by the ResTCN network before outputting them. The DA module is used to perform dimensional attention processing on the features output by the FA module to obtain dimensional attention weight coefficients, and then applies the dimensional attention weight coefficients to the features output by the FA module before outputting them to the softmax module.

[0026] The Softmax module is used to calculate the corresponding sentiment classification based on the output of the DA module.

[0027] The expressions for the positive global loss function and the reverse binding mapping loss function are as follows:

[0028] L b (G X→Y ) = L a (G X→Y D Y )+λ1L c (x, G) Y→X )+λ2L g (G X→Y )+λ3L m (y, G) X→Y );

[0029] L b (GY→X ) = L a (G Y→X D X )+λ1L c (y, G) X→Y )+λ2L g (G Y→X )+λ3L m (x, G) Y→X );

[0030] Where λ1, λ2, and λ3 are the fit balance coefficients; L b (G X→Y ) and L b (G Y→X ) represent the forward global loss function and the reverse global loss function, respectively. a (G X→Y D Y ) represents the first adversarial loss function, L a (G Y→X D X ) represents the second adversarial loss function; L c (x, G) Y→X ) and L c (y, G) X→Y ) represent the cycle consistency loss functions for the forward generator and the backward generator, respectively; L g (G X→Y ) and L g (G Y→X Let L represent the loss functions of the forward generator and the backward generator, respectively. m (y, G) X→Y ) and L m (x, G) Y→X ) represent the forward binding mapping loss function and the reverse binding mapping loss function, respectively.

[0031] The loss functions for the forward generator and the backward generator are:

[0032] L g (G X→Y ) = E x~X [L bce (G X→Y (x))];

[0033] L g (G Y→X ) = E y~Y [L bce (G Y→X (y))];

[0034] Among them, E x~X [L bce (G X→Y[x)] represents the mean of the crossover loss function of the forward generator, E y~Y [L bce (G Y→X [y)] represents the mean of the cross-loss function of the reverse generator.

[0035] The forward binding mapping loss function and the reverse binding mapping loss function are as follows:

[0036] L m (y, G) X→Y ) = E y~Y [L1(y,G X→Y (x))];

[0037] L m (x, G) Y→X ) = E x~X [L1(x,G Y→X (y))];

[0038] Where, L1(x, G) Y→X (y) represents the original articulatory feature x and the mapped articulatory feature G. Y→X L1 regularization of (y), L1(y, G) X→Y (x) represents the original acoustic feature y and the mapped acoustic feature G. X→Y L1 regularization of (x), E x~X and E y~Y This indicates that the average value is being calculated.

[0039] The first adversarial overall loss function and the second adversarial loss function are also used to adjust the weight parameters of the acoustic discriminator and the phonetic discriminator, respectively.

[0040] The ResTCN network includes multiple ResTCN layers. Each ResTCN layer includes a dilated convolutional layer, a normalized ReLU activation layer, and a Dropout layer. In each ResTCN layer, the input features are processed by the dilated convolutional layer, the normalized ReLU activation layer, and the Dropout layer, and then concatenated with the input features of the dilated convolutional layer before being input into the next ResTCN layer.

[0041] The FA module includes a global max pooling layer, a global average pooling layer, a convolutional layer, and a sigmoid layer. The features output by the ResTCN network are transposed and then passed through the global max pooling layer and the global average pooling layer, respectively. The outputs of the two are concatenated and then passed through the convolutional layer and the sigmoid layer in sequence to obtain the feature attention weights.

[0042] The DA module includes a global average pooling layer, a fully connected layer, and a Sigmoid layer. The features output by the FA module are sequentially passed through the global average pooling layer, the fully connected layer, and the Sigmoid layer to obtain the dimensional attention weight coefficients.

[0043] The formula for applying feature attention weights to the fused features output by the ResTCN network is as follows:

[0044]

[0045] The formula for applying the dimension attention weight coefficient to the features output by the FA module is as follows:

[0046]

[0047] Where V represents the fusion feature of the ResTCN network output, F f (V) represents the feature attention weight coefficient, F d (V′) represents the dimension attention weight coefficient, V′ represents the output signal of the FA module, and U represents the output signal of the DA module. This is element-wise multiplication.

[0048] Furthermore, this invention also provides a speech emotion recognition method based on Bi-MGAN and ResTCN-FDA networks, implemented using the aforementioned system, comprising the following steps:

[0049] S1. Obtain acoustic features and corresponding phonetic features to form a dataset;

[0050] S2. Divide the dataset into a training set and a test set; input the training set data into the Bi-MGAN network to train the network;

[0051] S4. After training is complete, input the test set into the Bi-MGAN network to test the network.

[0052] S5. After the test is completed, use the system to perform emotion recognition on acoustic features or phonetic features.

[0053] Compared with the prior art, the present invention has the following advantages:

[0054] This invention proposes a speech emotion recognition method and system based on Bi-MGAN and ResTCN-FDA networks. The Bi-MGAN network is used to map acoustic and articulatory features, making the model more widely applicable. Furthermore, the newly proposed bound mapping loss function improves the accuracy of the mapped features transformed by the generator. Considering the different contributions of emotional information carried by different features and different dimensional channels, this invention uses the ResTCN-FDA network for emotion classification and recognition. The FDA mechanism assigns different weight parameters to different dimensional channels of different features, avoiding the waste and loss of emotional information caused by using the same weight parameters, thus maximizing the utilization of emotional information. Given that no publicly available emotional database based on Mandarin Chinese and containing synchronous acoustic and articulatory kinematic signals has been found, this study designed and recorded the STEM-E2VA database, and added glottal data and facial micro-expression data, laying the foundation for studying the correlation between multimodal signals and emotional states. The impact of real features, mapped features, and the feature set after their fusion on emotion recognition was investigated. Experimental results show that the recognition rate of mapped features is generally lower than that of real features, but the fusion of the two improves the emotion recognition rate. This indicates that the emotional information contained in the mapped features can supplement the real features in terms of emotion, and the supplementation effect is different for different emotions. Therefore, this invention can improve the emotion recognition rate and the accuracy of speech emotion recognition. Attached Figure Description

[0055] Figure 1 A schematic diagram of the structure of a speech emotion recognition system based on Bi-MGAN and ResTCN-FDA networks provided for an embodiment of the present invention;

[0056] Figure 2 This is a diagram showing the data flow and loss function structure of the Bi-MGAN network in an embodiment of the present invention;

[0057] Figure 3 Here are schematic diagrams (a) of the overall structure of ResTCN-FDA and (b) of the dilated convolution principle in an embodiment of the present invention;

[0058] Figure 4 This is a schematic diagram of the feature-dimension attention mechanism in an embodiment of the present invention. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0060] Example 1

[0061] like Figure 1 As shown, Embodiment 1 of the present invention provides a speech emotion recognition system based on Bi-MGAN and ResTCN-FDA networks, including a Bi-MGAN network, a ResTCN-FDA network, and a Softmax module. The Bi-MGAN network includes a forward generator G. X→Y Inverse generator G Y→X Phonetic discriminator D X Harmony and acoustic discriminator D Y .

[0062] Specifically, in this embodiment, the forward generator is used to map the acoustic features of the phonetic features, the backward generator is used to map the phonetic features of the acoustic features, the phonetic discriminator is used to compare the real phonetic features and the mapped phonetic features, and uses a loss function to call back the weight parameters of the forward generator; the acoustic discriminator is used to compare the real acoustic features and the mapped acoustic features, and uses a loss function to call back the weight parameters of the backward generator.

[0063] Forward Generator G X→Y The goal of this method is to map the vocalization motion features to corresponding acoustic features, thereby enabling the acoustic discriminator D to... Y The model cannot accurately distinguish between mapped and true acoustic features. To reduce model redundancy, dense layers were used to construct upsampling and downsampling modules. The upsampling module expands the input 28-dimensional phonetic features to 512 dimensions, while the downsampling module maps the high-dimensional phonetic features to 60-dimensional acoustic features; where X represents the phonetic feature dataset, Y represents the acoustic feature dataset, x represents the phonetic feature, and y represents the acoustic feature.

[0064] Inverse Generator G Y→X The goal of using MFCC acoustic feature mapping to map out corresponding articulation motion features is to enable the articulation discriminator D to... x Unable to correctly distinguish between mapped and actual phonetic features;

[0065] Phonetic discriminator D XThe system discriminates and calculates the error between the real and mapped articulation kinematic features, and utilizes the loss function callback G. Y→X The weight parameters are adjusted to improve the accuracy of the mapped features, achieving a supervisory and feedback effect on the mapped phonetic features. D X Essentially a binary classifier, its purpose is to correctly distinguish between mapped and actual phonetic features, which happens to correspond to G. Y→X Conversely, such a mapping model will be in G Y→X and D X The global optimal solution is found through alternating iterative optimization.

[0066] Acoustic discriminator D Y : Distinguish between real and mapped acoustic features, and use the loss function to callback G X→Y The weight parameters are adjusted to improve the accuracy of the mapped features, achieving a supervisory and feedback effect on the mapped acoustic features. The purpose of the acoustic discriminator DY is to correctly distinguish between mapped features and true features.

[0067] The optimization of the loss function of Bi-MGAN is mainly reflected in two aspects: generator loss function and binding mapping loss function. That is, Bi-MGAN considers four types of loss during training, namely generator loss, adversarial loss, cycle consistency loss and binding mapping loss.

[0068] Generator loss function: Added L g As the foundational mapping function of the generator, it enhances the generator's mapping capabilities. Forward generator G X→Y The loss function is shown in equation (1):

[0069] L g (G X→Y ) = E x~X [L bce (G X→Y (x))]; (1)

[0070] Inverse Generator G Y→X The loss function is:

[0071] L g (G Y→X ) = E y~Y [L bce (G Y→X (y))]; (2)

[0072] L bce E represents the Binary Cross Entropy loss function. x~X [L bce (G X→Y[x] represents the mean of the crossover loss function of the forward generator, and similarly, E y~Y [L bce (G Y→X [y] represents the mean of the crossover loss function of the reverse generator. Let G... X→Y (x) After inputting, view L bce The determination of its validity. If the determination result is true, then it means that G... X→Y (x) is already difficult to distinguish from the true feature y; if it is judged as false, an error will occur.

[0073] Binding mapping loss: To achieve both forward and reverse mappings between acoustics and phonetics, relying solely on the aforementioned loss function cannot guarantee the accuracy of the mapped features. Therefore, this study combines x with... and y The L1 norm is added to the network as an additional regularization term to constrain the range of generated mapping features by reducing the number of errors generated during model training. The forward and reverse binding mapping loss functions are shown in equations (3) and (4):

[0074] L m (y, G) X→Y ) = E y~Y [L1(y,G X→Y (x))]; (3)

[0075] L m (x, G) Y→X ) = E x~X [L1(x,G Y→X (y))]; (4)

[0076] In the above formula, L1(x, G) Y→X (y) represents the original articulatory feature x and the mapped articulatory feature G. Y→X L1 regularization of (y), L1(y, G) X→Y (x) represents the original acoustic feature y and the mapped acoustic feature G. X→Y L1 regularization of (x). m (y, G) X→Y ) and L m (x, G) Y→X ) represent the forward binding mapping loss function and the reverse binding mapping loss function, respectively, E x~X and E y~Y This indicates that the average value is calculated, and the binding mapping loss function is the mean of the corresponding regularized values.

[0077] In summary, the forward and backward global loss functions of Bi-MGAN are as follows:

[0078] L b (G X→Y ) = L a (G X→Y D Y )+λ1L c (x, G) Y→X )+λ2L g (G X→Y )+λ3L m (y, G) X→Y (5)

[0079] L b (G Y→X ) = L a (G Y→X D X )+λ1L c (y, G) X→Y )+λ2L g (G Y→X )+λ3L m (x, G) Y→X (6)

[0080] In the above formula, λ1, λ2, and λ3 are the fitness balance coefficients, which are optimized through experimentation; L b (G X→Y ) and L b (G Y→X ) represent the forward global loss function and the reverse global loss function, respectively. a (G X→Y D Y ) represents the first adversarial loss function, which is calculated by the acoustic discriminator. Specifically, it is obtained by calculating the adversarial loss between the mapped phonetic features obtained by the inverse generator and the corresponding original phonetic features. L a (G Y→X D X ) represents the second adversarial loss function, which is calculated by the phonetic discriminator, specifically by calculating the adversarial loss between the mapped acoustic features obtained from the forward generator and the original acoustic features; L c (x, G) Y→X ) and L c (y, G) X→Y ) represent the cycle consistency loss functions for the forward generator and the backward generator, respectively.

[0081] Adversarial loss: used to measure the discriminability between mapped features and true features. The first adversarial loss function is as shown in equation (7):

[0082] L a (G X→Y D Y ) = Ex~X [log(1-D Y (G X→Y (x)))]+E y~Y [logD Y (y)]; (7)

[0083] When y∈Y, the acoustic discriminator classifies y, setting its loss value to 1 if it is identified as real data; when x∈X, the acoustic discriminator classifies G... X→Y (x) is used for discrimination, and if it is determined to be mapped data, the loss value of the data is set to 0. Similarly, the second adversarial loss function can be obtained, which is to replace X and Y in formula (7).

[0084] Cycle consistency loss: This function transforms mapped features into cyclic features, aiming to make the cyclic features approximate the true features, i.e., G. Y→X (G X→Y (x))≈X or G X→Y (G Y→X (y))≈Y.

[0085] L c (x, G) Y→X ) = E x~X [L1(x,G Y→X (G X→Y (x)))]; (8)

[0086] L c (y, G) X→Y ) = E y~Y [L1(y,G X→Y (G Y→X (y)))]; (9)

[0087] In the formula, L1 represents the L1 norm, which in the ideal case is G. Y→X (G X→Y (x))=X or G X→Y (G Y→X (y))=Y.

[0088] In this embodiment, a new generator loss function is added as the basic mapping function of the generator to enhance its mapping capability. Furthermore, to complete the acoustic and phonetic feature conversion task, relying solely on formulaic adversarial loss and cycle consistency loss cannot guarantee the accuracy of the mapped features. Therefore, this embodiment introduces the original phonetic feature x and the mapped phonetic feature, as well as the original acoustic feature y and the norm of the mapped acoustic feature, as additional regularization terms into the network. By reducing the generation of mapping features with large errors during model training, the range of generated mapping features is constrained, thereby improving the accuracy of the mapped features.

[0089] Specifically, such as Figure 2 As shown, in this embodiment, for the forward generator: after obtaining the mapped phonetic features through the phonetic discriminator, comparing them with the real phonetic features yields an adversarial loss. Simultaneously, the backward generator remaps the mapped phonetic features to obtain cyclic acoustic features, which are then compared with the real acoustic features to obtain a cyclic consistency loss. Furthermore, comparing the real phonetic features with the mapped phonetic features yields a binding mapping loss. The generator loss is obtained by taking the mean of the generator cross-loss function. In this embodiment, the forward and backward generators adjust their parameters using the forward and backward global loss functions, respectively, while the phonetic discriminator and acoustic discriminator adjust their parameters using their corresponding adversarial loss functions. Specifically, the forward and backward global loss functions adjust the parameters of their respective generators after gradient operations, and the adversarial loss function also adjusts the corresponding generator parameters after gradient operations.

[0090] The acoustic features mapped by the forward generator and the corresponding phonetic features are fused together and then output to the ResTCN-FDA network. The phonetic features mapped by the reverse generator and the corresponding acoustic features are fused together and then output to the ResTCN-FDA network.

[0091] like Figure 3 As shown, in this embodiment, the ResTCN-FDA network includes a ResTCN network and an FDA module. The ResTCN network is used to process the input fused features, and the processed data is sent to the FDA module. Specifically, the ResTCN network includes multiple ResTCN layers. Each ResTCN layer includes a dilated convolutional layer, a normalization layer, a ReLU activation layer, and a Dropout layer. In each ResTCN layer, the input features are processed by the dilated convolutional layer, the normalization layer, the ReLU activation layer, and the Dropout layer, and then concatenated with the input features of the dilated convolutional layer before being input into the next ResTCN layer, and finally output to the FDA module.

[0092] like Figure 4As shown, in this embodiment, the FDA module includes a FA module and a DA module. The FA module performs feature attention processing on the processed fused features to obtain feature attention weight coefficients, and applies these feature attention weights to the fused features output by the ResTCN network before outputting the weights. The DA module performs dimensional attention processing on the features output by the FA module to obtain dimensional attention weight coefficients, and applies these dimensional attention weight coefficients to the features output by the FA module before outputting the weights to the Softmax module. The Softmax module calculates the corresponding sentiment classification based on the output of the ResTCN-FDA network.

[0093] The formula for applying feature attention weights to the fused features output by the ResTCN network is as follows:

[0094]

[0095] The formula for applying the dimension attention weight coefficient to the features output by the FA module is as follows:

[0096]

[0097] Where V represents the fusion feature of the ResTCN network output, F f (V) represents the feature attention weight coefficient, F d (V′) represents the dimension attention weight coefficient, V′ represents the output signal of the FA module, and U represents the output signal of the DA module. This is element-wise multiplication.

[0098] Specifically, in this embodiment, the FA module includes a global max pooling layer, a global average pooling layer, a convolutional layer, and a sigmoid layer. The features output by the ResTCN network are transposed and then passed through the global max pooling layer and the global average pooling layer, respectively. The outputs of the two are then concatenated and then passed through the convolutional layer and the sigmoid layer in sequence to obtain the feature attention weights.

[0099] In multi-class emotion recognition tasks, combining features from multiple classes yields better classification results than using a single feature. However, feature vectors from different classes exhibit varying responsiveness to emotion recognition. To better extract emotional information from multi-class features, this embodiment calculates the attention weights for each type of feature in the input. First, the transposed feature vectors are passed through a global max-pooling layer S. max(t) and global average pooling layer S ave(t) The outputs of the two are concatenated, and then passed through a one-dimensional convolutional layer and a sigmoid layer to finally obtain the feature attention weights F. f .

[0100] The DA module includes a global average pooling layer, a fully connected layer, and a Sigmoid layer. The features output by the FA module are sequentially passed through the global average pooling layer, the fully connected layer, and the Sigmoid layer to obtain the dimensional attention weight coefficients.

[0101] When using convolutional layers to process one-dimensional sequence feature tasks, the emotional information contained in the dimensional channels of the convolutional layer is uneven. To better utilize the emotional information contained in the dimensional channels, this embodiment uses a dimensional channel attention mechanism to process the features. First, global average pooling is applied to the output signal V′ of the FA module to obtain the feature mean F under each dimensional channel. ave,c :

[0102]

[0103] Among them, V c ∈R T×1 This represents the T×1 feature map parameters in the c-th dimension channel.

[0104] Then, a dimensional attention mechanism is implemented using a fully connected layer and a sigmoid function. Finally, the weight coefficients of the dimensional attention mechanism are applied to V′, thereby assigning different weight coefficients to each dimension. The relevant calculation formula is as follows:

[0105] F d (V′)=sigmoid(ω×F ave (13)

[0106] In the formula, ω represents the mapping of the fully connected layer.

[0107] Experimental example:

[0108] The 3D Electromagnetic Articulography (EMA) AG501 was used to acquire vocalization motion and acoustic data. During recording, the EMA-AG501 acquired the Cartesian coordinates of sensors fixed to the vocal organs at a sampling rate of 250 Hz via electromagnetic coupling as vocalization kinematic data, and simultaneously recorded acoustic data, forming parallel acoustic and vocalization kinematic data. The specific details of the acquired data are shown in the table below:

[0109] Table 1 Data collected by EMA-AG501

[0110]

[0111] Comparative experiment:

[0112] 1. Ablation Experiment: To verify the effectiveness of the generator loss function and the bound mapping loss function, an ablation experiment of the transformation network was conducted in this experimental example. The comparison models were set as GAN, CycleGAN, Bi-MGAN(G) with generator loss function added, Bi-MGAN(M) with bound mapping loss function added, and Bi-MGAN(G+M) containing both of the above loss functions.

[0113] 2. Comparison with mainstream state-of-the-art networks: To verify the effectiveness of the proposed transformation network, this experimental example compares Bi-MGAN with traditional DNN, BLSTM, and the latest Deep Recurrent Hybrid Density Network (DRMDN) and Least Squares Support Vector Machine (PSO-LSSVM) based on particle swarm optimization algorithm.

[0114] By collecting data and using the system of this invention for mapping, experiments have shown that Bi-MGAN significantly improves the mapping effect, especially in reverse mapping where RMSE and MAE reach 0.683mm and 0.501mm, respectively. Although the improvement effect of Bi-MGAN in forward mapping is not as significant as in reverse mapping, it still improves MAE and RMSE by 0.275mm and 0.212mm, respectively, compared to the CycleGAN network. This is because Bi-MGAN calculates the error between the real features and the mapped features in each batch during training, thereby continuously adjusting the generator's weight parameters to make the mapped features closer to the real features. Experiments demonstrate that the Bi-MGAN proposed in this invention can generate high-precision mapped features in acoustic and phonetic mapping through a bounded mapping loss function.

[0115] Furthermore, this invention considers the varying contributions of emotional information carried by different features and different dimensional channels. It uses the ResTCN-FDA network for emotion classification and recognition. The FDA mechanism assigns different weight parameters to different dimensional channels of different features, avoiding the waste and loss of emotional information caused by identical weight parameters, thus maximizing the utilization of emotional information. Experimental results show that the recognition rate of mapped features is generally lower than that of true features, but fusing the two improves the emotion recognition rate. This indicates that the emotional information contained in the mapped features can supplement the emotions of the true features, and the supplementary effect varies for different emotions. Future research will consider feature enhancement methods in the mapped features to improve the accuracy of emotion recognition, and plans to introduce multimodal fusion and contrastive learning techniques to enable computers to understand multimodal emotional information.

[0116] Example 2

[0117] Embodiment 2 of the present invention provides a speech emotion recognition method based on Bi-MGAN and ResTCN-FDA networks, implemented using the system described in Embodiment 1, and including the following steps:

[0118] S1. Obtain acoustic features and corresponding phonetic features to form a dataset;

[0119] S2. Divide the dataset into a training set and a test set; input the training set data into the Bi-MGAN network to train the network;

[0120] S4. After training is complete, input the test set into the Bi-MGAN network to test the network.

[0121] S5. After the test is completed, use the system to perform emotion recognition on acoustic features or phonetic features.

[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A speech emotion recognition system based on Bi-MGAN and ResTCN-FDA network, characterized in that, The Bi-MGAN network comprises a forward generator, a backward generator, a pronunciation-acoustic discriminator, an acoustic discriminator and a loss function calculation module. The forward generator is configured to map acoustic features to pronunciation-acoustic features, the backward generator is configured to map acoustic features to pronunciation-acoustic features, the acoustic discriminator is configured to compare the real acoustic features with the mapped acoustic features to obtain a first adversarial loss function, and the pronunciation-acoustic discriminator is configured to compare the real pronunciation-acoustic features with the mapped pronunciation-acoustic features to obtain a second adversarial loss function. The loss calculation module is configured to calculate a forward overall loss function according to the original pronunciation-acoustic features, the mapped pronunciation-acoustic features and the first adversarial loss function, and calculate a backward overall loss function according to the original acoustic features, the mapped acoustic features and the second adversarial loss function, the forward overall loss function is configured to adjust the weight parameters of the forward generator, and the backward overall loss function is configured to adjust the weight parameters of the backward generator. The ResTCN-FDA network comprises a ResTCN network, an FA module and a DA module, the acoustic features mapped by the forward generator and the corresponding pronunciation-acoustic features are fused and then output to the ResTCN-FDA network, the pronunciation-acoustic features mapped by the backward generator and the corresponding acoustic features are fused and then output to the ResTCN-FDA network, the ResTCN network is configured to process the input fused features, the FA module is configured to perform feature attention processing on the processed fused features to obtain feature attention weight coefficients, and apply the feature attention weight coefficients to the fused features output by the ResTCN network for weight processing and then output, and the DA module is configured to perform dimension attention processing on the features output by the FA module to obtain dimension attention weight coefficients, and apply the dimension attention weight coefficients to the features output by the FA module for weight processing and then output to the softmax module. The softmax module is configured to calculate corresponding sentiment classification according to the output of the DA module.

2. The speech emotion recognition system based on Bi-MGAN and ResTCN-FDA network according to claim 1, wherein, The expressions of the forward overall loss function and the backward overall loss function are as follows: L b (G X→Y )=L a (G X→Y ,D Y )+λ1L c (x,G Y→X )+λ2L g (G X→Y )+λ3L m (y,G X→Y ); L b (G Y→X ) = L a (G Y→X , D X ) + λ1L c (y, G X→Y ) + λ2L g (G Y→X ) + λ3L m (x, G Y→X ); wherein λ1, λ2 and λ3 are adaptive balance coefficients; L b (G X→Y ) and L b (G Y→X ) represent forward and backward overall loss functions, respectively, L a (G X→Y , D Y ) represents a first adversarial loss function, L a (G Y→X , D X ) represents a second adversarial loss function; L c (x, G Y→X ) and L c (y, G X→Y ) represent cycle-consistency loss functions of forward and backward generators, respectively; L g (G X→Y ) and L g (G Y→X ) represent loss functions of forward and backward generators, respectively, L m (y, G X→Y ) and L m (x, G Y→X ) represent forward and backward binding mapping loss functions, respectively.

3. The speech emotion recognition system based on Bi-MGAN and ResTCN-FDA network according to claim 2, characterized in that, The loss functions of the forward generator and the backward generator are as follows: L g (G X→Y ) = E x~X [L bce (G X→Y (x))] ; L g (G Y→X ) = E y~Y [L bce (G Y→X (y))] ; where E x~X [L bce (G X→Y (x))] represents the mean of the forward generator cross-entropy loss function, E y~Y [L bce (G Y→X (y))] represents the mean of the backward generator cross-entropy loss function.

4. The speech emotion recognition system based on Bi-MGAN and ResTCN-FDA network according to claim 2, characterized in that, The forward and backward constrained mapping loss functions are as follows: L m (y, G X→Y ) = E y~Y [L1(y, G X→Y (x))] ; L m (x, G Y→X ) = E x~X [L1(x, G Y→X (y))] where L1(x, G Y→X (y)) represents the L1 regularization of the original phonetic features x and the mapped phonetic features G Y→X (y), and L1(y, G X→Y (x)) represents the L1 regularization of the original acoustic features y and the mapped acoustic features G X→Y (x). E x~X and E y~Y represent the averaging.

5. The speech emotion recognition system based on Bi-MGAN and ResTCN-FDA network according to claim 1, characterized in that, The first and second adversarial loss functions are also configured to adjust the weight parameters of the acoustic discriminator and the pronunciation-acoustic discriminator, respectively.

6. The speech emotion recognition system based on Bi-MGAN and ResTCN-FDA network according to claim 1, characterized in that, The ResTCN network comprises a plurality of ResTCN layers, each ResTCN layer comprises a dilated convolution layer, a normalization layer, a ReLU activation layer and a Dropout layer, and in each ResTCN layer, the input features are processed by the dilated convolution layer, the normalization layer, the ReLU activation layer and the Dropout layer, and then spliced with the input features of the dilated convolution layer and input to the next ResTCN layer.

7. The speech emotion recognition system based on Bi-MGAN and ResTCN-FDA network according to claim 1, characterized in that, The FA module comprises a global maximum pooling layer, a global average pooling layer, a convolution layer, and a Sigmoid layer; the features output by the ResTCN network are transposed, then passed through the global maximum pooling layer and the global average pooling layer respectively, the outputs of the two layers are spliced, then passed through the convolution layer and the Sigmoid layer in sequence to obtain feature attention weights; The DA module comprises a global average pooling layer, a fully connected layer, and a Sigmoid layer; the features output by the FA module are sequentially passed through the global average pooling layer, the fully connected layer, and the Sigmoid layer to obtain dimension attention weight coefficients.

8. The speech emotion recognition system based on Bi-MGAN and ResTCN-FDA network according to claim 1, characterized in that, The formula for applying the feature attention weights to weight processing of the fusion features output by the ResTCN network is: The formula for applying the dimension attention weight coefficients to weight processing of the features output by the FA module is: wherein V represents the fusion feature output by the ResTCN network, F f (V) represents the feature attention weight coefficient, F d (V') represents the dimension attention weight coefficient, V' represents the output signal of the FA module, U represents the output signal of the DA module, is an element-wise multiplication.

9. A speech emotion recognition method based on Bi-MGAN and ResTCN-FDA network, characterized in that, The system implementation of claim 1 comprises the following steps: S1, obtaining acoustic features and corresponding articulatory features to form a data set; S2, dividing the data set into a training set and a test set; inputting the training set data into the Bi-MGAN network to train the network; S4, after the training is completed, inputting the test set into the Bi-MGAN network to test the network; S5, after the testing is completed, using the system after the testing to perform emotion recognition on the acoustic features or the articulatory features.