Speech emotion recognition method and related device, electronic device and storage medium
By adopting a joint model in speech emotion recognition, combining emotion recognition network and domain recognition network, and using the labelless data collaborative training model of the second sample speech, the problem of low accuracy of speech emotion recognition under scarce labeled data is solved, and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202111363984.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-17
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-11-17
AI Technical Summary
In the case where sample data with accurate emotional category annotations are scarce, the accuracy of speech emotion recognition is affected.
Using a joint model, combining emotion recognition network and domain recognition network, by jointly training the first sample speech and the second sample speech, the labeled data of the second sample speech is used to assist the labeled data of the first sample speech, and the network model is coordinated.
In the case of scarce labeled data, the accuracy of speech emotion recognition is improved, and the advantages of supervised training and unsupervised training are combined to enhance the performance of the network model.
Smart Images

Figure CN114333786B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech recognition technology, and in particular to a speech emotion recognition method and related devices, electronic equipment and storage media. Background Art
[0002] Speech emotion recognition refers to identifying the speaker's immediate emotional state (i.e., emotion category) from the speaker's speech, such as but not limited to: happiness, sadness, anger, surprise, disgust, fear, etc. In many industries or scenarios that require interaction with people to conduct business, such as human-computer interaction, medical care, education, and telephone services, the role of speech emotion recognition is becoming increasingly prominent.
[0003] With the rapid development of deep learning, the use of network models for speech emotion recognition has gradually become one of the mainstream technologies. Generally speaking, the accurate recognition of network models depends on the annotation of a large amount of sample data. However, due to the lack of speech emotion data in real environments, the subjectivity of emotions and many other factors, sample data with accurate emotion category annotations are relatively scarce, which seriously restricts the accuracy of network models and affects speech emotion recognition. In view of this, how to improve the accuracy of speech emotion recognition when sample data with accurate emotion category annotations is relatively scarce has become an urgent problem to be solved. Summary of the invention
[0004] The main technical problem solved by the present application is to provide a speech emotion recognition method and related devices, electronic devices and storage media, which can improve the accuracy of speech emotion recognition when sample data with accurate emotion category annotations are relatively scarce.
[0005] In order to solve the above technical problems, the first aspect of the present application provides a speech emotion recognition method, including: obtaining a speech to be recognized; using an emotion recognition network to recognize the speech to be recognized, and obtaining the emotion category of the speech to be recognized; wherein the emotion recognition network is included in a joint model, and the joint model also includes a domain recognition network, and the joint model is based on the emotion recognition network's emotion classification loss for a first sample speech belonging to a first data domain category and the domain recognition network's domain classification loss for the first sample speech and the second sample speech respectively. Joint training is obtained, and the second sample speech belongs to the second data domain category, and the first sample speech is annotated with a sample emotion category.
[0006] In order to solve the above technical problems, the second method of the present application provides a speech emotion recognition device, including: a speech acquisition module and an emotion recognition module, the speech acquisition module is used to acquire the speech to be recognized; the emotion recognition module is used to use the emotion recognition network to recognize the speech to be recognized, and obtain the emotion category of the speech to be recognized; wherein the emotion recognition network is included in the joint model, and the joint model also includes a domain recognition network, and the joint model is based on the emotion classification loss of the emotion recognition network for a first sample speech belonging to a first data domain category and the domain classification loss of the domain recognition network for the first sample speech and the second sample speech respectively. Joint training is obtained, and the second sample speech belongs to the second data domain category, and the first sample speech is annotated with a sample emotion category.
[0007] In order to solve the above technical problems, the third aspect of the present application provides an electronic device, including a memory and a processor coupled to each other, the memory storing program instructions, and the processor being used to execute the program instructions to implement the speech emotion recognition method in the above first aspect.
[0008] In order to solve the above technical problems, the fourth aspect of the present application provides a computer-readable storage medium, which stores program instructions that can be executed by a processor, and the program instructions are used to implement the speech emotion recognition method in the above first aspect.
[0009] The above scheme obtains the speech to be recognized, and uses the emotion recognition network to recognize the speech to be recognized, and obtains the emotion category of the speech to be recognized, and the emotion recognition network is included in the joint model, and the joint model also includes a domain recognition network. The joint model is obtained by jointly training the emotion classification loss of the first sample speech belonging to the first data domain category based on the emotion recognition network and the domain classification loss of the first sample speech and the second sample speech respectively by the domain recognition network, and the second sample speech belongs to the second data domain category, and the first sample speech is annotated with a sample emotion category. On the one hand, the emotion recognition network is included in the joint model, and the emotion of the first sample speech annotated with the sample emotion category is obtained by jointly training the emotion recognition network. Classification loss, the first sample speech can be used to supervise the training of the network model. On the other hand, an unlabeled and sufficient second sample speech is introduced, and the joint model also includes a domain recognition network, and the domain classification loss of the first sample speech and the second sample speech is respectively performed through the domain recognition network, which is conducive to enabling the network model to process the speech data of the two data domains indiscriminately, thereby realizing the use of the second sample speech to assist the first sample speech to unsupervisedly train the network model, and then combining the advantages of supervised training and unsupervised training to collaboratively train the network model. Therefore, when the sample data with accurate emotion category annotations is relatively scarce, the accuracy of speech emotion recognition can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1This is a flow chart of an embodiment of a speech emotion recognition method of the present application;
[0011] Figure 2 It is a schematic diagram of the framework of an embodiment of a joint model;
[0012] Figure 3 It is a schematic diagram of the framework of an embodiment of a speech feature extraction subnetwork;
[0013] Figure 4 It is a schematic diagram of the framework of an embodiment of an image feature extraction network;
[0014] Figure 5 It is a schematic diagram of the framework of an embodiment of a speech emotion recognition device of the present application;
[0015] Figure 6 It is a schematic diagram of the framework of an embodiment of the electronic device of the present application;
[0016] Figure 7 It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION
[0017] The scheme of the embodiment of the present application is described in detail below in conjunction with the drawings of the specification.
[0018] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.
[0019] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. In addition, "many" in this article means two or more than two.
[0020] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of the speech emotion recognition method of the present application. Specifically, it may include the following steps:
[0021] Step S11: Acquire the speech to be recognized.
[0022] In an implementation scenario, as described below, the emotion category of the speech to be recognized is identified by using an emotion recognition network, and the emotion recognition network is included in a joint model, and the joint model is obtained by joint training based on a first sample speech belonging to a first data domain category and a second sample speech belonging to a second data domain category. On this basis, the speech to be recognized can belong to either the first data domain category or the second data domain category, which is not limited here.
[0023] In a specific implementation scenario, the data domain category can be defined by the data source. For example, voice data from a mobile phone call and voice data from an instant messaging software can be considered to belong to different data domain categories, or voice data from a video call and voice data from a live recording can be considered to belong to different data domain categories. Examples will not be given one by one here.
[0024] In a specific implementation scenario, the specific data domain categories of the first data domain category and the second data domain category can be set according to the actual application situation. For example, the voice data from a video call can be more accurately labeled with emotional categories because it can refer to the video screen, and the voice data from the Internet is generally sufficient and can be obtained more conveniently. In this case, the voice data from the video call can be used as the first sample voice, and the voice data from the Internet can be used as the second sample voice. At this time, the first data domain category can be regarded as a video call, and the second data domain category can be regarded as the Internet. It should be noted that the above-mentioned first sample voice, second sample voice, first data domain category and second data domain category are only one possible situation in the actual application process, and the actual application process is not limited to this example.
[0025] Step S12: using the emotion recognition network to recognize the speech to be recognized, and obtaining the emotion category of the speech to be recognized.
[0026] In the disclosed embodiment, the emotion recognition network is included in the joint model, and the joint model also includes a domain recognition network. The joint model is obtained by joint training based on the emotion classification loss of the emotion recognition network for the first sample speech belonging to the first data domain category and the domain classification loss of the domain recognition network for the first sample speech and the second sample speech, respectively, and the second sample speech belongs to the second data domain category, and the first sample speech is annotated with a sample emotion category. On this basis, after the training of the joint model converges, the emotion recognition network in the joint model can be obtained, and the emotion recognition network can be used to perform speech emotion recognition.
[0027] In one implementation scenario, the sample emotion categories marked by the first sample speech may include but are not limited to happiness, sadness, anger, surprise, disgust, fear, and are not limited here. In addition, as mentioned above, the second sample speech may not be marked with labels related to emotion categories.
[0028] In an implementation scenario, the emotion classification loss can represent the accuracy of the emotion recognition network in identifying the emotion of speech data. For example, the larger the emotion classification loss, the lower the accuracy of the emotion recognition network in identifying the emotion. Conversely, the smaller the emotion classification loss, the higher the accuracy of the emotion recognition network in identifying the emotion. In the joint training process, the emotion classification loss can be minimized to improve the accuracy of the emotion recognition network in identifying the emotion of speech data.
[0029] In an implementation scenario, the domain classification loss can characterize the accuracy of the domain recognition network in domain recognition of speech data. For example, the larger the domain classification loss, the lower the accuracy of the domain recognition network in domain recognition. Conversely, the smaller the domain classification loss, the higher the accuracy of the domain recognition network in domain recognition. In the joint training process, the domain classification loss can be maximized so that the network model can process the speech data of the two data domains indiscriminately. In this way, in the joint process, the scarcity of the labeled first sample speech can be used to compensate for the negative impact of the model training on the model training through the unlabeled and sufficient second sample speech, and then the advantages of supervised training and unsupervised training can be combined to synergistically train the network model.
[0030] In an implementation scenario, please refer to Figure 2 , Figure 2 This is a schematic diagram of the framework of an embodiment of the joint model of the present application. Figure 2 As shown, the emotion recognition network and the domain recognition network share the emotion feature extraction subnetwork. On this basis, the emotion feature extraction subnetwork can be used to extract emotion features from the first sample speech and the second sample speech respectively, to obtain the first emotion feature of the first sample speech and the second emotion feature of the second sample speech, and then to predict the emotion category based on the first emotion feature to obtain the predicted emotion category of the first sample speech, and to predict the domain category based on the first emotion feature and the second emotion feature respectively to obtain the first predicted domain category to which the first sample speech belongs and the second predicted domain category to which the second sample speech belongs, and then to obtain the first loss, the second loss and the third loss based on the difference between the first predicted domain category and the first data domain category, the difference between the second predicted domain category and the second data domain category, and the difference between the predicted emotion category and the sample emotion category, so as to obtain the total loss based on the first loss, the second loss and the third loss, and adjust the network parameters of the joint model based on the total loss. In the above method, the model performance of the network model is improved by simultaneously supervising the emotion classification loss of the first sample speech and the domain classification loss of the first sample speech and the second sample speech.
[0031] In a specific implementation scenario, such as Figure 2 As shown, the emotion feature extraction subnetwork may include a speech feature extraction subnetwork and a speech emotion coding subnetwork, and the speech feature extraction subnetwork is used to perform speech feature extraction, and the speech emotion coding subnetwork is used to perform emotion feature coding based on speech features to obtain emotion features. Specifically, taking the first sample speech as an example, the acoustic feature extraction may be performed on the first sample speech to obtain first sample acoustic features of several first sample audio frames. Exemplarily, the acoustic features may include but are not limited to SIFT (Scale Invariant Feature Transform) features, etc., which are not limited here. On this basis, the speech feature extraction subnetwork may be used to perform speech feature extraction on the first sample acoustic features of several first sample audio frames to obtain first sample speech features (denoted as ), and then use the speech emotion coding sub-network to encode the first sample speech features of the first sample audio frames to obtain the first emotion feature (denoted as ). In addition, the speech feature extraction subnetwork may include but is not limited to a time-delay neural network (TDNN), a long short-term memory network, etc., and the network structure of the speech feature extraction subnetwork is not limited here. Figure 3 , Figure 3 It is a framework diagram of an embodiment of a speech feature extraction subnetwork. Taking the speech feature extraction subnetwork including TDNN as an example, the speech feature extraction subnetwork may specifically include N layers (e.g., 5 layers, 6 layers, etc.) of sequentially connected TDNNs, which input acoustic features such as SIFT and output speech features of audio frames. In addition, the speech emotion coding subnetwork may include but is not limited to a statistical pooling layer and a fully connected layer. The statistical pooling layer is used to calculate the first-order and second-order statistics, i.e., the mean and standard deviation, of the first sample speech features of several first sample audio frames in the time dimension, and then obtain the first emotion feature through the fully connected layer. It should be noted that the process of extracting the second emotional features of the second sample speech can refer to the specific method of extracting the first emotional features of the first sample speech mentioned above, and the processing process of the statistical pooling layer can refer to its technical details, which will not be repeated here.
[0032] In a specific implementation scenario, please continue to refer to Figure 2, the emotion recognition network may also include an emotion classification subnetwork, and the emotion classification subnetwork is used to perform emotion category prediction. The emotion classification subnetwork may include but is not limited to a fully connected layer, etc., which is not limited here. On this basis, the first emotion feature can be input into the emotion classification subnetwork to obtain the predicted emotion category of the first sample voice. Specifically, the emotion classification subnetwork can output the predicted probability values that the first sample voice belongs to several preset emotion categories (such as happiness, sadness, anger, surprise, disgust, fear, etc.), and the preset emotion category corresponding to the maximum predicted probability value can be used as the predicted emotion category of the first sample voice. On this basis, based on the sample emotion category of the first sample voice, the predicted probability values that the first sample voice belongs to several preset emotion categories can be processed by using a loss function such as cross entropy to obtain the above-mentioned third loss. The specific calculation process of the third loss can refer to the technical details of loss functions such as cross entropy, which will not be repeated here. In the above manner, the emotion recognition network further includes an emotion classification subnetwork, and the emotion classification subnetwork is used to perform emotion category prediction, which is conducive to improving the efficiency of emotion prediction.
[0033] In a specific implementation scenario, please continue to refer to Figure 2 , the domain recognition network may also include a domain classification subnetwork, and the domain classification subnetwork is used to perform domain category prediction. The domain category prediction subnetwork may include but is not limited to a fully connected layer, etc., which is not limited here. On this basis, the first emotional feature can be input into the domain classification subnetwork to obtain the first data domain category of the first sample voice. Specifically, the domain classification subnetwork can output the predicted probability values that the first sample voice belongs to the first data domain category and the second data domain category respectively, and the data domain category corresponding to the maximum predicted probability value can be used as the first predicted domain category of the first sample voice. On this basis, based on the first data domain category of the first sample voice, the predicted probability values that the first sample voice belongs to the first data domain category and the second data domain category respectively can be processed using a loss function such as a binary cross entropy to obtain the above-mentioned first loss. The specific calculation process of the first loss can refer to the technical details of loss functions such as cross entropy, which will not be repeated here. In addition, the specific calculation process of the second loss is similar to the first loss, which will not be repeated here. In the above manner, the domain recognition network also includes a domain classification subnetwork, and the domain classification subnetwork is used to perform domain category prediction, which is conducive to improving the efficiency of domain category prediction.
[0034] In a specific implementation scenario, please continue to refer to Figure 2, the domain recognition network may include a first recognition network for domain recognition of a first sample speech and a second recognition network for domain recognition of a second sample speech, the first recognition network and the second recognition network have the same network structure and network parameters, that is, the network parameters and network structure of the first recognition network and the second recognition network remain consistent before and after adjustment. In this case, the first sample speech is input into the first recognition network to obtain a first predicted domain category, and the second sample speech is input into the second recognition network to obtain a second predicted domain category. For the specific recognition process, please refer to the aforementioned related description, which will not be repeated here. In addition, if Figure 2 As shown, the first recognition network and the emotion recognition network can share the emotion feature extraction subnetwork, that is, the first recognition network and the emotion recognition network can both include the aforementioned speech feature extraction subnetwork and the aforementioned speech emotion encoding subnetwork. In the above manner, by sharing the same network structure and even the same network parameters in the joint model, the data processing efficiency of the joint model can be improved.
[0035] In a specific implementation scenario, the first loss and the second loss are negatively correlated with the total loss, and the third loss is positively correlated with the total loss. In other words, the larger the first loss and the second loss, the smaller the total loss, and the smaller the third loss, the smaller the total loss. Therefore, by minimizing the total loss as the model optimization goal, as the model recognition accuracy improves, the model can be unable to distinguish the first emotional feature of the first sample speech and the second emotional feature of the second sample speech, that is, the feature data extracted from the speech data in different data domains tend to the same data distribution, thereby enabling the model to accurately perform emotion recognition in different data domain categories.
[0036] In a specific implementation scenario, the specific method of adjusting parameters can refer to the technical details of optimization methods such as gradient descent, which will not be repeated here.
[0037] In one implementation scenario, in order to further improve the model accuracy, the first sample speech can be separated from the sample video, and the sample video can also be separated into a sample face image. For example, the sample video can be separated into a sample image sequence, and the sample image sequence can include several sample face images. On this basis, the sample face images can be used to assist model training to enhance the expression of emotional features by using face image information. In addition, please refer to Figure 2, the joint model can further include an image feature extraction network, based on which the image feature extraction network can be used to extract image features from sample face images to obtain sample image features of the sample face images, and the first emotion features can be fused with the sample image features to obtain sample fusion features, so that emotion category prediction can be performed based on the sample fusion features to obtain predicted emotion categories. For example, the aforementioned emotion classification subnetwork can be used to predict emotion categories for sample fusion features to obtain predicted emotion categories. For the emotion classification subnetwork, please refer to the aforementioned related descriptions, which will not be repeated here. In the above method, the first sample voice is set to be separated from the sample video, and the sample face image can also be separated from the sample video. On this basis, the sample image features extracted from the sample face image are fused with the first emotion features extracted from the first sample voice to obtain sample fusion features, so as to perform emotion prediction based on the sample fusion features. The expression of speech emotion features can be enhanced through face image information, which is beneficial to improving the recognition accuracy of the model.
[0038] In a specific implementation scenario, in order to improve the auxiliary training effect of sample face images, the sample video is video data that has been verified for lip shape and voice consistency, that is, the audio data and image data at any time in the sample video are consistent in the pronunciation dimension. For example, at time t, the image data shows that the lips are opened with a large amplitude during pronunciation, and the audio data is expressed as the sound "a"; or, at time t+1, the image data shows that the lips are opened in a circular shape with a small amplitude during pronunciation, and the audio data is expressed as the sound "o", and so on, and no examples are given here one by one.
[0039] In a specific implementation scenario, the image feature extraction network may include but is not limited to a convolutional neural network, etc., and the network structure of the image feature extraction network is not limited here. Figure 4 , Figure 4 Schematic diagram of the framework of an embodiment of an image feature extraction network. Figure 4 As shown in the figure, the image feature extraction network contains 5 layers in total. The first and second layers contain convolutional layers (conv) and pooling layers (pool), the third and fourth layers contain convolutional layers, the fifth layer contains convolutional layers and pooling layers, the sixth layer contains convolutional layers, and the seventh layer contains a fully connected layer (fc). It should be noted that Figure 4 The convolutional layers shown are all three-dimensional convolutional layers, that is, the input image feature extraction network can be N consecutive frames (e.g., 5 frames) of sample face images randomly selected from the sample image sequence. In addition, Figure 4 In the figure, T represents the number of frames, W represents the image width, H represents the image height, and the last column represents the number of convolution kernels. Figure 4 What is shown is only a possible implementation method of the image feature extraction network in practical applications, and does not limit the specific structure of the image feature extraction network.
[0040] In a specific implementation scenario, the above features can be expressed in vector form. On this basis, the first emotion feature can be added to the sample image feature to obtain the sample fusion feature. Of course, the sample fusion feature can also be obtained by weighting the first emotion feature and the sample image feature, which is not limited here. For the convenience of description, the sample image feature can be recorded as The first emotional feature With sample image features The sample fusion feature f can be fused emo .
[0041] In an implementation scenario, there are differences in gender, age and other factors unrelated to emotions between different speakers, which may lead to huge differences in the distribution of emotional features between different speakers. In order to minimize the interference of speaker information on emotion recognition, please continue to refer to the figure. The joint model can further include a speaker recognition network, and the speaker recognition network can include a speaker feature extraction subnet. Based on this, the speaker feature extraction subnet can be used to extract speaker features from the first sample speech to obtain the speaker features of the first sample speech, and based on the mutual information between the first emotional feature and the speaker feature, the fourth loss is obtained, and based on the first loss, the second loss, the third loss and the fourth loss, the total loss is obtained, and the total loss is positively correlated with the fourth loss, that is, the larger the fourth loss, the larger the total loss, and vice versa, the smaller the fourth loss, the smaller the total loss. It should be noted that the fourth loss is also positively correlated with the amount of information of the mutual information, that is, the larger the amount of information, the larger the fourth loss, and vice versa, the smaller the amount of information, the smaller the fourth loss. The calculation process of the amount of information can refer to the relevant technical details of the mutual information, which will not be repeated here. In the above method, a speaker recognition network is set in the joint model, and the speaker recognition network includes a speaker feature extraction subnetwork, and the speaker feature extraction subnetwork is used to extract the speaker feature of the first sample speech to obtain the speaker feature of the first sample speech, thereby obtaining the fourth loss based on the mutual information between the first emotional feature and the speaker feature, and then by minimizing the total loss as the model optimization goal, the mutual information between the first emotional feature and the speaker feature is made as small as possible, that is, the correlation between the first emotional feature and the speaker feature is reduced as much as possible, which can reduce the interference of speaker information on emotion recognition as much as possible, which is conducive to improving the accuracy of emotion recognition.
[0042] In a specific implementation scenario, please continue to refer to Figure 2, the speaker feature extraction subnetwork can share the speech feature extraction subnetwork with the emotion feature extraction subnetwork. As mentioned above, the speech feature extraction subnetwork is used to perform speech feature extraction. The network structure and specific principles of the speech feature extraction subnetwork can refer to the aforementioned related descriptions and will not be repeated here. The speaker feature extraction subnetwork also includes a speaker encoding subnetwork, which is used to perform speaker encoding based on speech features to obtain speaker features. Specifically, as mentioned above, the first sample speech features (denoted as ), then the speaker encoding subnetwork can be used to perform speaker encoding on the first sample speech features of the first sample audio frames to obtain the speaker features of the first sample speech (denoted as f spk ). In addition, similar to the speech emotion coding subnetwork, the speaker coding subnetwork may include but is not limited to a statistical pooling layer and a fully connected layer. The statistical pooling layer is used to calculate the first-order and second-order statistics, i.e., the mean and standard deviation, of the first sample speech features of a plurality of first sample audio frames in the time dimension, and then process them through the fully connected layer to obtain the speaker feature f spk . It should be noted that the processing process of the statistical pooling layer can refer to its technical details, which will not be repeated here. In the above manner, the speaker feature extraction subnetwork and the emotion feature extraction subnetwork share the speech feature extraction subnetwork, the speech feature extraction subnetwork is used to perform speech feature extraction, the emotion feature extraction subnetwork also includes a speech emotion encoding subnetwork, the speech emotion encoding subnetwork is used to perform emotion feature encoding based on speech features to obtain emotion features, and the speaker feature extraction subnetwork also includes a speaker encoding subnetwork for performing speaker encoding based on speech features to obtain speaker features, which is conducive to reducing the complexity of the joint model as much as possible and improving model efficiency.
[0043] In a specific implementation scenario, please continue to refer to Figure 2, the speaker recognition network may also include a speaker classification subnetwork, and the first sample speech may also be labeled with a sample speaker (e.g., speaker A). It should be noted that the speaker classification subnetwork may include but is not limited to a fully connected layer, which is not limited here. On this basis, the speaker classification subnetwork may be used to perform speaker prediction on the speaker features to obtain the predicted speaker of the first sample speech, and based on the difference between the sample speaker and the predicted speaker, the fifth loss may be obtained, so that the total loss may be obtained based on the first loss, the second loss, the third loss, the fourth loss and the fifth loss, and the fifth loss is positively correlated with the total loss. Specifically, the speaker classification subnetwork may be used to perform speaker prediction on the speaker features to obtain the predicted probability values that the first sample speech belongs to several preset speakers (e.g., speaker A, speaker B, speaker C, etc.), and based on the sample speakers labeled by the first sample speech, the predicted probability values that the first sample speech belongs to several preset speakers are processed using loss functions such as cross entropy to obtain the fifth loss. It should be noted that the specific calculation process of the fifth loss can refer to the technical details of loss functions such as cross entropy, which will not be repeated here. The above method further predicts the predicted speaker of the first sample speech, and obtains the fifth loss based on the difference between the predicted speaker and the sample speaker marked by the first sample speech, so that the fifth loss can be further combined in the parameter adjustment process, and then the accuracy of speech features can be improved through the supervision of the total loss, which is conducive to improving the accuracy of emotion recognition.
[0044] In an implementation scenario, as described above and Figure 2 As shown, the emotion recognition network may include an emotion feature extraction subnetwork and an emotion classification subnetwork, and the emotion feature extraction subnetwork may include a speech feature extraction subnetwork and a speech emotion coding subnetwork. On this basis, the acoustic feature extraction may be performed on the speech to be recognized first to obtain the acoustic features of several audio frames (such as SIFT features, etc.). On this basis, the speech feature extraction subnetwork may be used to perform speech feature extraction on the acoustic features of several audio frames to obtain the speech features of several audio frames, and the speech emotion coding subnetwork may be used to perform emotion feature encoding on the speech features of several audio frames to obtain the emotion features of the speech to be recognized, and then the emotion classification subnetwork may be used to predict the emotion category of the emotion features to obtain the emotion category of the speech to be recognized. It should be noted that the specific methods of extracting acoustic features, extracting speech features, encoding emotion features, and predicting emotion categories can refer to the relevant descriptions in the aforementioned public embodiments, and will not be repeated here.
[0045] The above scheme obtains the speech to be recognized, and uses the emotion recognition network to recognize the speech to be recognized, and obtains the emotion category of the speech to be recognized, and the emotion recognition network is included in the joint model, and the joint model also includes a domain recognition network. The joint model is obtained by jointly training the emotion classification loss of the first sample speech belonging to the first data domain category based on the emotion recognition network and the domain classification loss of the first sample speech and the second sample speech respectively by the domain recognition network, and the second sample speech belongs to the second data domain category, and the first sample speech is annotated with a sample emotion category. On the one hand, the emotion recognition network is included in the joint model, and the emotion of the first sample speech annotated with the sample emotion category is obtained by jointly training the emotion recognition network. Classification loss, the first sample speech can be used to supervise the training of the network model. On the other hand, an unlabeled and sufficient second sample speech is introduced, and the joint model also includes a domain recognition network, and the domain classification loss of the first sample speech and the second sample speech is respectively performed through the domain recognition network, which is conducive to enabling the network model to process the speech data of the two data domains indiscriminately, thereby realizing the use of the second sample speech to assist the first sample speech to unsupervisedly train the network model, and then combining the advantages of supervised training and unsupervised training to collaboratively train the network model. Therefore, when the sample data with accurate emotion category annotations is relatively scarce, the accuracy of speech emotion recognition can be improved.
[0046] See also Figure 5 , Figure 5 It is a schematic diagram of a framework of an embodiment of a speech emotion recognition device 50 of the present application. The speech emotion recognition device 50 includes: a speech acquisition module 51 and an emotion recognition module 52, wherein the speech acquisition module 51 is used to acquire the speech to be recognized; the emotion recognition module 52 is used to recognize the speech to be recognized by using an emotion recognition network, and obtain the emotion category of the speech to be recognized; wherein the emotion recognition network is included in a joint model, and the joint model also includes a domain recognition network, and the joint model is obtained by joint training based on the emotion classification loss of the emotion recognition network for the first sample speech belonging to the first data domain category and the domain classification loss of the domain recognition network for the first sample speech and the second sample speech, respectively, and the second sample speech belongs to the second data domain category, and the first sample speech is annotated with a sample emotion category.
[0047] In the above scheme, on the one hand, the emotion recognition network is included in the joint model, and the emotion recognition network is used to perform emotion classification loss on the first sample speech labeled with the sample emotion category, and the first sample speech can be used to supervise the training of the network model. On the other hand, an unlabeled and sufficient second sample speech is introduced, and the joint model also includes a domain recognition network, and the domain classification loss of the first sample speech and the second sample speech is respectively performed through the domain recognition network, which is conducive to enabling the network model to process the speech data of the two data domains indiscriminately, thereby realizing the use of the second sample speech to assist the first sample speech to unsupervisedly train the network model, and then combining the advantages of supervised training and unsupervised training to collaboratively train the network model. Therefore, when the sample data with accurate emotion category annotations is relatively scarce, the accuracy of speech emotion recognition can be improved.
[0048] In some disclosed embodiments, the emotion recognition network and the domain recognition network share an emotion feature extraction subnetwork, and the speech emotion recognition device 50 includes an emotion feature extraction module, which is used to use the emotion feature extraction subnetwork to extract emotion features from the first sample speech and the second sample speech respectively, and obtain a first emotion feature of the first sample speech and a second emotion feature of the second sample speech; the speech emotion recognition device 50 includes an emotion category prediction module, which is used to predict the emotion category based on the first emotion feature and obtain a predicted emotion category of the first sample speech; the speech emotion recognition device 50 includes a domain category prediction module, which is used to predict the domain category based on the first emotion feature and the second emotion feature respectively, and obtain a first predicted domain category to which the first sample speech belongs and a second predicted domain category to which the second sample speech belongs; the speech emotion recognition device 50 includes a loss calculation module, which is used to obtain a first loss, a second loss and a third loss based on the difference between the first predicted domain category and the first data domain category, the difference between the second predicted domain category and the second data domain category, and the difference between the predicted emotion category and the sample emotion category; the speech emotion recognition device 50 includes a parameter adjustment module, which is used to obtain a total loss based on the first loss, the second loss and the third loss, and adjust the network parameters of the joint model based on the total loss.
[0049] Therefore, by simultaneously supervising the sentiment classification loss of the first sample speech and the domain classification loss of the first sample speech and the second sample speech, the model performance of the network model is improved.
[0050] In some disclosed embodiments, the first loss and the second loss are respectively negatively correlated with the total loss, and the third loss is positively correlated with the total loss.
[0051] Therefore, by taking minimizing the total loss as the model optimization goal, as the model recognition accuracy improves, the model can be made unable to distinguish between the first emotional features of the first sample speech and the second emotional features of the second sample speech, so that the feature data extracted from speech data in different data domains tend to the same data distribution, thereby enabling the model to accurately perform emotion recognition in different data domain categories.
[0052] In some disclosed embodiments, the emotion recognition network further includes an emotion classification subnetwork, which is used to perform emotion category prediction; and / or, the domain recognition network further includes a domain classification subnetwork, which is used to perform domain category prediction.
[0053] Therefore, the emotion recognition network further includes an emotion classification subnetwork, and the emotion classification subnetwork is used to perform emotion category prediction, which is conducive to improving the emotion prediction efficiency, while the domain recognition network also includes a domain classification subnetwork, and the domain classification subnetwork is used to perform domain category prediction, which is conducive to improving the domain category prediction efficiency.
[0054] In some disclosed embodiments, the first sample speech is separated from the sample video, and the sample video also separates a sample face image, the joint model also includes an image feature extraction network, the speech emotion recognition device 50 includes an image feature extraction module, which is used to use the image feature extraction network to extract image features of the sample face image to obtain sample image features of the sample face image; the speech emotion recognition device 50 includes a feature fusion module, which is used to fuse the first emotion feature with the sample image feature to obtain a sample fusion feature; the emotion category prediction module is specifically used to predict the emotion category based on the sample fusion feature to obtain a predicted emotion category.
[0055] Therefore, the first sample speech is set to be separated from the sample video, and the sample face image can also be separated from the sample video. On this basis, the sample image features extracted from the sample face image are fused with the first emotion features extracted from the first sample speech to obtain the sample fusion features. Emotion prediction is performed based on the sample fusion features, and the expression of speech emotion features can be enhanced through face image information, which is conducive to improving model recognition accuracy.
[0056] In some disclosed embodiments, the joint model also includes a speaker recognition network, the speaker recognition network includes a speaker feature extraction subnetwork, the speech emotion recognition device 50 includes a speaker feature extraction module, which is used to use the speaker feature extraction subnetwork to extract speaker features of the first sample speech to obtain speaker features of the first sample speech; the speech emotion recognition device 50 includes a mutual information calculation module, which is used to obtain a fourth loss based on the mutual information between the first emotion feature and the speaker feature; the parameter adjustment module is also used to obtain a total loss based on the first loss, the second loss, the third loss and the fourth loss; wherein the total loss is positively correlated with the fourth loss.
[0057] Therefore, by setting a speaker recognition network in the joint model, and the speaker recognition network includes a speaker feature extraction subnetwork, and using the speaker feature extraction subnetwork to extract speaker features of the first sample speech, the speaker features of the first sample speech are obtained, and then based on the mutual information between the first emotional feature and the speaker feature, the fourth loss is obtained, and then by minimizing the total loss as the model optimization goal, the mutual information between the first emotional feature and the speaker feature is made as small as possible, that is, the correlation between the first emotional feature and the speaker feature is reduced as much as possible, which can reduce the interference of speaker information on emotion recognition as much as possible, which is conducive to improving the accuracy of emotion recognition.
[0058] In some disclosed embodiments, the speaker recognition network also includes a speaker classification subnetwork, the first sample speech is also labeled with a sample speaker, and the speech emotion recognition device 50 includes a speaker prediction module, which is used to use the speaker classification subnetwork to perform speaker prediction on speaker features to obtain a predicted speaker of the first sample speech; the speech emotion recognition device 50 includes a speaker difference measurement module, which is used to obtain a fifth loss based on the difference between the sample speaker and the predicted speaker; the parameter adjustment module is also used to obtain a total loss based on the first loss, the second loss, the third loss, the fourth loss and the fifth loss; wherein the fifth loss is positively correlated with the total loss.
[0059] Therefore, by further predicting the predicted speaker of the first sample speech, and based on the difference between the predicted speaker and the sample speaker marked by the first sample speech, the fifth loss is obtained, so that the fifth loss can be further combined in the parameter adjustment process, and then the accuracy of speech features can be improved through the supervision of the total loss, which is conducive to improving the accuracy of emotion recognition.
[0060] In some disclosed embodiments, the speaker feature extraction subnetwork and the emotion feature extraction subnetwork share a speech feature extraction subnetwork, the speech feature extraction subnetwork is used to perform speech feature extraction, the emotion feature extraction subnetwork also includes a speech emotion encoding subnetwork, the speech emotion encoding subnetwork is used to perform emotion feature encoding based on speech features to obtain emotion features, and the speaker feature extraction subnetwork also includes a speaker encoding subnetwork for performing speaker encoding based on speech features to obtain speaker features.
[0061] Therefore, the speaker feature extraction subnetwork and the emotion feature extraction subnetwork share the speech feature extraction subnetwork. The speech feature extraction subnetwork is used to perform speech feature extraction. The emotion feature extraction subnetwork also includes a speech emotion encoding subnetwork. The speech emotion encoding subnetwork is used to perform emotion feature encoding based on speech features to obtain emotion features. The speaker feature extraction subnetwork also includes a speaker encoding subnetwork for performing speaker encoding based on speech features to obtain speaker features, which is conducive to reducing the complexity of the joint model as much as possible and improving model efficiency.
[0062] In some disclosed embodiments, the domain recognition network includes a first recognition network for performing domain recognition on a first sample speech and a second recognition network for performing domain recognition on a second sample speech, the first recognition network and the second recognition network have the same network structure and network parameters, and the first recognition network and the emotion recognition network share an emotion feature extraction subnetwork.
[0063] Therefore, by sharing the same network structure and even the same network parameters in the joint model, the data processing efficiency of the joint model can be improved.
[0064] See also Figure 6 , Figure 6 1 is a schematic diagram of a framework of an embodiment of an electronic device 60 of the present application. The electronic device 60 includes a memory 61 and a processor 62 coupled to each other, the memory 61 stores program instructions, and the processor 62 is used to execute the program instructions to implement the steps in any of the above-mentioned speech emotion recognition method embodiments. Specifically, the electronic device 60 may include but is not limited to: a desktop computer, a laptop computer, a server, a mobile phone, a tablet computer, etc., which are not limited here.
[0065] Specifically, the processor 62 is used to control itself and the memory 61 to implement the steps in any of the above-mentioned speech emotion recognition method embodiments. The processor 62 can also be called a CPU (Central Processing Unit). The processor 62 may be an integrated circuit chip with signal processing capabilities. The processor 62 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 62 can be implemented by an integrated circuit chip.
[0066] In the above scheme, on the one hand, the emotion recognition network is included in the joint model, and the emotion recognition network is used to perform emotion classification loss on the first sample speech labeled with the sample emotion category, and the first sample speech can be used to supervise the training of the network model. On the other hand, an unlabeled and sufficient second sample speech is introduced, and the joint model also includes a domain recognition network, and the domain classification loss of the first sample speech and the second sample speech is respectively performed through the domain recognition network, which is conducive to enabling the network model to process the speech data of the two data domains indiscriminately, thereby realizing the use of the second sample speech to assist the first sample speech to unsupervisedly train the network model, and then combining the advantages of supervised training and unsupervised training to collaboratively train the network model. Therefore, when the sample data with accurate emotion category annotations is relatively scarce, the accuracy of speech emotion recognition can be improved.
[0067] See also Figure 7 , Figure 7 Schematic diagram of a framework of an embodiment of a computer-readable storage medium 70 of the present application. The computer-readable storage medium 70 stores program instructions 71 that can be executed by a processor, and the program instructions 71 are used to implement the steps in any of the above-mentioned speech emotion recognition method embodiments.
[0068] In the above scheme, on the one hand, the emotion recognition network is included in the joint model, and the emotion recognition network is used to perform emotion classification loss on the first sample speech labeled with the sample emotion category, and the first sample speech can be used to supervise the training of the network model. On the other hand, an unlabeled and sufficient second sample speech is introduced, and the joint model also includes a domain recognition network, and the domain classification loss of the first sample speech and the second sample speech is respectively performed through the domain recognition network, which is conducive to enabling the network model to process the speech data of the two data domains indiscriminately, thereby realizing the use of the second sample speech to assist the first sample speech to unsupervisedly train the network model, and then combining the advantages of supervised training and unsupervised training to collaboratively train the network model. Therefore, when the sample data with accurate emotion category annotations is relatively scarce, the accuracy of speech emotion recognition can be improved.
[0069] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0070] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.
[0071] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0072] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0073] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0074] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
Claims
1. A speech emotion recognition method, characterized in that: include: Get the speech to be recognized; Recognize the speech to be recognized by using an emotion recognition network to obtain the emotion category of the speech to be recognized; Among them, the emotion recognition network is included in the joint model, and the joint model also includes a domain recognition network. The joint model is obtained by jointly training the emotion classification loss of the emotion recognition network on the first sample speech belonging to the first data domain category and the domain classification loss of the domain recognition network on the first sample speech and the second sample speech respectively, and the second sample speech belongs to the second data domain category, the first sample speech is annotated with a sample emotion category, the emotion classification loss is positively correlated with the total loss of the joint model, the domain classification loss is negatively correlated with the total loss of the joint model, and the domain classification loss includes the domain classification loss of the first sample speech and the domain classification loss of the second sample speech.
2. The method according to claim 1, characterized in that The emotion recognition network and the domain recognition network share an emotion feature extraction subnetwork, and the training steps of the joint model include: Using the emotion feature extraction subnetwork to extract emotion features from the first sample speech and the second sample speech respectively, to obtain a first emotion feature of the first sample speech and a second emotion feature of the second sample speech; Performing emotion category prediction based on the first emotion feature to obtain a predicted emotion category of the first sample speech, and performing domain category prediction based on the first emotion feature and the second emotion feature to obtain a first predicted domain category to which the first sample speech belongs and a second predicted domain category to which the second sample speech belongs; Obtaining a first loss, a second loss, and a third loss based on the difference between the first prediction domain category and the first data domain category, the difference between the second prediction domain category and the second data domain category, and the difference between the predicted emotion category and the sample emotion category, respectively; A total loss is obtained based on the first loss, the second loss, and the third loss, and a network parameter of the joint model is adjusted based on the total loss.
3. The method according to claim 2, characterized in that The emotion recognition network also includes an emotion classification subnetwork, and the emotion classification subnetwork is used to perform the emotion category prediction; And / or, the domain identification network further includes a domain classification subnetwork, and the domain classification subnetwork is used to perform the domain category prediction.
4. The method according to claim 2, characterized in that: The first sample speech is separated from a sample video, and a sample face image is also separated from the sample video. The joint model also includes an image feature extraction network. Before predicting the emotion category based on the first emotion feature to obtain the predicted emotion category of the first sample speech, the method also includes: Using the image feature extraction network to extract image features from the sample face image, to obtain sample image features of the sample face image; Fusing the first emotion feature with the sample image feature to obtain a sample fusion feature; The step of predicting the emotion category based on the first emotion feature to obtain the predicted emotion category of the first sample speech includes: Emotion category prediction is performed based on the sample fusion features to obtain the predicted emotion category.
5. The method according to claim 2, characterized in that: The joint model further includes a speaker recognition network, the speaker recognition network includes a speaker feature extraction subnetwork, and before obtaining the total loss based on the first loss, the second loss and the third loss, the method further includes: Using the speaker feature extraction subnetwork to extract speaker features from the first sample speech to obtain speaker features of the first sample speech; Obtaining a fourth loss based on mutual information between the first emotion feature and the speaker feature; The obtaining of a total loss based on the first loss, the second loss and the third loss comprises: The total loss is obtained based on the first loss, the second loss, the third loss and the fourth loss; wherein the total loss is positively correlated with the fourth loss.
6. The method according to claim 5, characterized in that The speaker recognition network further includes a speaker classification subnetwork, the first sample speech is further labeled with a sample speaker, and before obtaining the total loss based on the first loss, the second loss, the third loss, and the fourth loss, the method further includes: Using the speaker classification subnetwork to perform speaker prediction on the speaker features to obtain a predicted speaker of the first sample speech; Obtaining a fifth loss based on a difference between the sample speaker and the predicted speaker; The obtaining the total loss based on the first loss, the second loss, the third loss and the fourth loss comprises: The total loss is obtained based on the first loss, the second loss, the third loss, the fourth loss and the fifth loss; wherein the fifth loss is positively correlated with the total loss.
7. The method according to claim 5, characterized in that The speaker feature extraction subnetwork and the emotion feature extraction subnetwork share a speech feature extraction subnetwork, the speech feature extraction subnetwork is used to perform speech feature extraction, the emotion feature extraction subnetwork also includes a speech emotion encoding subnetwork, the speech emotion encoding subnetwork is used to perform emotion feature encoding based on speech features to obtain emotion features, and the speaker feature extraction subnetwork also includes a speaker encoding subnetwork for performing speaker encoding based on speech features to obtain speaker features.
8. The method according to claim 1, characterized in that The domain recognition network includes a first recognition network for performing domain recognition on the first sample speech and a second recognition network for performing domain recognition on the second sample speech, the first recognition network and the second recognition network have the same network structure and network parameters, and the first recognition network and the emotion recognition network share an emotion feature extraction subnetwork.
9. A speech emotion recognition device, characterized in that: include: A speech acquisition module, used to acquire the speech to be recognized; An emotion recognition module, used to recognize the speech to be recognized by using an emotion recognition network to obtain the emotion category of the speech to be recognized; Among them, the emotion recognition network is included in the joint model, and the joint model also includes a domain recognition network. The joint model is obtained by jointly training the emotion classification loss of the emotion recognition network on the first sample speech belonging to the first data domain category and the domain classification loss of the domain recognition network on the first sample speech and the second sample speech respectively, and the second sample speech belongs to the second data domain category, the first sample speech is annotated with a sample emotion category, the emotion classification loss is positively correlated with the total loss of the joint model, the domain classification loss is negatively correlated with the total loss of the joint model, and the domain classification loss includes the domain classification loss of the first sample speech and the domain classification loss of the second sample speech.
10. An electronic device, characterized in that: It comprises a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the speech emotion recognition method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the speech emotion recognition method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Emotion recognition method and device based on LSTM audio and video fusion and storage medium
CN110826466A
Emotion recognition model training method and device, computer equipment and storage medium
CN111933187A
Cross-database speech emotion recognition method and device based on multi-scale difference confrontation
CN112489689A
Emotion recognition method and device, computer equipment and storage medium
CN112949708A