Voice emotion recognition model training method, recognition method and device
By acquiring speech sample pairs from different languages and training a speech emotion recognition model using cross-entropy loss and local feature alignment algorithms, the accuracy problem of multilingual emotion recognition was solved, and the accuracy of cross-lingual emotion recognition was improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2023-03-20
- Publication Date
- 2026-05-22
AI Technical Summary
Existing voice emotion recognition technology is mainly applicable to a single language and has difficulty accurately recognizing emotions in multiple languages, resulting in low recognition accuracy.
By acquiring multiple speech sample pairs, including first speech samples with speech emotion labels and second speech samples without speech emotion labels in different languages, the initial speech emotion recognition model is trained using a feature extractor and a recognizer. The model parameters are then updated using cross-entropy loss and local feature alignment algorithms to improve the model's recognition ability.
The trained speech emotion recognition model can accurately identify speech emotions in different languages, improving the accuracy of the recognition results.
Smart Images

Figure CN116486785B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech recognition technology, and in particular to a training method, recognition method and apparatus for a speech emotion recognition model. Background Technology
[0002] With the continuous development of human-computer interaction, emotional expression in human-computer interaction is receiving more and more attention, especially for voice services. If intelligent interactive devices can accurately understand the meaning expressed by users without contact with them, they can provide users with more humanized services.
[0003] In existing technologies, speech emotion recognition technology is mostly applicable to monolingual scenarios. However, in practical applications, the same system needs to recognize more than one language, and different languages will affect the accuracy of speech emotion recognition. Therefore, how to accurately recognize speech emotions in different languages is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0004] This invention provides a training method, recognition method, and apparatus for a speech emotion recognition model, which can accurately recognize speech emotions in different languages, thereby improving the accuracy of the recognition results.
[0005] This application provides a training method for a speech emotion recognition model, including:
[0006] Multiple voice sample pairs are acquired. Each voice sample pair includes a first voice sample with a voice emotion label and a second voice sample without a voice emotion label. The first voice sample and the second voice sample belong to different languages.
[0007] For each speech sample pair, the speech features corresponding to the speech sample pair are input into the initial speech emotion recognition model to obtain the first prediction result corresponding to the first speech sample and the second prediction result corresponding to the second speech sample in the speech sample pair.
[0008] Based on the first prediction result, the second prediction result, and the voice emotion label corresponding to each voice sample, the model parameters of the initial voice emotion recognition model are updated to obtain the voice emotion recognition model.
[0009] According to the training method of the speech emotion recognition model provided in this application, the initial speech emotion recognition model includes a feature extractor and a recognizer. The step of inputting the speech features corresponding to the speech sample pair into the initial speech emotion recognition model to obtain a first prediction result corresponding to the first speech sample and a second prediction result corresponding to the second speech sample in the speech sample pair includes:
[0010] The speech features corresponding to the speech sample pairs are input into the feature extractor to obtain the first hidden layer features corresponding to the first speech sample and the second hidden layer features corresponding to the second speech sample.
[0011] The first hidden layer features and the second hidden layer features are respectively input into the recognizer to obtain the first prediction result corresponding to the first speech sample and the second prediction result corresponding to the second speech sample.
[0012] According to the training method of the speech emotion recognition model provided in this application, the step of updating the model parameters of the initial speech emotion recognition model based on the first prediction result, the second prediction result, and the speech emotion label for each speech sample includes:
[0013] Based on the first prediction result corresponding to each speech sample pair and the speech emotion label, construct the cross-entropy loss corresponding to the multiple speech sample pairs.
[0014] Based on the voice emotion label corresponding to each voice sample pair and the second prediction result, a local feature alignment algorithm is used to determine the similarity between the multiple voice sample pairs.
[0015] The model parameters of the initial speech emotion recognition model are updated based on the cross-entropy loss and similarity of the multiple speech sample pairs.
[0016] According to the training method of the speech emotion recognition model provided in this application, the step of determining the similarity of the multiple speech sample pairs based on the speech emotion label corresponding to each speech sample pair and the second prediction result using a local feature alignment algorithm includes:
[0017] Based on the voice emotion label corresponding to each voice sample pair and the second prediction result, a first target voice sample and a second target voice sample corresponding to each voice emotion category are determined from the plurality of voice sample pairs.
[0018] Based on the first target speech sample and the second target speech sample corresponding to each speech emotion category, a local feature alignment algorithm is used to determine the similarity of the multiple speech sample pairs.
[0019] According to the training method of the speech emotion recognition model provided in this application, the step of determining the similarity of the multiple speech sample pairs corresponding to the first target speech sample and the second target speech sample corresponding to each speech emotion category using a local feature alignment algorithm includes:
[0020] according to The similarity of the multiple speech sample pairs is determined by using a local feature alignment algorithm.
[0021] in, This represents the distance between multiple pairs of speech samples, where the distance is used to characterize the similarity. c represents the number of voice emotion categories. Voice emotion category c out of 1 voice emotion category, This represents the first set consisting of the first target speech samples corresponding to the speech emotion category c. This represents the second set consisting of the second target speech samples corresponding to the speech emotion category c. Represents the first set of... The first target speech sample, Indicates the first The weights corresponding to the first target speech sample Indicates the first The mapping from the first target speech sample to the regenerated Hilbert space RKHS. Represents the first in the second set A second target speech sample, Indicates the first The weights corresponding to each second target speech sample Indicates the first Mapping of a second target speech sample to RKHS.
[0022] According to the training method for a speech emotion recognition model provided in this application, the method further includes:
[0023] according to Determine the first The weights corresponding to the first target speech samples.
[0024] according to Determine the first The weights corresponding to each second target speech sample.
[0025] in, Indicates the first The probability value of the voice emotion label corresponding to each first target voice sample; Indicates the first The probability value of the second prediction result corresponding to each second target speech sample.
[0026] According to the training method for a speech emotion recognition model provided in this application, the method further includes:
[0027] The first and second speech samples in the speech sample pair are preprocessed using the log-Mel spectrum algorithm to obtain the speech features corresponding to the speech sample pair.
[0028] This application also provides a voice emotion recognition method, the method comprising:
[0029] Obtain the speech to be recognized.
[0030] The speech to be recognized is input into the speech emotion recognition model to obtain the speech emotion result corresponding to the speech to be recognized. The speech emotion recognition model is any of the speech emotion recognition models described above.
[0031] This application also provides a training device for a speech emotion recognition model, comprising:
[0032] The acquisition unit is used to acquire multiple speech sample pairs, each speech sample pair including a first speech sample with a speech emotion label and a second speech sample without a speech emotion label, wherein the first speech sample and the second speech sample belong to different languages.
[0033] The processing unit is configured to input the speech features corresponding to each speech sample pair into an initial speech emotion recognition model to obtain a first prediction result corresponding to the first speech sample and a second prediction result corresponding to the second speech sample in the speech sample pair.
[0034] The update unit is used to update the model parameters of the initial speech emotion recognition model according to the first prediction result, the second prediction result and the speech emotion label of each speech sample, so as to obtain the speech emotion recognition model.
[0035] According to the training apparatus for a speech emotion recognition model provided in this application, the initial speech emotion recognition model includes a feature extractor and a recognizer, and the processing unit is specifically used for:
[0036] The speech features corresponding to the speech sample pairs are input into the feature extractor to obtain the first hidden layer features corresponding to the first speech sample and the second hidden layer features corresponding to the second speech sample; the first hidden layer features and the second hidden layer features are respectively input into the recognizer to obtain the first prediction result corresponding to the first speech sample and the second prediction result corresponding to the second speech sample.
[0037] According to the training apparatus for a speech emotion recognition model provided in this application, the updating unit is specifically used for:
[0038] Based on the first prediction result and the voice emotion label corresponding to each voice sample pair, a cross-entropy loss is constructed for each voice sample pair; based on the voice emotion label and the second prediction result corresponding to each voice sample pair, a local feature alignment algorithm is used to determine the similarity of the voice sample pairs; based on the cross-entropy loss and the similarity of the voice sample pairs, the model parameters of the initial voice emotion recognition model are updated.
[0039] According to the training apparatus for a speech emotion recognition model provided in this application, the updating unit is specifically used for:
[0040] Based on the voice emotion label corresponding to each voice sample pair and the second prediction result, a first target voice sample and a second target voice sample corresponding to each voice emotion category are determined from the plurality of voice sample pairs; based on the first target voice sample and the second target voice sample corresponding to each voice emotion category, a local feature alignment algorithm is used to determine the similarity corresponding to the plurality of voice sample pairs.
[0041] According to the training apparatus for a speech emotion recognition model provided in this application, the updating unit is specifically used for:
[0042] according to The similarity of the multiple speech sample pairs is determined by using a local feature alignment algorithm.
[0043] in, This represents the distance between multiple pairs of speech samples, where the distance is used to characterize the similarity. c represents the number of voice emotion categories. Voice emotion category c out of 1 voice emotion category, This represents the first set consisting of the first target speech samples corresponding to the speech emotion category c. This represents the second set consisting of the second target speech samples corresponding to the speech emotion category c. Represents the first set of... The first target speech sample, Indicates the first The weights corresponding to the first target speech sample Indicates the first The mapping from the first target speech sample to the regenerated Hilbert space RKHS. Represents the first in the second set A second target speech sample, Indicates the first The weights corresponding to each second target speech sample Indicates the first Mapping of a second target speech sample to RKHS.
[0044] According to the training device for a speech emotion recognition model provided in this application, the device further includes a determination unit.
[0045] The determining unit is used to determine based on Determine the first The weights corresponding to the first target speech samples; based on Determine the first The weights corresponding to each second target speech sample.
[0046] in, Indicates the first The probability value of the voice emotion label corresponding to each first target voice sample; Indicates the first The probability value of the second prediction result corresponding to each second target speech sample.
[0047] According to the training device for a speech emotion recognition model provided in this application, the device further includes a preprocessing unit.
[0048] The preprocessing unit is used to preprocess the first speech sample and the second speech sample in the speech sample pair using the log-Mel spectrum algorithm to obtain the speech features corresponding to the speech sample pair.
[0049] This application also provides a voice emotion recognition device, the device comprising:
[0050] The acquisition unit is used to acquire the speech to be recognized.
[0051] The processing unit is used to input the speech to be recognized into the speech emotion recognition model to obtain the speech emotion result corresponding to the speech to be recognized, wherein the speech emotion recognition model is any of the speech emotion recognition models described above.
[0052] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a training method for the speech emotion recognition model as described above or the speech emotion recognition method.
[0053] This application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the training method of the speech emotion recognition model or the speech emotion recognition method as described above.
[0054] This application also provides a computer program product, including a computer program that, when executed by a processor, implements a training method for the speech emotion recognition model or the speech emotion recognition method as described above.
[0055] The training method, recognition method, and apparatus for the speech emotion recognition model provided in this application, when training the speech emotion recognition model, can first acquire multiple speech sample pairs. Each speech sample pair includes a first speech sample with a speech emotion label and a second speech sample without a speech emotion label. The first and second speech samples belong to different languages. For each speech sample pair, the speech features corresponding to the speech sample pair are input into the initial speech emotion recognition model to obtain a first prediction result corresponding to the first speech sample and a second prediction result corresponding to the second speech sample in the speech sample pair. Based on the first prediction result, the second prediction result, and the speech emotion label corresponding to each speech sample pair, the model parameters of the initial speech emotion recognition model are updated to obtain the trained speech emotion recognition model. The speech emotion recognition model trained in this way can obtain emotion recognition results from the initial speech emotion recognition model of a large language corpus even when there are no speech emotion labels in the small language corpus. This result can then participate in the training of the speech emotion recognition model, enabling the subsequently trained speech emotion recognition model to accurately recognize speech emotions in different languages, thereby improving the accuracy of the recognition results. Attached Figure Description
[0056] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0057] Figure 1 A flowchart illustrating the training method for the speech emotion recognition model provided in this application embodiment;
[0058] Figure 2 A flowchart illustrating a method for preprocessing speech samples provided in this application embodiment;
[0059] Figure 3 A schematic diagram of the architecture of the initial speech emotion recognition model provided in the embodiments of this application;
[0060] Figure 4 A flowchart illustrating the speech emotion recognition method provided in this application embodiment;
[0061] Figure 5 A schematic diagram of the structure of a training device for a speech emotion recognition model provided in an embodiment of this application;
[0062] Figure 6 This is a schematic diagram of the structure of the voice emotion recognition device provided in the embodiments of this application;
[0063] Figure 7 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0064] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0065] In the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone, where A and B can be singular or plural. In the textual description of this application, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0066] The technical solutions provided in this application can be applied to speech recognition scenarios, especially speech emotion recognition scenarios, such as common intelligent customer service and whole-house intelligent systems. Existing speech emotion recognition technologies are usually single-language speech emotion recognition, that is, they can only recognize speech emotions in one language.
[0067] However, in practical applications, users of the same system often speak more than one language. Currently, considering that minority language data is extremely difficult to train due to the scarcity of labels, and that there is not much work on multilingual speech emotion recognition, the few existing works are based on fine-tuning of speech emotion recognition models trained on large language corpora, and then using the fine-tuned speech emotion recognition models to recognize minority language speech.
[0068] For example, the language corpora currently widely used in the industry include IEMOCAP (English), EMO-DB (German), SAVEE (English), and EMOVO (Italian). Among these four language corpora, IEMOCAP has a larger data volume and can be referred to as a major language corpus; EMO-DB (German), SAVEE (English), and EMOVO (Italian) have smaller data volumes and can be referred to as minor language corpora.
[0069] Considering that different languages can affect the accuracy of speech emotion recognition, and that the fine-tuned speech emotion recognition model cannot accurately recognize the speech emotions of less common languages, the accuracy of the recognition results will be low.
[0070] Therefore, in order to accurately identify the emotions in speech in different languages, this application provides a method for training a speech emotion recognition model. By acquiring multiple speech sample pairs, including a first speech sample and a second speech sample with speech emotion labels, wherein the first speech sample and the second speech sample belong to different languages, and training an initial speech emotion recognition model based on multiple speech sample pairs including different languages, the trained speech emotion recognition model can accurately identify the emotions in speech in different languages, thereby improving the accuracy of the recognition results.
[0071] Before detailing the speech emotion recognition model training method provided in the embodiments of this application, the basic concepts involved in this application will be introduced first.
[0072] Voice emotion recognition refers to the simulation of how computers perceive and understand human speech information, analyzing speech information, and identifying the speaker's emotional information.
[0073] Transfer learning is a type of machine learning method that aims to reduce the differences between different data distributions, so that existing knowledge can be referenced when applying it to new data distributions.
[0074] Feature alignment is a method for measuring the similarity of different data distributions, which can bring the distributions of different data sources closer together in the latent space.
[0075] Pre-emphasis: This is a signal processing method that compensates for the high-frequency components of the input speech signal at the transmitting end. It keeps the low-frequency components of the speech signal unchanged and amplifies the high-frequency components to compensate for the excessive attenuation of the high-frequency components during transmission.
[0076] Framing is a technique used by computers to segment speech signals according to a specified length (time period or number of samples) in order to enable batch processing of speech signals.
[0077] Windowing is a technique that multiplies segmented frame data with a data segment of the same length to ensure the continuity of the speech signal. This data segment of the same length is the data of the window function over the entire period.
[0078] Fast Fourier Transform (FFT) is a general term for efficient and fast calculation methods of the Discrete Fourier Transform.
[0079] Mel filter bank: refers to multiple bandpass filters. At Mel frequencies, the passbands of the bandpass filters are of equal width. The purpose is to simulate human ear perception of sound by using nonlinear mapping to make the spectrum more resolution at lower frequencies and less resolution at higher frequencies.
[0080] The training method of the speech emotion recognition model provided in this application will be described in detail below through several specific embodiments. It is understood that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0081] Figure 1 This is a flowchart illustrating the training method for a speech emotion recognition model provided in an embodiment of this application. The training method for this speech emotion recognition model can be executed by software and / or hardware devices. For example, please refer to [link to example]. Figure 1 As shown, the training method for this speech emotion recognition model may include:
[0082] S101. Obtain multiple speech sample pairs, each speech sample pair including a first speech sample with a speech emotion label and a second speech sample without a speech emotion label, the first speech sample and the second speech sample belong to different languages.
[0083] In this context, multiple speech sample pairs can be understood as training samples from the language corpora used to train the speech emotion recognition model. The first speech sample and the second speech sample are speech samples from different languages within the speech sample pair, and the speech emotion label can be understood as the true speech emotion category to which the speech emotion in the first speech sample belongs.
[0084] For example, when acquiring multiple speech sample pairs, they can be obtained from four language corpora: IEMOCAP, EMO-DB (German), SAVEE (English), and EMOVO (Italian). For instance, the first speech sample and its associated emotion tag can be obtained from IEMOCAP, and the second speech sample from EMO-DB (German); or, for another example, the first speech sample and its associated emotion tag can be obtained from IEMOCAP, and the second speech sample from EMOVO (Italian), and so on.
[0085] It is understood that in the embodiments of this application, the aforementioned speech sample pairs have the same feature space, that is, they contain the same feature dimensions and number of dimensions; each speech sample pair has the same speech emotion label space, that is, the speech emotion classification task space is the same. For example, the speech emotion categories in the speech emotion classification task space may include four categories: happy, sad, neutral, and angry. It should be noted that, because each speech sample pair has a different probability distribution, even if each speech sample pair has the same feature space and speech emotion label space, the probability of the final identified speech emotion category will be different.
[0086] After obtaining multiple speech sample pairs, these multiple speech sample pairs can be combined to train and obtain a speech emotion recognition model, that is, to execute the following S102 and S103:
[0087] S102. For each speech sample pair, input the speech features corresponding to the speech sample pair into the initial speech emotion recognition model to obtain the first prediction result corresponding to the first speech sample and the second prediction result corresponding to the second speech sample in the speech sample pair.
[0088] Typically, before inputting the speech features corresponding to a speech sample pair into the initial speech emotion recognition model, it is necessary to first obtain the speech features corresponding to the speech sample pair. These speech features refer to the feature vectors that characterize speech emotions.
[0089] For example, in the embodiments of this application, when obtaining the speech features corresponding to a speech sample pair, the Log-Mel Spectrum (Filter Bank) algorithm can be used to preprocess the first speech sample and the second speech sample in the speech sample pair respectively to obtain the speech features corresponding to the speech sample pair; the MFCC algorithm can also be used to preprocess the first speech sample and the second speech sample in the speech sample pair respectively to obtain the speech features corresponding to the speech sample pair; or the LPC algorithm can also be used to preprocess the first speech sample and the second speech sample in the speech sample pair respectively to obtain the speech features corresponding to the speech sample pair, etc. The specific settings can be made according to actual needs.
[0090] Taking the preprocessing of either the first or second speech sample in a speech sample pair using the log-Mel spectrum algorithm as an example, see [link to example]. Figure 2 As shown, Figure 2 This is a flowchart illustrating a method for preprocessing speech samples provided in an embodiment of this application, combined with... Figure 2As shown, when preprocessing speech sample pairs using the log-Mel spectrum algorithm, the speech samples can first be pre-emphasized to keep the low-frequency part of the speech signal unchanged and increase the energy of the high-frequency part of the speech signal. For example, a first-order high-pass filter can be used for pre-emphasis processing to obtain pre-emphasized speech features. The pre-emphasized speech features are then framed to obtain framed speech features. The framed speech features are then windowed with a window length of 1024 points and an overlap length of 512 points to obtain windowed speech features. The windowed speech features are then subjected to a fast Fourier transform to obtain fast Fourier transform speech features. The fast Fourier transform speech features are then subjected to Mel filtering to obtain Mel-filtered speech features. Finally, a logarithmic operation is performed on the Mel-filtered speech features to obtain the speech features corresponding to the speech sample.
[0091] For example, in an embodiment of this application, the dimension size of the feature dimension form corresponding to the speech sample extracted by log-Mel spectrum is ,in This is the number of Mel filters, which is chosen as 512 here. It is the size of the time frame.
[0092] Based on the above description, after obtaining the speech features corresponding to the speech sample pairs, the speech features corresponding to the speech sample pairs can be input into the initial speech emotion recognition model. For example, in this embodiment, considering that directly using multiple speech sample pairs from this embodiment for training may easily cause the model to fail to converge, it is possible to first train based on the IEMOCAP large language corpus and then transfer the training to the speech emotion recognition model framework to obtain the initial speech emotion recognition model.
[0093] Considering that the part representing emotion in a speech segment may not appear in the specific part of the sentence, the initial speech emotion recognition model can employ a bidirectional long short-term memory (BiLSTM) network. The feature extractor can extract speech features corresponding to speech sample pairs from two directions, which is beneficial for extracting speech emotion features. The specific process may include: obtaining the log-Mel spectrum features of speech samples from the IEMOCAP large language corpus; inputting the log-Mel spectrum features into the BiLSTM model to learn the log-Mel spectrum features, obtaining the speech features of the speech samples; and inputting the learned speech features into an emotion classifier (fully connected layer) to obtain the speech emotion recognition result corresponding to the speech sample, thereby training the transfer to the speech emotion recognition model framework and obtaining the initial speech emotion recognition model.
[0094] For example, the BiLSTM model can be an attention mechanism that calculates weights for each frame of the input speech sample along with all other frames. The speech emotion recognition result encoded by BiLSTM has 256 dimensions.
[0095] For example, in an embodiment of this application, the initial speech emotion recognition model includes a feature extractor and a recognizer. For example, see [link to relevant documentation]. Figure 3 As shown, Figure 3 This is a schematic diagram of the architecture of the initial speech emotion recognition model provided in the embodiments of this application, wherein the feature extractor is used to extract the hidden layer features corresponding to the speech sample pairs, and the recognizer is used to recognize the prediction results corresponding to the speech samples in the speech sample pairs.
[0096] Combination Figure 3 As shown, when the initial speech emotion recognition model includes a feature extractor and a recognizer, the speech features corresponding to the speech sample pairs can be input into the feature extractor to obtain the first hidden layer features corresponding to the first speech sample and the second hidden layer features corresponding to the second speech sample. Then, the first hidden layer features and the second hidden layer features are respectively input into the recognizer to obtain the first prediction result corresponding to the first speech sample and the second prediction result corresponding to the second speech sample.
[0097] Based on the above description, the first prediction result and the second prediction result corresponding to each speech sample pair can be obtained, and then the following S103 is executed:
[0098] S103. Based on the first prediction result, the second prediction result, and the voice emotion label for each voice sample, update the model parameters of the initial voice emotion recognition model to obtain the voice emotion recognition model.
[0099] For example, in this embodiment of the application, when updating the model parameters of the initial speech emotion recognition model based on the first prediction result, the second prediction result, and the speech emotion label corresponding to each speech sample pair, multiple speech sample pairs can be constructed according to the first prediction result and the speech emotion label corresponding to each speech sample pair; and a local feature alignment algorithm is used to determine the similarity of multiple speech sample pairs according to the speech emotion label and the second prediction result corresponding to each speech sample pair; and the model parameters of the initial speech emotion recognition model are updated according to the cross-entropy loss and similarity of multiple speech sample pairs.
[0100] Among them, cross-entropy loss can characterize the degree of closeness between the first prediction result and the voice emotion label. The similarity can be represented by the distance between the vectors of the voice sample pairs. The closer the distance, the greater the similarity.
[0101] For example, in this embodiment of the application, when constructing the cross-entropy loss corresponding to multiple speech sample pairs based on the first prediction result and speech emotion label corresponding to each speech sample pair, for each speech sample pair, the cross-entropy loss corresponding to that speech sample is constructed based on the first prediction result and speech emotion label corresponding to the speech sample; the average value of the cross-entropy loss corresponding to each speech sample is calculated to obtain the cross-entropy loss corresponding to multiple speech sample pairs.
[0102] For example, in the embodiments of this application, when determining the similarity of multiple speech sample pairs using a local feature alignment algorithm based on the speech emotion labels and second prediction results corresponding to each speech sample pair, the first target speech sample and the second target speech sample corresponding to each speech emotion category can be determined from the multiple speech sample pairs firstly based on the speech emotion labels and second prediction results corresponding to each speech sample pair; then, based on the first target speech sample and the second target speech sample corresponding to each speech emotion category, the local feature alignment algorithm is used to determine the similarity of the multiple speech sample pairs.
[0103] For example, suppose multiple speech sample pairs include first speech sample 1-1, first speech sample 1-2, ..., first speech sample 1-11, first speech sample 1-12, second speech sample 2-1, second speech sample 2-2, ..., second speech sample 2-11, second speech sample 2-12. Based on the speech emotion labels corresponding to each of the first speech samples 1-1, first speech sample 1-2, ..., first speech sample 1-11, and first speech sample 1-12, and the second prediction results corresponding to each of the second speech samples 2-1, second speech sample 2-2, ..., second speech sample 2-11, and second speech sample 2-12, the first and second speech samples corresponding to the speech emotion categories "happy," "sad," "neutral," and "angry" can be determined from the above 24 speech samples. For distinction, the determined first and second speech samples can be denoted as the first target speech sample and the second target speech sample, respectively.
[0104] Based on the above description, assuming that the first target speech samples corresponding to the speech emotion category "happy" include first speech sample 1-1, first speech sample 1-2, and first speech sample 1-3, and the corresponding second target speech samples include second speech sample 2-10, second speech sample 2-11, and second speech sample 2-12; the first target speech samples corresponding to the speech emotion category "sad" include first speech sample 1-4, first speech sample 1-5, and first speech sample 1-6, and the corresponding second target speech samples include second speech sample 2-7, second speech sample 2-8, and second speech sample 2-9; and the speech emotion category "neutral" corresponds to... The first target speech samples include first speech samples 1-7, 1-8, and 1-9, and the corresponding second target speech samples include second speech samples 2-4, 2-5, and 2-6. The first target speech samples corresponding to the speech emotion category "anger" include first speech samples 1-10, 1-11, and 1-12, and the corresponding second target speech samples include second speech samples 2-1, 2-2, and 2-3, thereby determining the first and second target speech samples corresponding to each speech emotion category.
[0105] For example, when determining the similarity of multiple speech sample pairs using a local feature alignment algorithm based on the first and second target speech samples corresponding to each speech emotion category, a fine-grained local feature alignment method can be used, such as the Local Maximum Mean Discrepancy (LMMD) algorithm, as shown below:
[0106] according to A local feature alignment algorithm is used to determine the similarity of multiple speech sample pairs.
[0107] in, This represents the distance between multiple speech sample pairs; the distance is used to characterize similarity. c represents the number of voice emotion categories. Voice emotion category c out of 1 voice emotion category, This represents the first set consisting of the first target speech samples corresponding to the speech emotion category c. This represents the second set consisting of the second target speech samples corresponding to the speech emotion category c. Represents the first set of elements. The first target speech sample, Indicates the first The weights corresponding to the first target speech sample Indicates the first The mapping from the first target speech sample to the Reproducing Kernel Hilbert Space (RKHS) Represents the first in the second set A second target speech sample, Indicates the first The weights corresponding to each second target speech sample Indicates the first Mapping of a second target speech sample to RKHS.
[0108] It should be noted that when the above formula uses the local feature alignment algorithm to determine the similarity of multiple speech sample pairs, in order to reduce the error introduced by the second prediction result corresponding to the second speech sample, i.e., the pseudo label, a corresponding weight is set for each speech sample.
[0109] The above method employs a fine-grained feature alignment approach, such as the Local Maximum Mean Discrepancy (LMMD) algorithm, when determining the similarity between multiple speech sample pairs. This approach fully considers the inconsistent edge distribution of different emotion categories, enabling the acquisition of emotion recognition results from the initial speech emotion recognition model of the large language corpus even when there are no speech emotion labels in the small language corpus. These results then participate in the training of the speech emotion recognition model, allowing the subsequently trained speech emotion recognition model to accurately identify speech emotions in different languages, thereby effectively improving the accuracy of the recognition results.
[0110] For example, when determining the weights corresponding to the first target speech sample, it can be based on Determine the first The weights corresponding to the first target speech sample; when determining the weights corresponding to the second target speech sample, it can be based on Determine the first The weights corresponding to each second target speech sample.
[0111] in, Indicates the first The probability value of the voice emotion label corresponding to each first target voice sample; Indicates the first The probability value of the second prediction result corresponding to each second target speech sample.
[0112] After determining the cross-entropy loss and similarity of multiple speech sample pairs, the model parameters of the initial speech emotion recognition model can be updated based on the cross-entropy loss and similarity of the multiple speech sample pairs.
[0113] For example, when updating the model parameters of the initial speech emotion recognition model based on the cross-entropy loss and similarity of multiple speech samples, it can be determined whether the updated speech emotion recognition model meets preset conditions. If the preset conditions are met, the updated speech emotion recognition model is determined as the final speech emotion recognition model; if the preset conditions are not met, the updated speech emotion recognition model is used as the new initial speech emotion recognition model, and the model parameters of the new initial speech emotion recognition model are updated again until the preset conditions are met. For example, the preset conditions may include the number of model updates reaching a preset threshold, and / or the updated speech emotion recognition model converging.
[0114] As can be seen in this embodiment, when training the speech emotion recognition model, multiple speech sample pairs with first and second speech samples having speech emotion labels are obtained. The first and second speech samples belong to different languages. The initial speech emotion recognition model is trained based on multiple speech sample pairs including different languages, so that the trained speech emotion recognition model can accurately recognize speech emotions in different languages, thereby improving the accuracy of the recognition results.
[0115] Based on the above Figure 1 In the embodiments shown, in order to better verify that the technical solution provided in this application can accurately identify the emotional expression of speech in different languages, the embodiments of this application will be verified by combining three sets of experimental data.
[0116] The first set of experimental data includes a first speech sample and its associated emotion tag obtained from IEMOCAP, and a second speech sample obtained from EMO-DB (German). The second set of experimental data includes a first speech sample and its associated emotion tag obtained from IEMOCAP, and a second speech sample obtained from SAVEE (English). The third set of experimental data includes a first speech sample and its associated emotion tag obtained from IEMOCAP, and a second speech sample obtained from EMOVO (Italian). Compared to existing technologies that perform speech emotion recognition without feature alignment, or that use traditional Maximum Mean Discrepancy (MMD) algorithms for speech emotion recognition, this embodiment employs a fine-grained feature alignment method, such as Local Maximum Mean Discrepancy (LMMD), to perform feature alignment before speech emotion recognition. This fully considers the issue of edge distribution, thereby effectively improving the accuracy of speech emotion recognition results.
[0117] For example, referring to Table 1 below, for the three sets of experimental data, speech emotion recognition was performed directly without feature alignment, speech emotion recognition was performed using a traditional feature alignment algorithm, and speech emotion recognition was performed after fine-grained local feature alignment in this application. The Adam optimizer was used and the learning rate was set to 0.0001 to obtain the speech emotion recognition results.
[0118] Table 1
[0119] Experimental data Accuracy obtained from direct identification Accuracy obtained by traditional feature alignment algorithms The accuracy obtained in this application EMO-DB 40.71% 42.48% 54.28% SAVEE 42.33% 49.67% 55.33% EMOVO 28.87% 44.35% 45.24%
[0120] As shown in Table 1, when the experimental data includes the second speech sample obtained from EMO-DB (German), the speech emotion recognition result obtained by directly performing speech emotion recognition without feature alignment is 40.71%, the speech emotion recognition result obtained by performing speech emotion recognition using the traditional feature alignment algorithm is 42.48%, and the speech emotion recognition result obtained by performing speech emotion recognition using the fine-grained local feature alignment provided in this application is 54.28%. Therefore, it can be clearly seen that compared with the direct recognition result and the speech emotion recognition result obtained by using the traditional feature alignment algorithm, the speech emotion recognition result obtained by the technical solution provided in this application has an improvement of 13.57% and 11.80% respectively, and the accuracy is higher.
[0121] When the experimental data included a second speech sample obtained from SAVEE (in English), the speech emotion recognition result obtained by directly performing speech emotion recognition without feature alignment was 42.33%, the speech emotion recognition result obtained by using the traditional feature alignment algorithm was 49.67%, and the speech emotion recognition result obtained by using the fine-grained local feature alignment provided in this application before performing speech emotion recognition was 55.33%. Therefore, it can be clearly seen that compared with the direct recognition result and the speech emotion recognition result obtained by using the traditional feature alignment algorithm, the speech emotion recognition result obtained by the technical solution provided in this application has an improvement of 13.00% and 5.66% respectively, and the accuracy is higher.
[0122] When the experimental data included a second speech sample obtained from EMOVO (Italian), the speech emotion recognition result obtained by performing speech emotion recognition directly without feature alignment was 28.87%, the speech emotion recognition result obtained by performing speech emotion recognition using the traditional feature alignment algorithm was 44.35%, and the speech emotion recognition result obtained by performing speech emotion recognition using the fine-grained local feature alignment provided in this application was 45.24%. Therefore, it can be clearly seen that compared with the direct recognition result and the speech emotion recognition result obtained by using the traditional feature alignment algorithm, the speech emotion recognition result obtained by the technical solution provided in this application has an improvement of 16.37% and 0.89% respectively, and the accuracy is higher.
[0123] Based on the above description, it is clear that when the three sets of experiments performed speech emotion recognition directly without feature alignment, the accuracy was low. When the three sets of experiments used the traditional feature alignment algorithm for speech emotion recognition, the accuracy improved to varying degrees compared to the direct recognition results, but still did not exceed 50%, indicating that the traditional feature alignment algorithm has limited effectiveness in speech emotion recognition. The speech emotion recognition accuracy obtained by the technical solution provided in this application is improved compared to both the direct recognition results and the results obtained using the traditional feature alignment algorithm, demonstrating the effectiveness of the technical solution provided in this application.
[0124] For example, see Figure 4 As shown, Figure 4 This is a flowchart illustrating the voice emotion recognition method provided in an embodiment of this application. The voice emotion recognition method may include:
[0125] S401. Obtain the speech to be recognized.
[0126] S402. Input the speech to be recognized into the speech emotion recognition model to obtain the speech emotion result corresponding to the speech to be recognized. The speech emotion recognition model is the speech emotion recognition model of any of the above embodiments.
[0127] It should be noted that the specific implementation of inputting the speech to be recognized into the speech emotion recognition model to obtain the speech emotion result corresponding to the speech to be recognized is the same as described above. Figure 1 In the embodiment shown, the specific implementation of inputting the speech features corresponding to the speech sample pair into the initial speech emotion recognition model to obtain the first prediction result corresponding to the first speech sample and the second prediction result corresponding to the second speech sample in the speech sample pair is similar, and can be referred to the above-mentioned relevant description. Here, the embodiments of this application will not be repeated.
[0128] As can be seen in this embodiment, when performing voice emotion recognition, the voice to be recognized can be input into the voice emotion recognition model. Since the voice emotion recognition model can accurately recognize the voice emotions of different languages, the voice emotion recognition model can accurately identify the voice emotion result corresponding to the voice to be recognized, thereby improving the accuracy of the recognition result.
[0129] The training device and the speech emotion recognition device of the speech emotion recognition model provided in this application are described below. The training device of the speech emotion recognition model described below can be referred to in correspondence with the training method of the speech emotion recognition model described above, and the speech emotion recognition device can be referred to in correspondence with the speech emotion recognition method described above.
[0130] Figure 5This is a schematic diagram of the structure of the training device 50 for the speech emotion recognition model provided in the embodiments of this application. For example, please refer to [link to relevant documentation]. Figure 5 As shown, the training device 50 for the speech emotion recognition model may include:
[0131] The acquisition unit 501 is used to acquire multiple speech sample pairs. Each speech sample pair includes a first speech sample with a speech emotion label and a second speech sample without a speech emotion label. The first speech sample and the second speech sample belong to different languages.
[0132] The processing unit 502 is used to input the speech features corresponding to each speech sample pair into the initial speech emotion recognition model to obtain the first prediction result corresponding to the first speech sample and the second prediction result corresponding to the second speech sample in the speech sample pair.
[0133] The update unit 503 is used to update the model parameters of the initial speech emotion recognition model based on the first prediction result, the second prediction result, and the speech emotion label of each speech sample, so as to obtain the speech emotion recognition model.
[0134] For example, in an embodiment of this application, the initial speech emotion recognition model includes a feature extractor and a recognizer.
[0135] The processing unit 502 is specifically used to input the speech features corresponding to the speech sample pair into the feature extractor to obtain the first hidden layer features corresponding to the first speech sample and the second hidden layer features corresponding to the second speech sample; and to input the first hidden layer features and the second hidden layer features into the recognizer respectively to obtain the first prediction result corresponding to the first speech sample and the second prediction result corresponding to the second speech sample.
[0136] For example, in this embodiment of the application, the update unit 503 is specifically used to construct the cross-entropy loss corresponding to multiple speech sample pairs based on the first prediction result and speech emotion label corresponding to each speech sample pair; determine the similarity of multiple speech sample pairs based on the speech emotion label and second prediction result corresponding to each speech sample pair using a local feature alignment algorithm; and update the model parameters of the initial speech emotion recognition model based on the cross-entropy loss and similarity of multiple speech sample pairs.
[0137] For example, in this embodiment of the application, the updating unit 503 is specifically used to determine the first target speech sample and the second target speech sample corresponding to each speech emotion category from multiple speech sample pairs based on the speech emotion label and the second prediction result corresponding to each speech sample pair; and to determine the similarity of multiple speech sample pairs based on the first target speech sample and the second target speech sample corresponding to each speech emotion category using a local feature alignment algorithm.
[0138] For example, in this embodiment of the application, the update unit 503 is specifically used to update according to A local feature alignment algorithm is used to determine the similarity between multiple speech sample pairs. This represents the distance between multiple speech sample pairs; the distance is used to characterize similarity. c represents the number of voice emotion categories. Voice emotion category c out of 1 voice emotion category, This represents the first set consisting of the first target speech samples corresponding to the speech emotion category c. This represents the second set consisting of the second target speech samples corresponding to the speech emotion category c. Represents the first set of elements. The first target speech sample, Indicates the first The weights corresponding to the first target speech sample Indicates the first The mapping from the first target speech sample to the regenerated Hilbert space RKHS. Represents the first in the second set A second target speech sample, Indicates the first The weights corresponding to each second target speech sample Indicates the first Mapping of a second target speech sample to RKHS.
[0139] For example, in an embodiment of this application, the training device 50 for the voice emotion recognition model further includes a determining unit.
[0140] Determine the unit, specifically used according to Determine the first The weights corresponding to the first target speech samples; based on Determine the first The weights corresponding to each second target speech sample. Among them, Indicates the first The probability value of the voice emotion label corresponding to each first target voice sample; Indicates the first The probability value of the second prediction result corresponding to each second target speech sample.
[0141] For example, in an embodiment of this application, the training device 50 for the voice emotion recognition model further includes a preprocessing unit.
[0142] The preprocessing unit is specifically used to preprocess the first and second speech samples in the speech sample pair using the log-Mel spectrum algorithm to obtain the speech features corresponding to the speech sample pair.
[0143] The training device 50 for the speech emotion recognition model provided in this application embodiment can execute the technical solution of the training method for the speech emotion recognition model in any of the above embodiments. Its implementation principle and beneficial effects are similar to those of the training method for the speech emotion recognition model. Please refer to the implementation principle and beneficial effects of the training method for the speech emotion recognition model. It will not be repeated here.
[0144] Figure 6 This is a schematic diagram of the structure of the voice emotion recognition device 60 provided in the embodiments of this application. For example, please refer to [link to relevant documentation]. Figure 6 As shown, the voice emotion recognition device 60 may include:
[0145] Acquisition unit 601 is used to acquire the speech to be recognized.
[0146] The processing unit 602 is used to input the speech to be recognized into the speech emotion recognition model to obtain the speech emotion result corresponding to the speech to be recognized. The speech emotion recognition model is the speech emotion recognition model described in any of the above methods.
[0147] The voice emotion recognition device 60 provided in this application embodiment can execute the technical solution of the voice emotion recognition method in any of the above embodiments. Its implementation principle and beneficial effects are similar to those of the voice emotion recognition method. Please refer to the implementation principle and beneficial effects of the voice emotion recognition method. It will not be repeated here.
[0148] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740. The processor 710, communications interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions from the memory 730 to execute a training method for a speech emotion recognition model, or a speech emotion recognition method.
[0149] The training method for the speech emotion recognition model includes: acquiring multiple speech sample pairs, each speech sample pair including a first speech sample with a speech emotion label and a second speech sample without a speech emotion label, the first speech sample and the second speech sample belonging to different languages; for each speech sample pair, inputting the speech features corresponding to the speech sample pair into the initial speech emotion recognition model to obtain a first prediction result corresponding to the first speech sample and a second prediction result corresponding to the second speech sample in the speech sample pair; updating the model parameters of the initial speech emotion recognition model according to the first prediction result, the second prediction result and the speech emotion label corresponding to each speech sample pair to obtain the trained speech emotion recognition model.
[0150] The speech emotion recognition method includes: acquiring the speech to be recognized; inputting the speech to be recognized into the speech emotion recognition model to obtain the speech emotion result corresponding to the speech to be recognized.
[0151] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0152] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the training method of the speech emotion recognition model or the speech emotion recognition method provided by the above methods.
[0153] The training method for the speech emotion recognition model includes: acquiring multiple speech sample pairs, each speech sample pair including a first speech sample with a speech emotion label and a second speech sample without a speech emotion label, the first speech sample and the second speech sample belonging to different languages; for each speech sample pair, inputting the speech features corresponding to the speech sample pair into the initial speech emotion recognition model to obtain a first prediction result corresponding to the first speech sample and a second prediction result corresponding to the second speech sample in the speech sample pair; updating the model parameters of the initial speech emotion recognition model according to the first prediction result, the second prediction result and the speech emotion label corresponding to each speech sample pair to obtain the trained speech emotion recognition model.
[0154] The speech emotion recognition method includes: acquiring the speech to be recognized; inputting the speech to be recognized into the speech emotion recognition model to obtain the speech emotion result corresponding to the speech to be recognized.
[0155] In another aspect, this application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a training method for the speech emotion recognition model provided by the above methods, or a speech emotion recognition method.
[0156] The training method for the speech emotion recognition model includes: acquiring multiple speech sample pairs, each speech sample pair including a first speech sample with a speech emotion label and a second speech sample without a speech emotion label, the first speech sample and the second speech sample belonging to different languages; for each speech sample pair, inputting the speech features corresponding to the speech sample pair into the initial speech emotion recognition model to obtain a first prediction result corresponding to the first speech sample and a second prediction result corresponding to the second speech sample in the speech sample pair; updating the model parameters of the initial speech emotion recognition model according to the first prediction result, the second prediction result and the speech emotion label corresponding to each speech sample pair to obtain the trained speech emotion recognition model.
[0157] The speech emotion recognition method includes: acquiring the speech to be recognized; inputting the speech to be recognized into the speech emotion recognition model to obtain the speech emotion result corresponding to the speech to be recognized.
[0158] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0159] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A training method for a speech emotion recognition model, characterized in that, include: Multiple voice sample pairs are acquired. Each voice sample pair includes a first voice sample with a voice emotion label and a second voice sample without a voice emotion label. The first voice sample and the second voice sample belong to different languages. For each speech sample pair, the speech features corresponding to the speech sample pair are input into the initial speech emotion recognition model to obtain the first prediction result corresponding to the first speech sample and the second prediction result corresponding to the second speech sample in the speech sample pair. Based on the first prediction result, the second prediction result, and the voice emotion label corresponding to each voice sample, the model parameters of the initial voice emotion recognition model are updated to obtain the voice emotion recognition model. The step of updating the model parameters of the initial speech emotion recognition model based on the first prediction result, the second prediction result, and the speech emotion label for each speech sample to obtain the speech emotion recognition model includes: Based on the first prediction result corresponding to each speech sample pair and the speech emotion label, construct the cross-entropy loss corresponding to the multiple speech sample pairs; Based on the voice emotion label corresponding to each voice sample pair and the second prediction result, a local feature alignment algorithm is used to determine the similarity between the multiple voice sample pairs; the step of determining the similarity between the multiple voice sample pairs based on the voice emotion label corresponding to each voice sample pair and the second prediction result includes: determining a first target voice sample and a second target voice sample corresponding to each voice emotion category from the multiple voice sample pairs based on the voice emotion label corresponding to each voice sample pair and the second prediction result; and determining the similarity between the multiple voice sample pairs based on the first target voice sample and the second target voice sample corresponding to each voice emotion category using a local feature alignment algorithm. Based on the cross-entropy loss and similarity of the multiple speech sample pairs, the model parameters of the initial speech emotion recognition model are updated to obtain the speech emotion recognition model.
2. The training method for the speech emotion recognition model according to claim 1, characterized in that, The initial speech emotion recognition model includes a feature extractor and a recognizer. The step of inputting the speech features corresponding to the speech sample pair into the initial speech emotion recognition model to obtain a first prediction result corresponding to the first speech sample and a second prediction result corresponding to the second speech sample in the speech sample pair includes: The speech features corresponding to the speech sample pair are input into the feature extractor to obtain the first hidden layer features corresponding to the first speech sample and the second hidden layer features corresponding to the second speech sample. The first hidden layer features and the second hidden layer features are respectively input into the recognizer to obtain the first prediction result corresponding to the first speech sample and the second prediction result corresponding to the second speech sample.
3. The training method for the speech emotion recognition model according to claim 1, characterized in that, The step of determining the similarity of the multiple speech sample pairs based on the first target speech sample and the second target speech sample corresponding to each speech emotion category using a local feature alignment algorithm includes: according to The similarity of the multiple speech sample pairs is determined by using a local feature alignment algorithm. in, This represents the distance between multiple pairs of speech samples, where the distance is used to characterize the similarity. c represents the number of voice emotion categories. Voice emotion category c out of 1 voice emotion category, This represents the first set consisting of the first target speech samples corresponding to the speech emotion category c. This represents the second set consisting of the second target speech samples corresponding to the speech emotion category c. Represents the first set of... The first target speech sample, Indicates the first The weights corresponding to the first target speech sample Indicates the first The mapping from the first target speech sample to the regenerated Hilbert space RKHS. Represents the first in the second set A second target speech sample, Indicates the first The weights corresponding to each second target speech sample Indicates the first Mapping of a second target speech sample to the regenerated Hilbert space RKHS.
4. The training method for the speech emotion recognition model according to claim 3, characterized in that, The method further includes: according to Determine the first The weights corresponding to the first target speech samples; according to Determine the first The weights corresponding to each second target speech sample; in, Indicates the first The probability value of the voice emotion label corresponding to each first target voice sample; Indicates the first The probability value of the second prediction result corresponding to each second target speech sample.
5. A voice emotion recognition method, characterized in that, include: Acquire the speech to be recognized; The speech to be recognized is input into the speech emotion recognition model to obtain the speech emotion result corresponding to the speech to be recognized. The speech emotion recognition model is the speech emotion recognition model according to any one of claims 1-4 above.
6. A training device for a speech emotion recognition model, characterized in that, include: The acquisition unit is used to acquire multiple speech sample pairs, each speech sample pair including a first speech sample with a speech emotion label and a second speech sample without a speech emotion label, wherein the first speech sample and the second speech sample belong to different languages; The processing unit is configured to input the speech features corresponding to each speech sample pair into an initial speech emotion recognition model to obtain a first prediction result corresponding to the first speech sample and a second prediction result corresponding to the second speech sample in the speech sample pair. An update unit is configured to update the model parameters of the initial speech emotion recognition model based on the first prediction result, the second prediction result, and the speech emotion label corresponding to each speech sample pair, to obtain a speech emotion recognition model. Specifically, the update unit is configured to construct a cross-entropy loss corresponding to each speech sample pair based on the first prediction result and the speech emotion label corresponding to each speech sample pair; and to determine the similarity of the multiple speech sample pairs using a local feature alignment algorithm based on the speech emotion label and the second prediction result corresponding to each speech sample pair. The step of determining the similarity of the multiple speech sample pairs using a local feature alignment algorithm based on the speech emotion label and the second prediction result corresponding to each speech sample pair includes: determining a first target speech sample and a second target speech sample corresponding to each speech emotion category from the multiple speech sample pairs based on the speech emotion label and the second prediction result corresponding to each speech sample pair. Based on the first target speech sample and the second target speech sample corresponding to each speech emotion category, a local feature alignment algorithm is used to determine the similarity of the multiple speech sample pairs; based on the cross-entropy loss and the similarity of the multiple speech sample pairs, the model parameters of the initial speech emotion recognition model are updated to obtain the speech emotion recognition model.
7. A voice emotion recognition device, characterized in that, include: The acquisition unit is used to acquire the speech to be recognized; The processing unit is configured to input the speech to be recognized into a speech emotion recognition model to obtain the speech emotion result corresponding to the speech to be recognized, wherein the speech emotion recognition model is the speech emotion recognition model described in any one of claims 1-4.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the training method for the speech emotion recognition model as described in any one of claims 1 to 4, or the speech emotion recognition method as described in claim 5.