Training method of emotion recognition model, emotion recognition method, device, electronic device, storage medium and program product

By introducing speaker encoder and decoder into the speech emotion recognition model and using multimodal training data set for staged joint training, the problem of low accuracy in recognition among different speech emotion recognition systems is solved, and the generalization and robustness of the model is improved.

CN118245803BActive Publication Date: 2025-05-16MOORE THREADS TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410445146.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-12
Publication Date
2025-05-16
Estimated Expiration
2044-04-12

AI Technical Summary

Technical Problem

When the existing speech emotion recognition system uses speakers different from the training set in the test stage, incorrect emotion recognition results often occur, and because emotion recognition has a certain subjectivity, the recognition accuracy is low.

Method used

A multimodal training data set is used to train a multimodal emotion recognition model including text encoder, audio encoder, fusion module and multimodal decoder, and a speaker encoder and speaker decoder are added during the model training process. Through a phased joint training method, speaker adaptive information is injected.

Benefits of technology

The generalization and robustness of the speech emotion recognition model is improved, allowing the model to more accurately identify the emotions of different speakers and reduce the dependence on specific speaker data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118245803B_ABST
    Figure CN118245803B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a training method for an emotion recognition model, an emotion recognition method, an apparatus, an electronic device, a storage medium and a program product. The method comprises: obtaining a multimodal training data set, wherein the multimodal training data set comprises multiple multimodal training data; using the multimodal training data set to train a first multimodal emotion recognition model, wherein the first multimodal emotion recognition model comprises a text encoder, an audio encoder, a fusion module and a multimodal decoder; in response to the first multimodal emotion recognition model being trained to meet a first preset training condition, adding a speaker encoder and a speaker decoder to the first multimodal emotion recognition model to obtain a second multimodal emotion recognition model; using the multimodal training data set to train the second multimodal emotion recognition model until the second preset training condition is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a training method for an emotion recognition model, an emotion recognition method, a training device for an emotion recognition model, an emotion recognition device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Speech is one of the main communication media in daily life. It not only conveys thoughts but also expresses the speaker's emotional state. Speech emotion recognition can detect the user's mental health status and provide corresponding feedback and responses based on the user's emotional changes in different scenarios (for example, digital humans, customer service, etc.). In addition, speech emotion recognition can also help parents understand their children's mental health status.

[0003] However, in speech emotion recognition systems, labeling emotional data is a challenging task, which makes speech emotion recognition models more susceptible to the emotions of speakers in the training set. In other words, if we use speakers different from those in the training set during the test phase, the system often produces incorrect emotion recognition results. In addition, speech emotion recognition is subjective because different users may have different emotional understandings of the same sentence. That is, some people may think that a sentence carries a special or strong emotion, while others may think that it expresses a calm emotion.

[0004] How to improve the accuracy of emotion recognition is a technical problem that needs to be solved urgently. Summary of the invention

[0005] The present invention provides a technical solution for emotion recognition.

[0006] According to one aspect of the present disclosure, a method for training an emotion recognition model is provided, comprising:

[0007] Obtaining a multimodal training data set, wherein the multimodal training data set includes multiple multimodal training data;

[0008] Using the multimodal training data set to train a first multimodal emotion recognition model, wherein the first multimodal emotion recognition model includes a text encoder, an audio encoder, a fusion module and a multimodal decoder;

[0009] In response to the first multimodal emotion recognition model being trained to satisfy a first preset training condition, a speaker encoder and a speaker decoder are added to the first multimodal emotion recognition model to obtain a second multimodal emotion recognition model;

[0010] The second multimodal emotion recognition model is trained using the multimodal training data set until a second preset training condition is met.

[0011] In a possible implementation, during the training of the second multimodal emotion recognition model, the parameters of the text encoder and the audio encoder are fixed, and the parameters of the speaker encoder, the speaker decoder, the fusion module and the multimodal decoder are updated.

[0012] In a possible implementation, any item of multimodal training data includes a sentiment label, training text, and training audio;

[0013] The adopting the multimodal training data set to train a first multimodal emotion recognition model comprises:

[0014] For any item of multimodal training data in the multimodal training data set, extracting a first text feature vector corresponding to a training text in the multimodal training data by the text encoder, and extracting a first audio feature vector corresponding to a training audio in the multimodal training data by the audio encoder;

[0015] fusing the first text feature vector and the first audio feature vector through the fusion module to obtain a first fused feature vector corresponding to the multimodal training data;

[0016] Inputting the first fused feature vector into the multimodal decoder, and outputting a first emotion category prediction result corresponding to the multimodal training data through the multimodal decoder;

[0017] The first multimodal emotion recognition model is trained according to the first emotion category prediction result and the emotion label in the multimodal training data.

[0018] In a possible implementation, the first multimodal emotion recognition model further includes an adaptation module;

[0019] The fusing the first text feature vector and the first audio feature vector by the fusion module to obtain a first fused feature vector corresponding to the multimodal training data includes:

[0020] Performing dimensionality reduction on the first audio feature vector by the adaptation module to obtain a first dimensionality reduction feature vector corresponding to the first audio feature vector;

[0021] The first text feature vector and the first dimension reduction feature vector are fused through the fusion module to obtain a first fused feature vector corresponding to the multimodal training data.

[0022] In a possible implementation, fusing the first text feature vector and the first dimension reduction feature vector by the fusion module to obtain a first fused feature vector corresponding to the multimodal training data includes:

[0023] Processing the first dimensionality reduction feature vector by the fusion module based on a cross attention mechanism to obtain a second dimensionality reduction feature vector corresponding to the first audio feature vector;

[0024] The first text feature vector and the second dimension reduction feature vector are concatenated or added by the fusion module to obtain a first fused feature vector corresponding to the multimodal training data.

[0025] In one possible implementation, any multimodal training data also includes speaker labels;

[0026] The adopting the multimodal training data set to train the second multimodal emotion recognition model includes:

[0027] For any item of multimodal training data in the multimodal training data set, extracting a second text feature vector corresponding to the training text in the multimodal training data by the text encoder, extracting a second audio feature vector corresponding to the training audio in the multimodal training data by the audio encoder, and extracting a first speaker feature vector corresponding to the training audio by the speaker encoder;

[0028] fusing the first text feature vector, the first audio feature vector and the first speaker feature vector through the fusion module to obtain a second fused feature vector corresponding to the multimodal training data;

[0029] Inputting the second fused feature vector into the multimodal decoder, and outputting a second emotion category prediction result corresponding to the multimodal training data through the multimodal decoder;

[0030] Inputting the first speaker feature vector into the speaker decoder, and outputting a second speaker prediction result corresponding to the multimodal training data through the speaker decoder;

[0031] The second multimodal emotion recognition model is trained according to difference information between the second emotion category prediction result and the emotion label in the multimodal training data, and difference information between the second speaker prediction result and the speaker label in the multimodal training data.

[0032] In a possible implementation, before adopting the multimodal training dataset to train the first multimodal emotion recognition model, the method further includes:

[0033] For any one of the multimodal training data in the multimodal training data set, input the training audio in the multimodal training data into a speaker recognition model, and output a first speaker prediction result corresponding to the training audio through the speaker recognition model; wherein the speaker recognition model includes the speaker encoder, the speaker decoder and the speaker classifier;

[0034] The speaker recognition model is trained according to the first speaker prediction result and the speaker labels in the multimodal training data.

[0035] In a possible implementation, before adopting the multimodal training dataset to train the first multimodal emotion recognition model, the method further includes:

[0036] The audio encoder is trained by using training audio in the multimodal training data in the multimodal training data set.

[0037] In a possible implementation, before adopting the multimodal training dataset to train the first multimodal emotion recognition model, the method further includes:

[0038] The text encoder is trained by using training text in the multimodal training data in the multimodal training data set.

[0039] In a possible implementation, the method further includes:

[0040] In response to the second multimodal emotion recognition model being trained to satisfy the second preset training condition, the speaker decoder in the second multimodal emotion recognition model is deleted to obtain a third multimodal emotion recognition model.

[0041] According to one aspect of the present disclosure, there is provided an emotion recognition method, comprising:

[0042] Acquire a third multimodal emotion recognition model obtained by the emotion recognition model training method;

[0043] Obtaining multimodal data to be processed, wherein the multimodal data to be processed includes audio to be processed and text to be processed;

[0044] The audio to be processed and the text to be processed are input into the third multimodal emotion recognition model to obtain an emotion category prediction result corresponding to the multimodal data to be processed.

[0045] In a possible implementation, obtaining the multimodal data to be processed includes:

[0046] Get the audio to be processed;

[0047] Perform speech recognition on the audio to be processed to obtain text to be processed.

[0048] According to one aspect of the present disclosure, a training device for an emotion recognition model is provided, characterized in that it includes:

[0049] A first acquisition module is used to obtain a multimodal training data set, wherein the multimodal training data set includes multiple multimodal training data;

[0050] A first training module, configured to train a first multimodal emotion recognition model using the multimodal training data set, wherein the first multimodal emotion recognition model comprises a text encoder, an audio encoder, a fusion module and a multimodal decoder;

[0051] An adding module, configured to add a speaker encoder and a speaker decoder to the first multimodal emotion recognition model in response to the first multimodal emotion recognition model being trained to meet a first preset training condition, so as to obtain a second multimodal emotion recognition model;

[0052] The second training module is used to adopt the multimodal training data set to train the second multimodal emotion recognition model until a second preset training condition is met.

[0053] In a possible implementation, during the training of the second multimodal emotion recognition model, the parameters of the text encoder and the audio encoder are fixed, and the parameters of the speaker encoder, the speaker decoder, the fusion module and the multimodal decoder are updated.

[0054] In a possible implementation, any item of multimodal training data includes a sentiment label, training text, and training audio;

[0055] The first training module is used to:

[0056] For any item of multimodal training data in the multimodal training data set, extracting a first text feature vector corresponding to a training text in the multimodal training data by the text encoder, and extracting a first audio feature vector corresponding to a training audio in the multimodal training data by the audio encoder;

[0057] fusing the first text feature vector and the first audio feature vector through the fusion module to obtain a first fused feature vector corresponding to the multimodal training data;

[0058] Inputting the first fused feature vector into the multimodal decoder, and outputting a first emotion category prediction result corresponding to the multimodal training data through the multimodal decoder;

[0059] The first multimodal emotion recognition model is trained according to the first emotion category prediction result and the emotion label in the multimodal training data.

[0060] In a possible implementation, the first multimodal emotion recognition model further includes an adaptation module;

[0061] The first training module is used to:

[0062] Performing dimensionality reduction on the first audio feature vector by the adaptation module to obtain a first dimensionality reduction feature vector corresponding to the first audio feature vector;

[0063] The first text feature vector and the first dimension reduction feature vector are fused through the fusion module to obtain a first fused feature vector corresponding to the multimodal training data.

[0064] In a possible implementation, the first training module is used to:

[0065] Processing the first dimensionality reduction feature vector by the fusion module based on a cross attention mechanism to obtain a second dimensionality reduction feature vector corresponding to the first audio feature vector;

[0066] The first text feature vector and the second dimension reduction feature vector are concatenated or added by the fusion module to obtain a first fused feature vector corresponding to the multimodal training data.

[0067] In one possible implementation, any multimodal training data also includes speaker labels;

[0068] The second training module is used to:

[0069] For any item of multimodal training data in the multimodal training data set, extracting a second text feature vector corresponding to the training text in the multimodal training data by the text encoder, extracting a second audio feature vector corresponding to the training audio in the multimodal training data by the audio encoder, and extracting a first speaker feature vector corresponding to the training audio by the speaker encoder;

[0070] fusing the first text feature vector, the first audio feature vector and the first speaker feature vector through the fusion module to obtain a second fused feature vector corresponding to the multimodal training data;

[0071] Inputting the second fused feature vector into the multimodal decoder, and outputting a second emotion category prediction result corresponding to the multimodal training data through the multimodal decoder;

[0072] Inputting the first speaker feature vector into the speaker decoder, and outputting a second speaker prediction result corresponding to the multimodal training data through the speaker decoder;

[0073] The second multimodal emotion recognition model is trained according to difference information between the second emotion category prediction result and the emotion label in the multimodal training data, and difference information between the second speaker prediction result and the speaker label in the multimodal training data.

[0074] In a possible implementation manner, the device further includes:

[0075] A speaker prediction module, for any one of the multimodal training data in the multimodal training data set, inputting the training audio in the multimodal training data into a speaker recognition model, and outputting a first speaker prediction result corresponding to the training audio through the speaker recognition model; wherein the speaker recognition model includes the speaker encoder, the speaker decoder and the speaker classifier;

[0076] The third training module is used to train the speaker recognition model according to the first speaker prediction result and the speaker label in the multimodal training data.

[0077] In a possible implementation manner, the device further includes:

[0078] The fourth training module is used to use the training audio in the multimodal training data in the multimodal training data set to train the audio encoder.

[0079] In a possible implementation manner, the device further includes:

[0080] The fifth training module is used to use the training text in the multimodal training data in the multimodal training data set to train the text encoder.

[0081] In a possible implementation manner, the device further includes:

[0082] A deleting module is used to delete the speaker decoder in the second multimodal emotion recognition model in response to the second multimodal emotion recognition model being trained to meet the second preset training condition, so as to obtain a third multimodal emotion recognition model.

[0083] According to one aspect of the present disclosure, there is provided an emotion recognition device, comprising:

[0084] An acquisition module, used to acquire a third multimodal emotion recognition model obtained by the training method of the emotion recognition model;

[0085] A second obtaining module is used to obtain multimodal data to be processed, wherein the multimodal data to be processed includes audio to be processed and text to be processed;

[0086] The emotion recognition module is used to input the audio to be processed and the text to be processed into the third multimodal emotion recognition model to obtain the emotion category prediction result corresponding to the multimodal data to be processed.

[0087] In a possible implementation manner, the second obtaining module is used to:

[0088] Get the audio to be processed;

[0089] Perform speech recognition on the audio to be processed to obtain text to be processed.

[0090] According to one aspect of the present disclosure, an electronic device is provided, comprising: one or more processors; a memory for storing executable instructions; wherein the one or more processors are configured to call the executable instructions stored in the memory to execute the above method.

[0091] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and the computer program instructions implement the above method when executed by a processor.

[0092] According to one aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes the above method.

[0093] In an embodiment of the present disclosure, a multimodal training data set is obtained, wherein the multimodal training data set includes multiple multimodal training data, and the multimodal training data set is used to train a first multimodal emotion recognition model, wherein the first multimodal emotion recognition model includes a text encoder, an audio encoder, a fusion module and a multimodal decoder. In response to the first multimodal emotion recognition model being trained to meet a first preset training condition, a speaker encoder and a speaker decoder are added to the first multimodal emotion recognition model to obtain a second multimodal emotion recognition model, and the multimodal training data set is used to train the second multimodal emotion recognition model until the second preset training condition is met. Thus, a staged joint training method is proposed, which uses multimodal data to improve the effect of the emotion recognition model, and injects speaker adaptive information, so that the emotion recognition model has stronger generalization and is more robust to different speakers. The multimodal emotion recognition model trained in the embodiment of the present disclosure can realize speaker adaptive multimodal emotion recognition.

[0094] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure.

[0095] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and are used to illustrate the technical solutions of the present disclosure together with the specification.

[0097] Figure 1 A flowchart of a method for training an emotion recognition model provided by an embodiment of the present disclosure is shown.

[0098] Figure 2 A schematic diagram of a speaker recognition model provided by an embodiment of the present disclosure is shown.

[0099] Figure 3 A schematic diagram of a first multimodal emotion recognition model provided by an embodiment of the present disclosure is shown.

[0100] Figure 4 A schematic diagram of a second multimodal emotion recognition model provided by an embodiment of the present disclosure is shown.

[0101] Figure 5 A schematic diagram of a third multimodal emotion recognition model provided by an embodiment of the present disclosure is shown.

[0102] Figure 6 A block diagram of a training device for an emotion recognition model provided by an embodiment of the present disclosure is shown.

[0103] Figure 7 A block diagram of an electronic device 1900 provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0104] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0105] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0106] The term "and / or" herein is only a description of the association relationship of the associated objects, indicating that there may be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the term "at least one" herein represents any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set consisting of A, B, and C.

[0107] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present disclosure.

[0108] When training speech emotion recognition models, it is crucial to introduce speaker adaptation, especially in practical applications. The introduction of speaker adaptation helps to improve the accuracy of speech emotion recognition and enables the speech emotion recognition model to better adapt to the emotional expressions of different speakers. Due to the limited speech emotion data, the emotion data for a specific speaker may not be sufficient to provide sufficient robustness and generalization for the speech emotion recognition model. Therefore, if there is not enough data to train the speech emotion recognition model to adapt to a specific speaker, the effect of speaker adaptation may be limited.

[0109] In the speech emotion recognition system in the related art, speaker adaptation is mainly achieved through the following two schemes.

[0110] The first solution is to use the voiceprint model to pre-acquire the voiceprint of each speaker when training the original speech emotion recognition model, and then train different speech emotion recognition models for different speakers.

[0111] The second solution is to constrain different speakers during the training process of the speech emotion recognition model, so that the speech emotion recognition model can learn speaker-independent emotional features.

[0112] Among them, the disadvantage of the first solution is that it is necessary to model the voiceprint features to emotion classification at the same time in the initialization stage of training, which cannot be quickly iterated for different users. In addition, different speech emotion recognition models need to be trained for different speakers, which is costly.

[0113] Although the second solution can effectively improve the robustness of the speech emotion recognition system for emotion recognition of different speakers, it cannot quickly adapt to a certain user. For example, if the emotional expression of a certain user is not very obvious, it may not be possible to perform emotion recognition well.

[0114] In order to solve the technical problems similar to those described above, the embodiment of the present disclosure provides a method for training an emotion recognition model, by obtaining a multimodal training data set, wherein the multimodal training data set includes multiple multimodal training data, using the multimodal training data set to train a first multimodal emotion recognition model, wherein the first multimodal emotion recognition model includes a text encoder, an audio encoder, a fusion module and a multimodal decoder, in response to the first multimodal emotion recognition model being trained to meet a first preset training condition, a speaker encoder and a speaker decoder are added to the first multimodal emotion recognition model to obtain a second multimodal emotion recognition model, and the multimodal training data set is used to train the second multimodal emotion recognition model until the second preset training condition is met, thereby proposing a staged joint training method, using multimodal data to improve the effect of the emotion recognition model, and injecting speaker adaptive information, so that the emotion recognition model has stronger generalization and is more robust to different speakers. The multimodal emotion recognition model trained by the embodiment of the present disclosure can realize speaker adaptive multimodal emotion recognition.

[0115] The following is a detailed description of the training method of the emotion recognition model provided by the embodiment of the present disclosure in conjunction with the accompanying drawings.

[0116] Figure 1 A flow chart of a method for training an emotion recognition model provided by an embodiment of the present disclosure is shown. In one possible implementation, the executor of the method for training an emotion recognition model may be a training device for the emotion recognition model. For example, the method for training an emotion recognition model may be executed by a terminal device or a server or other electronic device. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device or a wearable device. In some possible implementations, the method for training an emotion recognition model may be implemented by a processor calling computer-readable instructions stored in a memory. For example Figure 1 As shown, the training method of the emotion recognition model includes steps S11 to S14.

[0117] In step S11, a multimodal training data set is obtained, wherein the multimodal training data set includes multiple items of multimodal training data.

[0118] In step S12, the multimodal training data set is used to train a first multimodal emotion recognition model, wherein the first multimodal emotion recognition model includes a text encoder, an audio encoder, a fusion module and a multimodal decoder.

[0119] In step S13, in response to the first multimodal emotion recognition model being trained to meet a first preset training condition, a speaker encoder and a speaker decoder are added to the first multimodal emotion recognition model to obtain a second multimodal emotion recognition model.

[0120] In step S14, the multimodal training data set is used to train the second multimodal emotion recognition model until a second preset training condition is met.

[0121] In an embodiment of the present disclosure, the multimodal training data may represent data used to train a multimodal emotion recognition model. Any item of the multimodal training data may include an emotion label, training text, and training audio. In a possible implementation, any item of the multimodal training data may also include a speaker label.

[0122] In one possible implementation, open source speech emotion data can be collected, or speech emotion data can be crawled or annotated based on the web. Among them, the speech emotion data can include audio and emotion tags. In some examples, the speech emotion data can also include speaker tags. The speech emotion data can be screened to select training audios of multiple emotion categories corresponding to multiple speakers. Speech recognition can be performed on the training audio to obtain training text corresponding to the training audio. Multimodal training data can be obtained based on the training audio, training text, emotion tags and speaker tags. Each item of multimodal training data can form a multimodal training data set.

[0123] In a possible implementation, before adopting the multimodal training data set to train the first multimodal emotion recognition model, the method further includes: for any multimodal training data in the multimodal training data set, inputting the training audio in the multimodal training data into a speaker recognition model, and outputting a first speaker prediction result corresponding to the training audio through the speaker recognition model; wherein the speaker recognition model includes the speaker encoder, the speaker decoder and a speaker classifier; and training the speaker recognition model according to the first speaker prediction result and the speaker label in the multimodal training data.

[0124] Figure 2 FIG. 1 is a schematic diagram showing a speaker recognition model provided by an embodiment of the present disclosure. Figure 2As shown, the speaker recognition model may include a speaker encoder, a speaker decoder and a classifier. The speaker encoder may represent an encoder in the speaker recognition model, and the speaker decoder may represent a decoder in the speaker recognition model. The speaker recognition model may adopt different model structures.

[0125] The output of the speaker encoder can be recorded as a speaker feature vector. In some application scenarios, the speaker encoder can also be called a voiceprint embedding extractor, a speaker feature vector extractor, etc., which is not limited here. After the speaker recognition model training is completed, the speaker encoder can be selected as the module for extracting the speaker feature vector. In an example, the pre-trained speaker encoder can be represented by S_encoder.

[0126] In a possible implementation, before using the multimodal training data set to train the first multimodal emotion recognition model, the method further includes: using training audio in the multimodal training data in the multimodal training data set to train the audio encoder.

[0127] In this implementation, the training audio in the multimodal training data in the multimodal training data set can be used to pre-train an unsupervised audio coding model (such as a Wav2vec2.0 model). The pre-trained audio coding model can be used as an audio encoder. In some application scenarios, the audio encoder can also be called an audio embedding extractor, an audio feature vector extractor, etc., which is not limited here.

[0128] In this implementation, during the pre-training of the unsupervised audio coding model, audio other than the training audio in the multimodal training data set may also be used to improve the pre-training effect of the audio coding model.

[0129] In an example, the audio encoder obtained through pre-training can be represented by W_encoder.

[0130] In a possible implementation, before using the multimodal training data set to train the first multimodal emotion recognition model, the method further includes: using training text in the multimodal training data in the multimodal training data set to train the text encoder.

[0131] In this implementation, the training text in the multimodal training data in the multimodal training data set can be used to pre-train a language model (such as a model such as BERT). The pre-trained language model can be used as a text encoder. In some application scenarios, the text encoder can also be called a text embedding extractor, a text feature vector extractor, etc., which is not limited here. In this implementation, in the process of pre-training the language model, a part of the text can be masked, and then the masked part can be predicted.

[0132] In this implementation, in the process of pre-training the language model, texts other than the training texts in the multimodal training data set may also be used, which is not limited here. For example, a large amount of texts crawled from the web may be used to pre-train the language model.

[0133] In this implementation, the output of the classifier of the language model can be used as a text feature vector. In one example, the text encoder obtained by pre-training can be represented by T_encoder.

[0134] In the disclosed embodiment, after the speaker encoder, audio encoder and text encoder are obtained through pre-training, the joint training phase may be entered.

[0135] In the joint training stage, first, a multimodal training data set can be used to train a first multimodal emotion recognition model, wherein the first multimodal emotion recognition model includes a text encoder, an audio encoder, a fusion module and a multimodal decoder. The text encoder can be a text encoder T_encoder obtained through pre-training, and the audio encoder can be an audio encoder W_encoder obtained through pre-training. The multimodal decoder can represent a decoder in the first multimodal emotion recognition model. The input of the first multimodal emotion recognition model can include training audio and training text.

[0136] In a possible implementation, any item of multimodal training data includes an emotion label, training text, and training audio; the use of the multimodal training data set to train a first multimodal emotion recognition model includes: for any item of multimodal training data in the multimodal training data set, extracting a first text feature vector corresponding to the training text in the multimodal training data through the text encoder, and extracting a first audio feature vector corresponding to the training audio in the multimodal training data through the audio encoder; fusing the first text feature vector and the first audio feature vector through the fusion module to obtain a first fused feature vector corresponding to the multimodal training data; inputting the first fused feature vector into the multimodal decoder, and outputting a first emotion category prediction result corresponding to the multimodal training data through the multimodal decoder; and training the first multimodal emotion recognition model according to the first emotion category prediction result and the emotion label in the multimodal training data.

[0137] In this implementation, the text encoder in the first multimodal emotion recognition model can be used to extract the first text feature vector corresponding to the training text, the audio encoder in the first multimodal emotion recognition model can be used to extract the first audio feature vector corresponding to the training audio, and the fusion module in the first multimodal emotion recognition model can be used to fuse the first text feature vector and the first audio feature vector, and input the fused first fused feature vector into the multimodal decoder.

[0138] The multimodal decoder can output a first emotion category prediction result corresponding to the multimodal training data. The value of the first loss function corresponding to the first multimodal emotion recognition model can be determined based on the difference information between the first emotion category prediction result and the emotion label in the multimodal training data, and the parameters of the first multimodal emotion recognition model can be updated based on the value of the first loss function.

[0139] In a possible implementation, the first multimodal emotion recognition model also includes an adaptation module; the fusing the first text feature vector and the first audio feature vector through the fusion module to obtain a first fused feature vector corresponding to the multimodal training data includes: reducing the dimension of the first audio feature vector through the adaptation module to obtain a first reduced dimension feature vector corresponding to the first audio feature vector; fusing the first text feature vector and the first reduced dimension feature vector through the fusion module to obtain a first fused feature vector corresponding to the multimodal training data.

[0140] Figure 3 FIG. 2 is a schematic diagram of a first multimodal emotion recognition model provided by an embodiment of the present disclosure. Figure 3As shown, the first multimodal emotion recognition model may include a text encoder, an audio encoder, an adaptation module, a fusion module and a multimodal decoder.

[0141] The first audio feature vector may be a multi-dimensional feature vector, and the dimension of the first audio feature vector may be related to the number of audio frames. The first text feature vector may be a one-dimensional feature vector.

[0142] In one example, the first audio feature vector may be reduced from 16 dimensions to 4 dimensions through an adaptation module, and then reduced from 4 dimensions to 1 dimension through a fusion module.

[0143] In this implementation, the dimension of the first audio feature vector is reduced by introducing an adaptation module, thereby facilitating subsequent feature fusion.

[0144] In a possible implementation, the first text feature vector and the first reduced dimensionality feature vector are fused through the fusion module to obtain a first fused feature vector corresponding to the multimodal training data, including: processing the first reduced dimensionality feature vector based on a cross-attention mechanism through the fusion module to obtain a second reduced dimensionality feature vector corresponding to the first audio feature vector; and concatenating or adding the first text feature vector and the second reduced dimensionality feature vector through the fusion module to obtain a first fused feature vector corresponding to the multimodal training data.

[0145] In this implementation, the fusion module may process the first reduced dimension feature vector based on the cross attention mechanism to obtain a second reduced dimension feature vector. As an example of this implementation, the dimension of the second reduced dimension feature vector may be smaller than the dimension of the first reduced dimension feature vector. For example, the first reduced dimension feature vector may be 4-dimensional, and the second reduced dimension feature vector may be 1-dimensional.

[0146] In the disclosed embodiment, in response to the first multimodal emotion recognition model being trained to meet the first preset training condition, a speaker encoder and a speaker decoder can be added to the first multimodal emotion recognition model to obtain a second multimodal emotion recognition model. The first preset training condition may be that the first multimodal emotion recognition model is trained to a preset number of iterations, the parameters of the first multimodal emotion recognition model reach a preset convergence standard, the performance of the first multimodal emotion recognition model in the validation set reaches a preset requirement, etc., which are not limited here.

[0147] Figure 4 FIG. 2 is a schematic diagram of a second multimodal emotion recognition model provided by an embodiment of the present disclosure. Figure 4As shown, the second multimodal emotion recognition model may include a text encoder, an audio encoder, an adaptation module, a speaker encoder, a speaker decoder, a fusion module and a multimodal decoder. The speaker encoder and the speaker decoder may be the speaker encoder and the speaker decoder in the speaker recognition model obtained by pre-training. The input of the text encoder may be a training text, and the input of the audio encoder and the speaker encoder may be a training audio.

[0148] In a possible implementation, in the process of training the second multimodal emotion recognition model, the parameters of the text encoder and the audio encoder are fixed, and the parameters of the speaker encoder, the speaker decoder, the fusion module, and the multimodal decoder are updated. In this implementation, in the process of training the second multimodal emotion recognition model, the parameters of the text encoder, the audio encoder, and the adaptation module can be fixed, and only the parameters of the speaker encoder, the speaker decoder, the fusion module, and the multimodal decoder are updated. In this implementation, by fixing the parameters of the text encoder and the audio encoder and updating the parameters of the speaker encoder, the speaker decoder, the fusion module, and the multimodal decoder during the training of the second multimodal emotion recognition model, it is helpful to inject speaker information into multimodal emotion recognition.

[0149] In a possible implementation, any item of multimodal training data also includes a speaker label; the multimodal training data set is used to train the second multimodal emotion recognition model, including: for any item of multimodal training data in the multimodal training data set, a second text feature vector corresponding to the training text in the multimodal training data is extracted by the text encoder, a second audio feature vector corresponding to the training audio in the multimodal training data is extracted by the audio encoder, and a first speaker feature vector corresponding to the training audio is extracted by the speaker encoder; the first text feature vector, the first audio feature vector and the first speaker feature vector are fused by the fusion module. The feature vectors are fused to obtain a second fused feature vector corresponding to the multimodal training data; the second fused feature vector is input into the multimodal decoder, and the second emotion category prediction result corresponding to the multimodal training data is output through the multimodal decoder; the first speaker feature vector is input into the speaker decoder, and the second speaker prediction result corresponding to the multimodal training data is output through the speaker decoder; the second multimodal emotion recognition model is trained according to the difference information between the second emotion category prediction result and the emotion label in the multimodal training data, and the difference information between the second speaker prediction result and the speaker label in the multimodal training data.

[0150] In this implementation, the loss function corresponding to the second multimodal emotion recognition model can be called a second loss function, and the second loss function can include a first sub-loss function and a second sub-loss function. The value of the first sub-loss function can be determined based on the difference information between the second emotion category prediction result and the emotion label in the multimodal training data, and the value of the second sub-loss function can be determined based on the difference information between the second speaker prediction result and the speaker label in the multimodal training data. The weighted sum of the value of the first sub-loss function and the value of the second sub-loss function can be calculated to obtain the value of the second loss function.

[0151] In the disclosed embodiment, the multimodal training data set may be used to train the second multimodal emotion recognition model until the second preset training condition is met. The second preset training condition may be that the second multimodal emotion recognition model is trained to a preset number of iterations, the parameters of the second multimodal emotion recognition model reach a preset convergence standard, the performance of the second multimodal emotion recognition model in the validation set reaches a preset requirement, etc., which are not limited here.

[0152] In a possible implementation, the method further includes: in response to the second multimodal emotion recognition model being trained to satisfy the second preset training condition, deleting the speaker decoder in the second multimodal emotion recognition model to obtain a third multimodal emotion recognition model.

[0153] Figure 5 FIG. 2 is a schematic diagram of a third multimodal emotion recognition model provided by an embodiment of the present disclosure. Figure 5 As shown, the third multimodal emotion recognition model may include a text encoder, an audio encoder, an adaptation module, a speaker encoder, a fusion module and a multimodal decoder.

[0154] The disclosed embodiment also provides an emotion recognition method, comprising: obtaining a third multimodal emotion recognition model obtained by the training method of the emotion recognition model; obtaining multimodal data to be processed, wherein the multimodal data to be processed includes audio to be processed and text to be processed; inputting the audio to be processed and the text to be processed into the third multimodal emotion recognition model to obtain an emotion category prediction result corresponding to the multimodal data to be processed.

[0155] In the disclosed embodiment, the third text feature vector corresponding to the text to be processed can be extracted by the text encoder in the third multimodal emotion recognition model, the third audio feature vector corresponding to the audio to be processed can be extracted by the audio encoder in the third multimodal emotion recognition model, and the second speaker feature vector corresponding to the audio to be processed can be extracted by the speaker encoder in the third multimodal emotion recognition model. The third audio feature vector can be reduced in dimension by the adaptation module in the third multimodal emotion recognition model to obtain a third reduced-dimensional feature vector. The third text feature vector, the third reduced-dimensional feature vector and the second speaker feature vector can be fused by the fusion module to obtain a third fused feature vector. The third fused feature vector can be processed by the multimodal decoder in the third multimodal emotion recognition model to obtain an emotion category prediction result corresponding to the multimodal data to be processed.

[0156] In a possible implementation, the obtaining of the multimodal data to be processed includes: obtaining audio to be processed; and performing speech recognition on the audio to be processed to obtain text to be processed. In this implementation, speech recognition can be performed on the audio to be processed by a speech recognition model to obtain the text to be processed.

[0157] The training method of the emotion recognition model and the emotion recognition method provided in the embodiments of the present disclosure can also be migrated to other multimodal recognition tasks, such as speaker recognition, intent understanding, punctuation recognition and other tasks based on text and audio.

[0158] The emotion recognition model training method and emotion recognition method provided by the embodiments of the present disclosure can be applied to technical fields such as artificial intelligence, multimodal emotion recognition, speaker adaptation, etc., and are not limited here.

[0159] The following describes an embodiment of the present disclosure through a specific application scenario.

[0160] 1. Training Phase

[0161] 1. Obtain a multimodal training dataset

[0162] (1) Open source speech emotion data can be collected, or speech emotion data can be crawled or annotated based on the Internet.

[0163] (2) The speech emotion data can be screened to select training audios of multiple emotion categories corresponding to multiple speakers. Speech recognition can be performed on the training audio to obtain training text corresponding to the training audio. Multimodal training data can be obtained based on the training audio, training text, and emotion labels and speaker labels in the speech emotion data. The various multimodal training data can form a multimodal training data set.

[0164] 2. Pre-training stage

[0165] (1) A speaker recognition model may be pre-trained based on the training audio and speaker labels in the multimodal training data in the multimodal training dataset, wherein the speaker recognition model may include a speaker encoder, a speaker decoder, and a classifier.

[0166] (2) An unsupervised audio coding model may be pre-trained based on the training audio in the multimodal training data in the multimodal training dataset. The pre-trained audio coding model may be used as an audio encoder.

[0167] (3) A language model may be pre-trained based on the training text in the multimodal training data set. The pre-trained language model may be used as a text encoder.

[0168] 3. Joint training phase

[0169] (1) Based on the multimodal training data set, a first multimodal emotion recognition model is trained, wherein the first multimodal emotion recognition model includes a text encoder, an audio encoder, an adaptation module, a fusion module and a multimodal decoder.

[0170] (2) In response to the first multimodal emotion recognition model being trained to meet the first preset training condition, a speaker encoder and a speaker decoder are added to the first multimodal emotion recognition model to obtain a second multimodal emotion recognition model. The second multimodal emotion recognition model includes a text encoder, an audio encoder, an adaptation module, a speaker encoder, a speaker decoder, a fusion module and a multimodal decoder.

[0171] Based on the multimodal training data set, a second multimodal emotion recognition model is trained. In the process of training the second multimodal emotion recognition model, the parameters of the text encoder, the audio encoder and the adaptation module can be fixed, and only the parameters of the speaker encoder, the speaker decoder, the fusion module and the multimodal decoder are updated.

[0172] (3) In response to the second multimodal emotion recognition model being trained to meet the second preset training condition, the speaker decoder in the second multimodal emotion recognition model is deleted to obtain a third multimodal emotion recognition model.

[0173] 2. Deployment Phase

[0174] 1. Voice Recognition

[0175] The speech recognition model is used to perform speech recognition on the audio to be processed to obtain the text to be processed corresponding to the audio to be processed.

[0176] 2. Multimodal emotion recognition

[0177] The audio to be processed and the text to be processed are input into the third multimodal emotion recognition model, and the emotion category prediction results corresponding to the audio to be processed and the text to be processed are obtained through the third multimodal emotion recognition model.

[0178] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, the present disclosure will not repeat them. It can be understood by those skilled in the art that in the above-mentioned method of the specific implementation method, the specific execution order of each step should be determined according to its function and possible internal logic.

[0179] In addition, the present disclosure also provides a training device for an emotion recognition model, an emotion recognition device, an electronic device, a computer-readable storage medium, and a computer program product, all of which can be used to implement any emotion recognition model training method or emotion recognition method provided by the present disclosure. The corresponding technical solutions and technical effects can be found in the corresponding records in the method part and will not be repeated here.

[0180] Figure 6 A block diagram of a training device for an emotion recognition model provided by an embodiment of the present disclosure is shown. Figure 6 As shown, the training device of the emotion recognition model includes:

[0181] A first obtaining module 61 is used to obtain a multimodal training data set, wherein the multimodal training data set includes multiple multimodal training data;

[0182] A first training module 62, configured to train a first multimodal emotion recognition model using the multimodal training data set, wherein the first multimodal emotion recognition model includes a text encoder, an audio encoder, a fusion module and a multimodal decoder;

[0183] An adding module 63 is used for adding a speaker encoder and a speaker decoder to the first multimodal emotion recognition model in response to the first multimodal emotion recognition model being trained to meet a first preset training condition, so as to obtain a second multimodal emotion recognition model;

[0184] The second training module 64 is used to use the multimodal training data set to train the second multimodal emotion recognition model until a second preset training condition is met.

[0185] In a possible implementation, during the training of the second multimodal emotion recognition model, the parameters of the text encoder and the audio encoder are fixed, and the parameters of the speaker encoder, the speaker decoder, the fusion module and the multimodal decoder are updated.

[0186] In a possible implementation, any item of multimodal training data includes a sentiment label, training text, and training audio;

[0187] The first training module 62 is used to:

[0188] For any item of multimodal training data in the multimodal training data set, extracting a first text feature vector corresponding to a training text in the multimodal training data by the text encoder, and extracting a first audio feature vector corresponding to a training audio in the multimodal training data by the audio encoder;

[0189] fusing the first text feature vector and the first audio feature vector through the fusion module to obtain a first fused feature vector corresponding to the multimodal training data;

[0190] Inputting the first fused feature vector into the multimodal decoder, and outputting a first emotion category prediction result corresponding to the multimodal training data through the multimodal decoder;

[0191] The first multimodal emotion recognition model is trained according to the first emotion category prediction result and the emotion label in the multimodal training data.

[0192] In a possible implementation, the first multimodal emotion recognition model further includes an adaptation module;

[0193] The first training module 62 is used to:

[0194] Performing dimensionality reduction on the first audio feature vector by the adaptation module to obtain a first dimensionality reduction feature vector corresponding to the first audio feature vector;

[0195] The first text feature vector and the first dimension reduction feature vector are fused through the fusion module to obtain a first fused feature vector corresponding to the multimodal training data.

[0196] In a possible implementation, the first training module 62 is used to:

[0197] Processing the first dimensionality reduction feature vector by the fusion module based on a cross attention mechanism to obtain a second dimensionality reduction feature vector corresponding to the first audio feature vector;

[0198] The first text feature vector and the second dimension reduction feature vector are concatenated or added by the fusion module to obtain a first fused feature vector corresponding to the multimodal training data.

[0199] In one possible implementation, any item of multimodal training data also includes a speaker label;

[0200] The second training module 64 is used to:

[0201] For any item of multimodal training data in the multimodal training data set, extracting a second text feature vector corresponding to the training text in the multimodal training data by the text encoder, extracting a second audio feature vector corresponding to the training audio in the multimodal training data by the audio encoder, and extracting a first speaker feature vector corresponding to the training audio by the speaker encoder;

[0202] fusing the first text feature vector, the first audio feature vector and the first speaker feature vector through the fusion module to obtain a second fused feature vector corresponding to the multimodal training data;

[0203] Inputting the second fused feature vector into the multimodal decoder, and outputting a second emotion category prediction result corresponding to the multimodal training data through the multimodal decoder;

[0204] Inputting the first speaker feature vector into the speaker decoder, and outputting a second speaker prediction result corresponding to the multimodal training data through the speaker decoder;

[0205] The second multimodal emotion recognition model is trained according to difference information between the second emotion category prediction result and the emotion label in the multimodal training data, and difference information between the second speaker prediction result and the speaker label in the multimodal training data.

[0206] In a possible implementation manner, the device further includes:

[0207] A speaker prediction module, for any one of the multimodal training data in the multimodal training data set, inputting the training audio in the multimodal training data into a speaker recognition model, and outputting a first speaker prediction result corresponding to the training audio through the speaker recognition model; wherein the speaker recognition model includes the speaker encoder, the speaker decoder and the speaker classifier;

[0208] The third training module is used to train the speaker recognition model according to the first speaker prediction result and the speaker label in the multimodal training data.

[0209] In a possible implementation manner, the device further includes:

[0210] The fourth training module is used to use the training audio in the multimodal training data in the multimodal training data set to train the audio encoder.

[0211] In a possible implementation manner, the device further includes:

[0212] The fifth training module is used to use the training text in the multimodal training data in the multimodal training data set to train the text encoder.

[0213] In a possible implementation manner, the device further includes:

[0214] A deleting module is used to delete the speaker decoder in the second multimodal emotion recognition model in response to the second multimodal emotion recognition model being trained to meet the second preset training condition, so as to obtain a third multimodal emotion recognition model.

[0215] The present disclosure also provides an emotion recognition device, including:

[0216] An acquisition module, used to acquire a third multimodal emotion recognition model obtained by the training method of the emotion recognition model;

[0217] A second obtaining module is used to obtain multimodal data to be processed, wherein the multimodal data to be processed includes audio to be processed and text to be processed;

[0218] The emotion recognition module is used to input the audio to be processed and the text to be processed into the third multimodal emotion recognition model to obtain the emotion category prediction result corresponding to the multimodal data to be processed.

[0219] In a possible implementation manner, the second obtaining module is used to:

[0220] Get the audio to be processed;

[0221] Perform speech recognition on the audio to be processed to obtain text to be processed.

[0222] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. Its specific implementation and technical effects can refer to the description of the above method embodiments. For the sake of brevity, they will not be repeated here.

[0223] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium may be a non-volatile computer-readable storage medium, or may be a volatile computer-readable storage medium.

[0224] The embodiment of the present disclosure further provides a computer program, including a computer-readable code. When the computer-readable code is executed in an electronic device, a processor in the electronic device executes the above method.

[0225] The embodiments of the present disclosure also provide a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes the above method.

[0226] An embodiment of the present disclosure also provides an electronic device, comprising: one or more processors; and a memory for storing executable instructions; wherein the one or more processors are configured to call the executable instructions stored in the memory to execute the above method.

[0227] The electronic device may be provided as a terminal, a server, or a device in other forms.

[0228] Figure 7 1 is a block diagram of an electronic device 1900 provided in an embodiment of the present disclosure. For example, the electronic device 1900 may be provided as a server. Figure 7 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0229] The electronic device 1900 may also include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as a Microsoft Server operating system (Windows Server 2003). TM ), a graphical user interface operating system launched by Apple (MacOS X TM ), a multi-user, multi-process computer operating system (Unix TM ), a free and open source Unix-like operating system (Linux TM ), an open source Unix-like operating system (FreeBSD TM ) or similar.

[0230] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.

[0231] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0232] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.

[0233] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0234] The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0235] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0236] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0237] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0238] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions.

[0239] The computer program product may be implemented in hardware, software or a combination thereof. In one optional embodiment, the computer program product is embodied as a computer storage medium, and in another optional embodiment, the computer program product is embodied as a software product, such as a software development kit (SDK) and the like.

[0240] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.

[0241] If the technical solution of the embodiments of the present disclosure involves personal information, the product using the technical solution of the embodiments of the present disclosure has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of the embodiments of the present disclosure involves sensitive personal information, the product using the technical solution of the embodiments of the present disclosure has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the requirement of "explicit consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to collect his or her personal information; or on the device for processing personal information, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0242] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A training method for an emotion recognition model, characterized in that: include: Obtaining a multimodal training data set, wherein the multimodal training data set includes multiple multimodal training data; Using the multimodal training data set to train a first multimodal emotion recognition model, wherein the first multimodal emotion recognition model includes a text encoder, an audio encoder, a fusion module and a multimodal decoder; In response to the first multimodal emotion recognition model being trained to satisfy a first preset training condition, a speaker encoder and a speaker decoder are added to the first multimodal emotion recognition model to obtain a second multimodal emotion recognition model; Using the multimodal training data set, training the second multimodal emotion recognition model until a second preset training condition is met, wherein, during the training of the second multimodal emotion recognition model, the parameters of the text encoder and the audio encoder are fixed, and the parameters of the speaker encoder, the speaker decoder, the fusion module, and the multimodal decoder are updated; In response to the second multimodal emotion recognition model being trained to satisfy the second preset training condition, the speaker decoder in the second multimodal emotion recognition model is deleted to obtain a third multimodal emotion recognition model.

2. The method according to claim 1, characterized in that Any multimodal training data includes sentiment labels, training texts, and training audios; The adopting the multimodal training data set to train a first multimodal emotion recognition model comprises: For any item of multimodal training data in the multimodal training data set, extracting a first text feature vector corresponding to a training text in the multimodal training data by the text encoder, and extracting a first audio feature vector corresponding to a training audio in the multimodal training data by the audio encoder; fusing the first text feature vector and the first audio feature vector through the fusion module to obtain a first fused feature vector corresponding to the multimodal training data; Inputting the first fused feature vector into the multimodal decoder, and outputting a first emotion category prediction result corresponding to the multimodal training data through the multimodal decoder; The first multimodal emotion recognition model is trained according to the first emotion category prediction result and the emotion label in the multimodal training data.

3. The method according to claim 2, characterized in that The first multimodal emotion recognition model also includes an adaptation module; The fusing the first text feature vector and the first audio feature vector by the fusion module to obtain a first fused feature vector corresponding to the multimodal training data includes: Performing dimensionality reduction on the first audio feature vector by the adaptation module to obtain a first dimensionality reduction feature vector corresponding to the first audio feature vector; The first text feature vector and the first dimension reduction feature vector are fused through the fusion module to obtain a first fused feature vector corresponding to the multimodal training data.

4. The method according to claim 3, characterized in that The step of fusing the first text feature vector and the first dimension reduction feature vector by the fusion module to obtain a first fused feature vector corresponding to the multimodal training data includes: Processing the first dimensionality reduction feature vector by the fusion module based on a cross attention mechanism to obtain a second dimensionality reduction feature vector corresponding to the first audio feature vector; The first text feature vector and the second dimension reduction feature vector are concatenated or added by the fusion module to obtain a first fused feature vector corresponding to the multimodal training data.

5. The method according to claim 2, characterized in that: Any multimodal training data also includes speaker labels; The adopting the multimodal training data set to train the second multimodal emotion recognition model includes: For any item of multimodal training data in the multimodal training data set, extracting a second text feature vector corresponding to the training text in the multimodal training data by the text encoder, extracting a second audio feature vector corresponding to the training audio in the multimodal training data by the audio encoder, and extracting a first speaker feature vector corresponding to the training audio by the speaker encoder; fusing the first text feature vector, the first audio feature vector and the first speaker feature vector through the fusion module to obtain a second fused feature vector corresponding to the multimodal training data; Inputting the second fused feature vector into the multimodal decoder, and outputting a second emotion category prediction result corresponding to the multimodal training data through the multimodal decoder; Inputting the first speaker feature vector into the speaker decoder, and outputting a second speaker prediction result corresponding to the multimodal training data through the speaker decoder; The second multimodal emotion recognition model is trained according to difference information between the second emotion category prediction result and the emotion label in the multimodal training data, and difference information between the second speaker prediction result and the speaker label in the multimodal training data.

6. The method according to claim 1, characterized in that Before adopting the multimodal training data set to train the first multimodal emotion recognition model, the method further includes: For any one of the multimodal training data in the multimodal training data set, input the training audio in the multimodal training data into a speaker recognition model, and output a first speaker prediction result corresponding to the training audio through the speaker recognition model; wherein the speaker recognition model includes the speaker encoder, the speaker decoder and the speaker classifier; The speaker recognition model is trained according to the first speaker prediction result and the speaker labels in the multimodal training data.

7. The method according to claim 1, characterized in that Before adopting the multimodal training data set to train the first multimodal emotion recognition model, the method further includes: The audio encoder is trained by using training audio in the multimodal training data in the multimodal training data set.

8. The method according to claim 1, characterized in that: Before adopting the multimodal training data set to train the first multimodal emotion recognition model, the method further includes: The text encoder is trained by using training text in the multimodal training data in the multimodal training data set.

9. An emotion recognition method, characterized in that: include: Obtaining a third multimodal emotion recognition model obtained by the emotion recognition model training method according to any one of claims 1 to 8; Obtaining multimodal data to be processed, wherein the multimodal data to be processed includes audio to be processed and text to be processed; The audio to be processed and the text to be processed are input into the third multimodal emotion recognition model to obtain an emotion category prediction result corresponding to the multimodal data to be processed.

10. The method according to claim 9, characterized in that The obtaining of the multimodal data to be processed comprises: Get the audio to be processed; Perform speech recognition on the audio to be processed to obtain text to be processed.

11. A training device for an emotion recognition model, characterized in that: include: A first acquisition module is used to obtain a multimodal training data set, wherein the multimodal training data set includes multiple multimodal training data; A first training module, configured to train a first multimodal emotion recognition model using the multimodal training data set, wherein the first multimodal emotion recognition model comprises a text encoder, an audio encoder, a fusion module and a multimodal decoder; An adding module, configured to add a speaker encoder and a speaker decoder to the first multimodal emotion recognition model in response to the first multimodal emotion recognition model being trained to meet a first preset training condition, so as to obtain a second multimodal emotion recognition model; A second training module, used to train the second multimodal emotion recognition model using the multimodal training data set until a second preset training condition is met, wherein, during the training of the second multimodal emotion recognition model, the parameters of the text encoder and the audio encoder are fixed, and the parameters of the speaker encoder, the speaker decoder, the fusion module, and the multimodal decoder are updated; A deleting module is used to delete the speaker decoder in the second multimodal emotion recognition model in response to the second multimodal emotion recognition model being trained to meet the second preset training condition, so as to obtain a third multimodal emotion recognition model.

12. An emotion recognition device, characterized in that: include: An acquisition module, used to acquire a third multimodal emotion recognition model obtained by the emotion recognition model training method according to claim 11; A second obtaining module is used to obtain multimodal data to be processed, wherein the multimodal data to be processed includes audio to be processed and text to be processed; The emotion recognition module is used to input the audio to be processed and the text to be processed into the third multimodal emotion recognition model to obtain the emotion category prediction result corresponding to the multimodal data to be processed.

13. An electronic device, characterized in that: include: one or more processors; a memory for storing executable instructions; The one or more processors are configured to call the executable instructions stored in the memory to execute the method according to any one of claims 1 to 10.

14. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 10 is implemented.

15. A computer program product comprising computer readable code, or a non-volatile computer readable storage medium carrying computer readable code, characterized in that: When the computer readable code is executed in an electronic device, a processor in the electronic device executes the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Multi-mode depressive emotion recognition method and device

    CN115641543A

  • Method for training emotion recognition model and emotion recognition method and device

    CN115713797A

  • Speech emotion recognition method and device, electronic equipment and storage medium

    CN116386674A