Emotion recognition model determination method and apparatuses, and emotion type recognition method and apparatuses
By obtaining model training data, conducting data training and creating emotion recognition models, combining speech encoder and emotion joint network, the shortcomings in the accuracy and application scope of the existing emotion recognition model are solved, and efficient frame-level speech emotion recognition is achieved.
Patent Information
- Application Number
- PCT/CN2025/070198
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2025-01-02
- Publication Date
- 2025-07-17
AI Technical Summary
The existing emotion recognition model has shortcomings in recognition accuracy, especially in frame-level speech recognition, and its application scenarios and range are limited.
By obtaining model training data, data training is performed to obtain an emotional joint network, the loss function is determined based on the emotional prediction probability, and an emotion recognition model is created based on the loss function and the emotional joint network, and fine-grained frame-level speech emotion recognition is combined with a speech encoder, a text predictor and an emotion joint network.
It greatly improves the recognition accuracy of the emotion recognition model, especially in frame-level speech recognition, and expands the application scenarios and scope of the emotion recognition model.
Smart Images

Figure CN2025070198_17072025_PF_FP_ABST
Abstract
Description
Method for determining emotion recognition model, method and device for identifying emotion type
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 11, 2024, with application number "202410044481.X", the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of speech recognition technology, and in particular to a method for determining an emotion recognition model, and a method and device for identifying emotion types. Background Art
[0003] Speech emotion recognition refers to the process of outputting the emotion label corresponding to a given speech. At present, emotion recognition models can be used to identify the user's emotion type by recognizing the user's speech, but existing emotion recognition models have technical problems such as low recognition accuracy. Summary of the Invention
[0004] This application aims to solve at least one of the technical problems existing in the prior art or related art.
[0005] To this end, the first aspect of this application is to propose a method for determining an emotion recognition model.
[0006] The second aspect of this application is to propose a method for identifying emotion types.
[0007] The third aspect of this application is to propose a device for determining an emotion recognition model.
[0008] The fourth aspect of the present application is to propose another device for determining an emotion recognition model.
[0009] The fifth aspect of the present application is to provide a device for identifying emotion types.
[0010] The sixth aspect of the present application is to propose another emotion type recognition device.
[0011] The seventh aspect of the present application is to provide a readable storage medium.
[0012] An eighth aspect of the present application is to provide a computer program product.
[0013] In view of this, according to the first aspect of the present application, a method for determining an emotion recognition model is proposed, and the method for determining an emotion recognition model includes: obtaining model training data; performing data training on the model training data to obtain an emotion joint network, and the emotion joint network stores the emotion prediction probability corresponding to the model training data; determining a loss function based on the emotion prediction probability; and creating an emotion recognition model based on the loss function and the emotion joint network.
[0014] The method for determining the emotion recognition model in this technical solution greatly improves the recognition accuracy of the emotion recognition model, while also improving the recognition accuracy of the emotion recognition model for frame-level speech, and expanding the application scenarios and scope of the emotion recognition model.
[0015] According to the second aspect of the present application, a method for identifying emotion types is proposed, and the method for identifying emotion types includes: acquiring speech data and an emotion recognition model, where the emotion recognition model is an emotion recognition model determined by the method for determining the emotion recognition model defined in the first aspect; determining first feature data corresponding to the speech data based on an emotion joint network of the emotion recognition model; and determining emotion type information corresponding to the speech data based on the first feature data.
[0016] The emotion type recognition method in this technical solution greatly improves the recognition accuracy of speech data through the emotion recognition model, improves the recognition accuracy of frame-level speech, ensures the accuracy of emotion type information, and expands the application scenarios and scope of the emotion recognition model.
[0017] According to the third aspect of the present application, a device for determining an emotion recognition model is proposed, and the device for determining the emotion recognition model includes: a first processing module for acquiring model training data; the first processing module is also used to perform data training on the model training data to obtain an emotion joint network, and the emotion joint network stores the emotion prediction probability corresponding to the model training data; the first processing module is also used to determine the loss function based on the emotion prediction probability; the first processing module is also used to create an emotion recognition model based on the loss function and the emotion joint network.
[0018] The device for determining the emotion recognition model in this technical solution greatly improves the recognition accuracy of the emotion recognition model, while also improving the recognition accuracy of the emotion recognition model for frame-level speech, and expanding the application scenarios and scope of the emotion recognition model.
[0019] According to a fourth aspect of the present application, another device for determining an emotion recognition model is provided, comprising a processor and a memory, wherein the memory stores a program or instructions that, when executed by the processor, implement the steps of the method for determining an emotion recognition model as described in any of the above-mentioned technical solutions. Therefore, this device for determining an emotion recognition model has all the beneficial effects of the method for determining an emotion recognition model as described in any of the above-mentioned technical solutions, and no further details are given here.
[0020] According to the fifth aspect of the present application, an emotion type recognition device is proposed, and the emotion type recognition device includes: a second processing module, used to obtain speech data and an emotion recognition model, the emotion recognition model is an emotion recognition model determined by the emotion recognition model determination method defined in the first aspect; the second processing module is also used to determine the first feature data corresponding to the speech data based on the emotion joint network of the emotion recognition model; the second processing module is also used to determine the emotion type information corresponding to the speech data based on the first feature data.
[0021] The emotion type recognition device in this technical solution greatly improves the recognition accuracy of voice data through the emotion recognition model, improves the recognition accuracy of frame-level speech, ensures the accuracy of emotion type information, and expands the application scenarios and application scope of the emotion recognition model.
[0022] According to a sixth aspect of the present application, another device for identifying emotion types is provided, comprising a processor and a memory, wherein the memory stores a program or instructions that, when executed by the processor, implement the steps of the emotion type identification method described in any of the above technical solutions. Therefore, this device for identifying emotion types possesses all the beneficial effects of the emotion type identification method described in any of the above technical solutions, and no further description is given here.
[0023] According to a seventh aspect of the present application, a readable storage medium is provided, on which a program or instruction is stored. When executed by a processor, the program or instruction implements the method for determining an emotion recognition model or the method for identifying an emotion type as described in any of the above technical solutions. Therefore, the readable storage medium has all the beneficial effects of the method for determining an emotion recognition model or the method for identifying an emotion type as described in any of the above technical solutions, and no further description is given here.
[0024] According to an eighth aspect of the present application, a computer program product is provided, comprising computer instructions. When executed by a processor, the computer instructions implement the method for determining an emotion recognition model or the method for identifying an emotion type as described in any of the aforementioned technical solutions. Therefore, this computer program product has all the beneficial effects of the method for determining an emotion recognition model or the method for identifying an emotion type as described in any of the aforementioned technical solutions, and no further description is given herein.
[0025] Additional aspects and advantages of the present application will become apparent in the following description or may be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0027] FIG1 shows a flow chart of a method for determining an emotion recognition model in an embodiment of the present application;
[0028] FIG2 shows one schematic diagram of an emotion recognition model in an embodiment of the present application;
[0029] FIG3 shows a second schematic diagram of an emotion recognition model in an embodiment of the present application;
[0030] FIG4 shows a second flow chart of a method for determining an emotion recognition model in an embodiment of the present application;
[0031] FIG5 shows a third flow chart of a method for determining an emotion recognition model in an embodiment of the present application;
[0032] FIG6 shows one of the flowcharts of the method for identifying emotion types in an embodiment of the present application;
[0033] FIG7 shows a second flow chart of the method for identifying emotion types in an embodiment of the present application;
[0034] FIG8 shows a third flow chart of the method for identifying emotion types in an embodiment of the present application;
[0035] FIG9 shows a fourth flow chart of the method for identifying emotion types in an embodiment of the present application;
[0036] FIG10 shows a fifth flow chart of the method for identifying emotion types in an embodiment of the present application;
[0037] FIG11 shows a sixth flow chart of the method for identifying emotion types in an embodiment of the present application;
[0038] FIG12 shows a seventh flow chart of the method for identifying emotion types in an embodiment of the present application;
[0039] FIG13 shows an eighth flow chart of the method for identifying emotion types in an embodiment of the present application;
[0040] FIG14 shows a ninth flowchart of the method for identifying emotion types in an embodiment of the present application;
[0041] FIG15 shows a tenth flowchart of the method for identifying emotion types in an embodiment of the present application;
[0042] FIG16 is a schematic diagram showing a method for identifying emotion types in an embodiment of the present application;
[0043] FIG17 shows one structural block diagram of a device for determining an emotion recognition model in an embodiment of the present application;
[0044] FIG18 shows one structural block diagram of an apparatus for identifying emotion types in an embodiment of the present application;
[0045] FIG19 shows a second structural block diagram of the device for determining an emotion recognition model in an embodiment of the present application;
[0046] FIG20 shows a second structural block diagram of the device for identifying emotion types in an embodiment of the present application. DETAILED DESCRIPTION
[0047] In order to more clearly understand the above-mentioned objects, features and advantages of the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features therein can be combined with each other in the absence of conflict.
[0048] In the following description, many specific details are set forth to facilitate a full understanding of the present application. However, the present application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present application is not limited to the specific embodiments disclosed below.
[0049] 1 to 20 , the method for determining the emotion recognition model, the method and device for identifying the emotion type provided in the embodiments of the present application are described in detail through specific embodiments and their application scenarios.
[0050] The technical solution for the method for determining an emotion recognition model provided in this application can be implemented by a determination device, or can be determined based on actual usage requirements, which is not specifically limited here. In order to more clearly describe the method for determining an emotion recognition model provided in this application, the following description is based on the determination device as the implementation subject.
[0051] As shown in FIG1 , an embodiment of the present application provides a method for determining an emotion recognition model. The method for determining an emotion recognition model includes:
[0052] Step 102: Obtain model training data;
[0053] Step 104: performing data training on the model training data to obtain an emotion joint network;
[0054] Step 106, determining a loss function based on the emotion prediction probability;
[0055] Step 108: Create an emotion recognition model based on the loss function and the emotion joint network.
[0056] In this embodiment, a method for determining an emotion recognition model is proposed, where a determining device obtains model training data, wherein the model training data is data used for model training.
[0057] Exemplarily, the model training data may include voice training data and emotion type data.
[0058] The determination device performs data training on the model training data to obtain an emotion joint network, wherein the emotion joint network is a deep learning network that can identify emotion types, and the emotion joint network stores the emotion prediction probability corresponding to the model training data, and the emotion prediction probability is the emotion prediction probability output by the emotion joint network.
[0059] Exemplarily, as shown in FIG2 , the emotion joint network may include an emotion grid. The emotion grid is constructed during the training phase, and each node in the emotion grid represents the probability distribution of emotions after outputting a number of texts in the current frame.
[0060] Exemplarily, the emotion prediction probability may include the target emotion posterior probability and the neutral emotion posterior probability, wherein the target emotion posterior probability is the posterior probability of emotions such as joy, anger, sadness, and happiness, and the neutral emotion posterior probability is the posterior probability of neutral emotion (a state without emotion).
[0061] The determining device determines a loss function based on the emotion prediction probability, wherein the loss function is a function that limits the model output.
[0062] Exemplarily, the loss function is determined by the maximum pooling loss of the sentiment grid.
[0063] The determination device creates an emotion recognition model based on the loss function and the emotion joint network.
[0064] Exemplarily, a recognition model including an emotion joint network is created, and the model output is optimized through a loss function to obtain an emotion recognition model, wherein the emotion recognition model is a model that can recognize the emotion type corresponding to the speech.
[0065] Exemplarily, as shown in Figure 2, the emotion recognition model may include a speech encoder, a text predictor, an emotion joint network and a text joint network, wherein the speech encoder is used to take the speech signal as input and output the encoded speech features, the text predictor is used to take the text as input and output the encoded text features, the text joint network is used to integrate speech features and text features, and output joint features for preset speech recognition text symbol prediction, and the emotion joint network is used to integrate speech features and text features, and output joint features for fine-grained frame-level emotion prediction.
[0066] Exemplarily, as shown in FIG3 , the emotion recognition model may include a speech encoder, a text predictor, a space symbol predictor, a space symbol joint network, an emotion joint network, and a text joint network. This structure can jointly perform fine-grained frame-level speech emotion recognition and speech recognition, and decouple the space symbol from the vocabulary, while making it a text separator and emotion discriminator, and using it as an accumulation variable of the current speech information and text information, so as to more effectively utilize the ending space symbol to judge the frame-level speech emotion.
[0067] The space symbol prediction network is used to separate space symbol prediction from vocabulary prediction by employing two independent predictors. The text predictor and text union network predict the text vocabulary set excluding the special space symbol. A separate space symbol predictor is used to encode historical text and output encoded features that contain both text semantic information and frame-level sentiment information. The space symbol union network integrates speech features and text features. The output joint features are used to calculate the probability of space symbols and, together with the text vocabulary, form the probability distribution of the next text symbol.
[0068] The method for determining the emotion recognition model in this embodiment greatly improves the recognition accuracy of the emotion recognition model, while also improving the recognition accuracy of the emotion recognition model for frame-level speech, and expanding the application scenarios and scope of the emotion recognition model.
[0069] In some embodiments, optionally, as shown in FIG4 , a method for determining an emotion recognition model is proposed, which determines a loss function based on emotion prediction probability, including:
[0070] Step 402, calculating the maximum value of multiple first probabilities to obtain a first target probability;
[0071] Step 404, calculating the maximum value among the plurality of second probabilities to obtain a second target probability;
[0072] Step 406: Calculate the difference between the first target probability and the second target probability to obtain a loss function.
[0073] In this embodiment, the emotion prediction probability includes multiple first probabilities and multiple second probabilities, where the first probability is the posterior probability of the target emotion, and the second probability is the posterior probability of the neutral emotion. It should be noted that the posterior probability of the target emotion is the posterior probability of emotions such as joy, anger, sadness, and happiness, and the posterior probability of the neutral emotion is the posterior probability of the neutral emotion (a state without emotion).
[0074] The determination device calculates the maximum value of multiple first probabilities to obtain a first target probability, and calculates the maximum value of multiple second probabilities to obtain a second target probability, wherein the first target probability is the maximum value of the multiple first probabilities, and the second target probability is the maximum value of the multiple second probabilities.
[0075] The determining device then calculates the difference between the first target probability and the second target probability to obtain a loss function.
[0076] For example, the loss function can be calculated as:
[0077] Among them, Llattice is the loss function, max is the operator for finding the maximum value, The posterior probability of the target emotion predicted by the joint network after the neural converter inputs t frames and outputs u characters, The posterior probability of neutral sentiment predicted by the joint network after the neural converter inputs t frames and outputs u characters.
[0078] For example, the loss function can also be expressed as:
[0079] Among them, Llattice is the loss function, P is the target emotional area in the set emotional grid, N is the neutral emotional area in the set emotional grid, The posterior probability of the target emotion predicted by the joint network after the neural converter inputs t frames and outputs u characters, The posterior probability of neutral sentiment predicted by the joint network after the neural converter inputs t frames and outputs u characters.
[0080] The method for determining the emotion recognition model in this embodiment calculates the maximum value of multiple first probabilities to obtain a first target probability, calculates the maximum value of multiple second probabilities to obtain a second target probability, and then calculates the difference between the first target probability and the second target probability to obtain a loss function, thereby improving the data accuracy of the loss function and thus improving the recognition accuracy of the emotion recognition model.
[0081] In some embodiments, optionally, as shown in FIG5 , a method for determining an emotion recognition model is proposed, wherein the emotion recognition model is created based on a loss function and an emotion joint network, including:
[0082] Step 502: creating a data model based on the emotion joint network;
[0083] Step 504: Update the data model according to the loss function to obtain an emotion recognition model.
[0084] In this embodiment, the determination device creates a data model based on the emotion joint network, and then updates the data model according to the loss function to obtain an emotion recognition model, wherein the data model is an initial recognition model.
[0085] Exemplarily, the determining device constrains the output of the data model according to the loss function to determine the emotion recognition model.
[0086] The method for determining the emotion recognition model in this embodiment is based on the emotion joint network, creates a data model, and then updates the data model according to the loss function to obtain the emotion recognition model, thereby improving the model accuracy of the emotion recognition model and further improving the recognition accuracy of the emotion recognition model.
[0087] The technical solution for the emotion type recognition method provided in this application can be implemented by a recognition device, which can also be determined based on actual usage requirements and is not specifically limited here. In order to more clearly describe the emotion type recognition method provided in this application, the following description uses the recognition device as the implementation entity.
[0088] In some embodiments, optionally, as shown in FIG6 , a method for identifying an emotion type is proposed. The method for identifying an emotion type includes:
[0089] Step 602: Acquire speech data and emotion recognition model;
[0090] Step 604: determining first feature data corresponding to the speech data based on the emotion joint network of the emotion recognition model;
[0091] Step 606: Determine emotion type information corresponding to the voice data based on the first feature data.
[0092] In this embodiment, a method for identifying emotion types is proposed, and the identification device obtains voice data and an emotion recognition model, wherein the voice data is the voice data to be identified, and the emotion recognition model is the emotion recognition model determined by the method for determining the emotion recognition model in the above embodiment.
[0093] Exemplarily, the emotion recognition model may be the model shown in FIG2 .
[0094] Exemplarily, the emotion recognition model may be the model shown in FIG3 .
[0095] Exemplarily, the voice data may be fine-grained frame-level voice data.
[0096] The recognition device determines first feature data corresponding to the speech data based on the emotion association network of the emotion recognition model, wherein the first feature data is feature data output by the emotion association network.
[0097] Exemplarily, the first feature data is used for fine-grained frame-level emotion prediction.
[0098] The recognition device determines the emotion type information corresponding to the voice data based on the first feature data, wherein the emotion type information is used to represent information about the emotion type corresponding to the voice data.
[0099] For example, the emotion type information may include labels such as joy, anger, sadness, or happiness.
[0100] The emotion type recognition method in this embodiment greatly improves the recognition accuracy of speech data through the emotion recognition model, improves the recognition accuracy of frame-level speech, ensures the accuracy of emotion type information, and expands the application scenarios and scope of the emotion recognition model.
[0101] In some embodiments, optionally, as shown in FIG7 , a method for identifying emotion types is proposed, which determines first feature data corresponding to speech data based on an emotion joint network of an emotion recognition model, including:
[0102] Step 702: input the speech data into a speech encoder to obtain speech features output by the speech encoder;
[0103] Step 704: converting the speech data into text data based on the speech recognition module;
[0104] Step 706: input the text data into a text predictor to obtain a first text feature output by the text predictor;
[0105] Step 708: Input the speech feature and the first text feature into the emotion joint network to obtain first feature data output by the emotion joint network.
[0106] In this embodiment, the emotion recognition model further includes a speech encoder, a text predictor, and a speech recognition module, wherein the speech recognition module is used to recognize speech into text, the speech encoder is used to output features of the speech, and the text predictor is used to output features of the text.
[0107] Exemplarily, the speech encoder may be the speech encoder in FIG. 2 .
[0108] Exemplarily, the text predictor may be the text predictor in FIG. 2 .
[0109] The recognition device inputs the speech data into the speech encoder to obtain the speech features output by the speech encoder, wherein the speech features are feature data corresponding to the speech data.
[0110] Exemplarily, speech data is used as input to a speech encoder, which outputs encoded speech features.
[0111] The recognition device converts the voice data into text data through the voice recognition module, wherein the text data is the text data corresponding to the voice data.
[0112] Exemplarily, the speech recognition module may be a speech recognition model.
[0113] The recognition device inputs the text data into a text predictor to obtain a first text feature output by the text predictor, wherein the first text feature is feature data corresponding to the text data.
[0114] Exemplarily, the first text feature is used for prediction of a preset speech recognition text symbol.
[0115] The recognition device inputs the speech feature and the first text feature into the emotion association network to obtain the first feature data output by the emotion association network.
[0116] Exemplarily, the emotion joint network may integrate the speech feature and the first text feature, and then output the first feature data.
[0117] The emotion type recognition method in this embodiment outputs speech features through a speech encoder, outputs first text features through a text predictor, and then inputs text data into the text predictor to obtain the first text features output by the text predictor, thereby ensuring the data accuracy of the first text features and further ensuring the information accuracy of the emotion type information.
[0118] In some embodiments, optionally, as shown in FIG8 , a method for identifying an emotion type is proposed, and the method for identifying an emotion type further includes:
[0119] Step 802: Input the speech feature and the first text feature into a text joint network to obtain second feature data output by the text joint network;
[0120] Step 804: Determine vocabulary data corresponding to the speech data based on the second feature data.
[0121] In this embodiment, the emotion recognition model further includes a text association network, which is a deep learning network for outputting speech recognition text tokens.
[0122] The recognition device inputs the speech feature and the first text feature into the text joint network to obtain second feature data output by the text joint network, wherein the second feature data is feature data output by the text joint network.
[0123] Exemplarily, the text-joint network may integrate the speech feature and the first text feature, and then output second feature data.
[0124] The recognition device determines vocabulary data corresponding to the voice data based on the second feature data, wherein the vocabulary data is the vocabulary of the voice data.
[0125] Exemplarily, the vocabulary data may specifically be the text vocabulary table in FIG. 2 .
[0126] For example, vocabulary data can be used as emotion labels for speech data.
[0127] The emotion type recognition method in this embodiment obtains second feature data output by the text joint network by inputting speech features and first text features into the text joint network, and then determines the vocabulary data corresponding to the speech data based on the second feature data, thereby ensuring the data accuracy of the vocabulary data and further ensuring the information accuracy of the emotion type information.
[0128] In some embodiments, optionally, as shown in FIG9 , a method for identifying emotion types is proposed, which determines first feature data corresponding to speech data based on an emotion joint network of an emotion recognition model, including:
[0129] Step 902: input the speech data into a speech encoder to obtain speech features output by the speech encoder;
[0130] Step 904: converting the speech data into text data based on the speech recognition module;
[0131] Step 906: input the text data into a symbol predictor to obtain a second text feature output by the symbol predictor;
[0132] Step 908: Input the speech feature and the second text feature into the emotion joint network to obtain first feature data output by the emotion joint network.
[0133] In this embodiment, the emotion recognition model further includes a symbol predictor, a speech encoder, and a speech recognition module, wherein the symbol predictor is used to predict the empty symbol in the speech data.
[0134] Exemplarily, the symbol predictor may be the null symbol predictor in FIG3 .
[0135] Exemplarily, the speech encoder may be the speech encoder in FIG3 .
[0136] The recognition device inputs the speech data into the speech encoder and obtains the speech features output by the speech encoder.
[0137] Exemplarily, speech data is used as input to a speech encoder, which outputs encoded speech features.
[0138] The recognition device converts the voice data into text data through the voice recognition module, wherein the text data is the text data corresponding to the voice data.
[0139] Exemplarily, the speech recognition module may be a speech recognition model.
[0140] The recognition device inputs the text data into the symbol predictor to obtain a second text feature output by the symbol predictor, wherein the second text feature is feature data corresponding to the text data.
[0141] The recognition device inputs the speech feature and the second text feature into the emotion joint network to obtain first feature data output by the emotion joint network.
[0142] Exemplarily, the emotion joint network may integrate the speech feature and the second text feature to output the first feature data.
[0143] The emotion type recognition method in this embodiment ensures the data accuracy of the first feature data and thus ensures the information accuracy of the emotion type information by inputting speech data into a speech encoder to obtain speech features output by the speech encoder, inputting text data into a symbol predictor to obtain second text features output by the symbol predictor, and then inputting the speech features and the second text features into an emotion joint network to obtain first feature data output by the emotion joint network.
[0144] In some embodiments, optionally, as shown in FIG10 , a method for identifying an emotion type is proposed, and the method for identifying an emotion type further includes:
[0145] Step 1002: input text data into a text predictor to obtain a first text feature output by the text predictor;
[0146] Step 1004: input the first text feature and the speech feature into a text joint network to obtain second feature data output by the text joint network;
[0147] Step 1006: Input the speech feature and the second text feature into the symbol association network to obtain third feature data output by the symbol association network;
[0148] Step 1008: Determine vocabulary data corresponding to the speech data based on the second feature data and the third feature data.
[0149] In this embodiment, the emotion recognition model further includes a text joint network, a symbol joint network and a text predictor, wherein the symbol joint network is a deep learning network for identifying space symbols in text.
[0150] Exemplarily, the symbol-joint network may be the empty-symbol-joint network in FIG. 3 .
[0151] Exemplarily, the text joint network may be the text joint network in FIG3 .
[0152] Exemplarily, the text predictor may be the text predictor in FIG3 .
[0153] The recognition device inputs the text data into the text predictor and obtains the first text feature output by the text predictor.
[0154] Exemplarily, text data is taken as input and input into a text predictor, which outputs a first text feature.
[0155] The recognition device inputs the first text feature and the speech feature into the text joint network to obtain second feature data output by the text joint network.
[0156] For example, the text combination network may integrate the first text feature and the speech feature to output the second feature data.
[0157] The recognition device inputs the speech feature and the second text feature into the symbol association network to obtain third feature data output by the symbol association network, wherein the third feature data is used to calculate the probability of the empty symbol in the text data.
[0158] For example, the symbolic joint network may integrate the speech feature and the second text feature to output third feature data.
[0159] The recognition device determines vocabulary data corresponding to the speech data based on the second feature data and the third feature data.
[0160] Illustratively, the vocabulary data may include the text vocabulary in FIG. 3 .
[0161] The emotion type recognition method in this embodiment obtains the second feature data output by the text joint network and the third feature data output by the symbol joint network, and then determines the vocabulary data corresponding to the speech data based on the second feature data and the third feature data, thereby ensuring the data accuracy of the vocabulary data and the information accuracy of the emotion type information.
[0162] In some embodiments, optionally, as shown in FIG11 , a method for identifying an emotion type is proposed, which determines vocabulary data corresponding to the speech data based on the second feature data and the third feature data, including:
[0163] Step 1102, determining segmentation symbol information according to the third feature data;
[0164] Step 1104: Convert the second feature data into vocabulary data according to the segmentation symbol information.
[0165] In this embodiment, the recognition device determines segmentation symbol information based on the third feature data, wherein the segmentation symbol information is segmentation symbol reply distribution information.
[0166] Exemplarily, the segmentation symbol information may be distribution information representing the empty symbols in FIG. 3 .
[0167] The recognition device converts the second feature data into vocabulary data based on the segmentation symbol information.
[0168] Illustratively, the vocabulary data may include the text vocabulary in FIG. 3 .
[0169] The emotion type recognition method in this embodiment determines the segmentation symbol information based on the third feature data, and then converts the second feature data into vocabulary data based on the segmentation symbol information, thereby ensuring the data accuracy of the vocabulary data and further ensuring the information accuracy of the emotion type information.
[0170] In some embodiments, optionally, as shown in FIG12 , a method for identifying an emotion type is proposed, and the method for identifying an emotion type further includes:
[0171] Step 1202: Acquire the user's voice;
[0172] Step 1204 : Segment the user's voice according to the preset duration to obtain voice data.
[0173] In this embodiment, the recognition device obtains the user's voice, and divides the user's voice into data according to a preset time length to obtain voice data, wherein the preset time length is a preset time length.
[0174] Exemplarily, the preset duration may be 20 ms.
[0175] The emotion type recognition method in this embodiment obtains voice data by segmenting the user's voice according to preset time lengths, thereby ensuring the data accuracy of the voice data and further ensuring the information accuracy of the emotion type information.
[0176] In some embodiments, optionally, as shown in FIG13 , a method for identifying an emotion type is proposed. The method for identifying an emotion type includes:
[0177] Step 1302: Acquire speech data and emotion recognition model;
[0178] Step 1304: determining first feature data corresponding to the speech data based on the emotion joint network of the emotion recognition model;
[0179] Step 1306: Determine emotion type information corresponding to the voice data based on the first feature data;
[0180] Step 1308: Generate an emotion tag corresponding to the speech data according to the emotion type information.
[0181] In this embodiment, the recognition device generates an emotion tag corresponding to the speech data according to the emotion type information, wherein the emotion tag is used to express the emotion type corresponding to the speech data.
[0182] For example, the emotion label may be a label such as happy or sad.
[0183] The emotion type recognition method in this embodiment generates emotion labels corresponding to speech data based on emotion type information, thereby ensuring the data accuracy of the emotion labels.
[0184] In some embodiments, optionally, as shown in FIG14 , a method for identifying an emotion type is proposed. The method for identifying an emotion type includes:
[0185] Step 1402, input audio;
[0186] Step 1404: the audio encoder reads one frame.
[0187] Step 1406, the text predictor reads the 1 symbol;
[0188] Step 1408, the text association network outputs the next symbol;
[0189] Step 1410, whether it is a null symbol, if so, go to step 1412, if not, go to step 1406;
[0190] Step 1412, the emotion association network outputs the current emotion;
[0191] Step 1414: Check whether it is the last frame. If so, it is the end. If not, execute step 1404.
[0192] In this embodiment, the inference stage uses a neural converter. For each speech frame, several text symbols are output until a null symbol ends the frame's output. The text predictor's features from the last time step are combined with speech features and passed through an emotion association network to output the current frame's emotion. The recognition device inputs audio, the audio encoder reads one frame of audio, the text predictor reads one symbol, the text association network outputs the next symbol, and determines whether the current symbol is a null symbol. If so, the emotion association network outputs the current emotion. If not, the text predictor reads another symbol, and the recognition device determines whether the current frame is the last frame. If so, recognition ends; otherwise, the audio encoder reads another frame.
[0193] Exemplarily, as shown in Table 1, the emotion recognition model may be an emotion neural converter or a decomposed emotion neural converter, and the recognition accuracy of the emotion recognition model and other models is shown in Table 1.
[0194] Table 1
[0195] In some embodiments, optionally, as shown in FIG15 , a method for identifying an emotion type is proposed. The method for identifying an emotion type includes:
[0196] Step 1502, input audio;
[0197] Step 1504: the audio encoder reads one frame;
[0198] Step 1506: the text predictor and the space predictor read the 1 symbol;
[0199] Step 1508: The text joint network and the space symbol joint network output the next symbol;
[0200] Step 1510: Is it a null symbol? If yes, go to step 1512; if no, go to step 1506.
[0201] Step 1512, the emotion association network outputs the current emotion;
[0202] Step 1514: Check whether it is the last frame. If so, it is over. If not, execute step 1504.
[0203] In this embodiment, in the inference stage, for each speech frame, the probability distribution of the text symbol is calculated jointly by the empty symbol joint network and the text joint network, and several text symbols are output until the empty symbol ends the output content of the frame, and the features of the last time step of the empty symbol predictor are combined with the speech features through the emotion joint network to output the emotion of the current frame. The recognition device inputs audio, the audio encoder reads in 1 frame of audio, the text predictor and the empty symbol predictor read in 1 symbol, the text joint network and the empty symbol joint network output the next symbol, and judge whether the current symbol is an empty symbol. If it is an empty symbol, the emotion joint network outputs the current emotion. If it is not an empty symbol, the text predictor and the empty symbol predictor re-read in 1 symbol. The recognition device judges whether the current frame is the last frame. If it is, the recognition is ended. If not, the audio encoder re-reads 1 frame.
[0204] Exemplarily, as shown in Table 2, the emotion recognition model may be an emotion neural converter or a decomposed emotion neural converter, and the recognition accuracy of the emotion recognition model and other models is shown in Table 2.
[0205] Table 2
[0206] Existing speech emotion recognition methods mostly focus on sentence-level emotion recognition, ignoring the dynamic sequence characteristics of emotion in speech. Specifically, given a piece of emotion-labeled speech data, from a temporal perspective, most frames do not reflect the target emotion or are actually neutral, while only a small number of frames do. Fine-grained, frame-level speech emotion recognition aligns with emotional characteristics and is beneficial for real-time human-computer interaction applications. Furthermore, speech data inherently contains both acoustic and textual semantic information, and speech frames are aligned with the corresponding textual information.
[0207] From the perspective of existing methods, previous methods for fine-grained speech emotion recognition did not consider linguistic information, resulting in a certain amount of information loss, and the method of segmenting speech frames for emotion recognition had a contradiction between the length of speech frames and the semantic alignment of text; and previous methods that simultaneously utilized linguistic and acoustic information failed to consider the emotional fine-grainedness of speech emotion recognition, and most methods found it difficult to synchronously output text and corresponding emotions.
[0208] From a data perspective, most speech emotion datasets only provide sentence-level labels rather than frame-level labels, which makes frame-level speech emotion recognition a weakly supervised problem. How to enable the model to discover significant emotional intervals at the frame level with only sentence-level annotations is also an important problem that needs to be solved urgently.
[0209] This embodiment proposes a method that combines speech recognition with fine-grained emotion recognition to address the shortcomings of previous methods. Fine-grained speech emotion recognition and speech recognition are jointly modeled based on a neural converter. This allows the model to output several texts simultaneously with the input frame, followed by an emotion label as the frame's emotion label. A grid-maximum pooling technique is also proposed to perform fine-grained speech emotion recognition under the weak supervision of sentence-level labels in the dataset. As shown in Figure 16, Curve B shows the target emotion change curve predicted by the model, and Curve A shows the neutral emotion change curve predicted by the model.
[0210] The above methods may be implemented in various ways depending on the specific features and / or example applications. For example, these methods may be implemented through a combination of hardware, firmware, and / or software. For example, in a hardware implementation, the processor may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, electronic devices, other device units for performing the above functions, and / or combinations thereof.
[0211] As shown in FIG17 , an embodiment of the present application provides an apparatus 1700 for determining an emotion recognition model. The apparatus 1700 for determining an emotion recognition model includes:
[0212] The first processing module 1702 is used to obtain model training data;
[0213] The first processing module 1702 is further configured to perform data training on the model training data to obtain an emotion joint network, wherein the emotion joint network stores emotion prediction probabilities corresponding to the model training data;
[0214] The first processing module 1702 is further configured to determine a loss function based on the emotion prediction probability;
[0215] The first processing module 1702 is further configured to create an emotion recognition model based on the loss function and the emotion joint network.
[0216] In this embodiment, a device 1700 for determining an emotion recognition model is proposed. A first processing module 1702 obtains model training data, wherein the model training data is data used for model training.
[0217] Exemplarily, the model training data may include voice training data and emotion type data.
[0218] The first processing module 1702 performs data training on the model training data to obtain an emotion joint network, wherein the emotion joint network is a deep learning network that can identify emotion types, and the emotion joint network stores the emotion prediction probability corresponding to the model training data, and the emotion prediction probability is the emotion prediction probability output by the emotion joint network.
[0219] Exemplarily, as shown in FIG2 , the emotion joint network may include an emotion grid. The emotion grid is constructed during the training phase, and each node in the emotion grid represents the probability distribution of emotions after outputting a number of texts in the current frame.
[0220] Exemplarily, the emotion prediction probability may include the target emotion posterior probability and the neutral emotion posterior probability, wherein the target emotion posterior probability is the posterior probability of emotions such as joy, anger, sadness, and happiness, and the neutral emotion posterior probability is the posterior probability of neutral emotion (a state without emotion).
[0221] The first processing module 1702 determines a loss function based on the emotion prediction probability, where the loss function is a function that limits the model output.
[0222] Exemplarily, the loss function is determined by the maximum pooling loss of the sentiment grid.
[0223] The first processing module 1702 creates an emotion recognition model according to the loss function and the emotion joint network.
[0224] Exemplarily, a recognition model including an emotion joint network is created, and the model output is optimized through a loss function to obtain an emotion recognition model, wherein the emotion recognition model is a model that can recognize the emotion type corresponding to the speech.
[0225] The emotion recognition model determination device 1700 in this embodiment greatly improves the recognition accuracy of the emotion recognition model, while also improving the recognition accuracy of the emotion recognition model for frame-level speech, thereby expanding the application scenarios and scope of the emotion recognition model.
[0226] In some embodiments, optionally, the emotion recognition model determination device 1700 further includes:
[0227] The first processing module 1702 is further configured to calculate a maximum value among the plurality of first probabilities to obtain a first target probability;
[0228] The first processing module 1702 is further configured to calculate a maximum value among the plurality of second probabilities to obtain a second target probability;
[0229] The first processing module 1702 is further configured to calculate a difference between the first target probability and the second target probability to obtain a loss function.
[0230] The device 1700 for determining the emotion recognition model in this embodiment calculates the maximum value of multiple first probabilities to obtain a first target probability, calculates the maximum value of multiple second probabilities to obtain a second target probability, and then calculates the difference between the first target probability and the second target probability to obtain a loss function, thereby improving the data accuracy of the loss function and thus improving the recognition accuracy of the emotion recognition model.
[0231] In some embodiments, optionally, the emotion recognition model determination device 1700 further includes:
[0232] The first processing module 1702 is further configured to create a data model based on the emotion joint network;
[0233] The first processing module 1702 is further configured to update the data model according to the loss function to obtain an emotion recognition model.
[0234] The emotion recognition model determination device 1700 in this embodiment creates a data model based on the emotion joint network, and then updates the data model according to the loss function to obtain an emotion recognition model, thereby improving the model accuracy of the emotion recognition model and further improving the recognition accuracy of the emotion recognition model.
[0235] As shown in FIG18 , an embodiment of the present application provides an emotion type recognition device 1800 , which includes:
[0236] The second processing module 1802 is used to obtain speech data and emotion recognition models;
[0237] The second processing module 1802 is further configured to determine first feature data corresponding to the speech data based on an emotion joint network of the emotion recognition model;
[0238] The second processing module 1802 is further configured to determine emotion type information corresponding to the voice data based on the first feature data.
[0239] In this embodiment, an emotion type recognition device 1800 is proposed, and the second processing module 1802 obtains voice data and an emotion recognition model, wherein the voice data is the voice data to be recognized, and the emotion recognition model is the emotion recognition model determined by the emotion recognition model determination method in the above embodiment.
[0240] Exemplarily, the voice data may be fine-grained frame-level voice data.
[0241] The second processing module 1802 determines first feature data corresponding to the speech data based on the emotion association network of the emotion recognition model, wherein the first feature data is feature data output by the emotion association network.
[0242] Exemplarily, the first feature data is used for fine-grained frame-level emotion prediction.
[0243] The second processing module 1802 determines emotion type information corresponding to the voice data based on the first feature data, wherein the emotion type information is used to indicate information about the emotion type corresponding to the voice data.
[0244] For example, the emotion type information may include labels such as joy, anger, sadness, or happiness.
[0245] The emotion type recognition device 1800 in this embodiment greatly improves the recognition accuracy of voice data through the emotion recognition model, improves the recognition accuracy of frame-level speech, ensures the accuracy of emotion type information, and expands the application scenarios and application scope of the emotion recognition model.
[0246] In some embodiments, optionally, the emotion type recognition device 1800 further includes:
[0247] The second processing module 1802 is further configured to input the speech data into a speech encoder to obtain speech features output by the speech encoder;
[0248] The second processing module 1802 is further configured to convert the speech data into text data based on the speech recognition module;
[0249] The second processing module 1802 is further configured to input the text data into the text predictor to obtain a first text feature output by the text predictor;
[0250] The second processing module 1802 is further configured to input the speech feature and the first text feature into the emotion joint network to obtain first feature data output by the emotion joint network.
[0251] The emotion type recognition device 1800 in this embodiment outputs speech features through a speech encoder, outputs first text features through a text predictor, and then inputs text data into the text predictor to obtain the first text features output by the text predictor, thereby ensuring the data accuracy of the first text features and further ensuring the information accuracy of the emotion type information.
[0252] In some embodiments, optionally, the emotion type recognition device 1800 further includes:
[0253] The second processing module 1802 is further configured to input the speech feature and the first text feature into the text joint network to obtain second feature data output by the text joint network;
[0254] The second processing module 1802 is further configured to determine vocabulary data corresponding to the speech data based on the second feature data.
[0255] The emotion type recognition device 1800 in this embodiment inputs the voice feature and the first text feature into the text joint network to obtain the second feature data output by the text joint network, and then determines the vocabulary data corresponding to the voice data based on the second feature data, thereby ensuring the data accuracy of the vocabulary data and further ensuring the information accuracy of the emotion type information.
[0256] In some embodiments, optionally, the emotion type recognition device 1800 further includes:
[0257] The second processing module 1802 is further configured to input the speech data into a speech encoder to obtain speech features output by the speech encoder;
[0258] The second processing module 1802 is further configured to convert the speech data into text data based on the speech recognition module;
[0259] The second processing module 1802 is further configured to input the text data into the symbol predictor to obtain a second text feature output by the symbol predictor;
[0260] The second processing module 1802 is further configured to input the speech feature and the second text feature into the emotion joint network to obtain first feature data output by the emotion joint network.
[0261] The emotion type recognition device 1800 in this embodiment inputs speech data into a speech encoder to obtain speech features output by the speech encoder, inputs text data into a symbol predictor to obtain second text features output by the symbol predictor, and then inputs the speech features and the second text features into an emotion joint network to obtain first feature data output by the emotion joint network, thereby ensuring the data accuracy of the first feature data and thus ensuring the information accuracy of the emotion type information.
[0262] In some embodiments, optionally, the emotion type recognition device 1800 further includes:
[0263] The second processing module 1802 is further configured to input the text data into the text predictor to obtain a first text feature output by the text predictor;
[0264] The second processing module 1802 is further configured to input the first text feature and the speech feature into the text joint network to obtain second feature data output by the text joint network;
[0265] The second processing module 1802 is further configured to input the speech feature and the second text feature into the symbol association network to obtain third feature data output by the symbol association network;
[0266] The second processing module 1802 is further configured to determine vocabulary data corresponding to the speech data based on the second feature data and the third feature data.
[0267] The emotion type recognition device 1800 in this embodiment obtains the second feature data output by the text joint network and the third feature data output by the symbol joint network, and then determines the vocabulary data corresponding to the speech data based on the second feature data and the third feature data, thereby ensuring the data accuracy of the vocabulary data and the information accuracy of the emotion type information.
[0268] In some embodiments, optionally, the emotion type recognition device 1800 further includes:
[0269] The second processing module 1802 is further configured to determine segmentation symbol information based on the third feature data;
[0270] The second processing module 1802 is further configured to convert the second feature data into vocabulary data according to the segmentation symbol information.
[0271] The emotion type recognition device 1800 in this embodiment determines the segmentation symbol information based on the third feature data, and then converts the second feature data into vocabulary data based on the segmentation symbol information, thereby ensuring the data accuracy of the vocabulary data and further ensuring the information accuracy of the emotion type information.
[0272] In some embodiments, optionally, the emotion type recognition device 1800 further includes:
[0273] The second processing module 1802 is further used to obtain the user's voice;
[0274] The second processing module 1802 is further configured to segment the user's voice according to a preset duration to obtain voice data.
[0275] The emotion type recognition device 1800 in this embodiment obtains voice data by segmenting the user's voice according to preset time lengths, thereby ensuring the data accuracy of the voice data and further ensuring the information accuracy of the emotion type information.
[0276] In some embodiments, optionally, the emotion type recognition device 1800 further includes:
[0277] The second processing module 1802 is further configured to generate an emotion tag corresponding to the speech data according to the emotion type information.
[0278] The emotion type recognition device 1800 in this embodiment generates an emotion label corresponding to the speech data according to the emotion type information, thereby ensuring the data accuracy of the emotion label.
[0279] In some embodiments, optionally, as shown in FIG. 19 , an apparatus 1900 for determining an emotion recognition model is proposed. The apparatus 1900 includes a processor 1902 and a memory 1904. The memory 1904 stores a program or instruction that, when executed by the processor 1902, implements the steps of the method for determining an emotion recognition model in any of the aforementioned technical solutions. Therefore, the apparatus 1900 for determining an emotion recognition model has all the benefits of the method for determining an emotion recognition model in any of the aforementioned technical solutions, and will not be further elaborated upon here.
[0280] In some embodiments, as shown in FIG20 , an emotion type recognition device 2000 is provided. The emotion type recognition device 2000 includes a processor 2002 and a memory 2004. The memory 2004 stores a program or instruction. When executed by the processor 2002, the program or instruction implements the steps of the emotion type recognition method described in any of the above technical solutions. Therefore, the emotion type recognition device 2000 has all the beneficial effects of the emotion type recognition method described in any of the above technical solutions, and will not be further described here.
[0281] In some embodiments, optionally, a readable storage medium is provided on which a program or instruction is stored. When the program or instruction is executed by a processor, the method for determining an emotion recognition model or the method for identifying an emotion type in any of the above embodiments is implemented, thereby having all the beneficial technical effects of the method for determining an emotion recognition model or the method for identifying an emotion type in any of the above embodiments.
[0282] The readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0283] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium can be, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: a portable computer floppy disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory card, a floppy disk, an encoding mechanical device (such as a punched card or a groove with a raised structure on which instructions are recorded), and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be understood as a transmission signal per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium, or electrical signals transmitted through wires.
[0284] In some embodiments, optionally, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement the method for determining an emotion recognition model in any of the above embodiments or the method for identifying an emotion type in any of the above embodiments, thereby having all the beneficial technical effects of the method for determining an emotion recognition model in any of the above embodiments or the method for identifying an emotion type in any of the above embodiments.
[0285] It should be clarified that in the claims, specification and drawings of this application, the term "plurality" refers to two or more. Unless otherwise clearly defined, the orientation or positional relationship indicated by the terms "upper" and "lower" is based on the orientation or positional relationship shown in the drawings. It is only for the purpose of more conveniently describing this application and making the description process simpler, and is not intended to indicate or imply that the device or element referred to must have the specific orientation described, be constructed and operated in a specific orientation. Therefore, these descriptions cannot be understood as limitations on this application. The terms "connect", "install", "fix" and the like should be understood in a broad sense. For example, "connection" can be a fixed connection between multiple objects, or a detachable connection between multiple objects, or an integral connection; it can be a direct connection between multiple objects, or an indirect connection between multiple objects through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood based on the specific circumstances of the above data.
[0286] In the claims, specification, and drawings of this application, the terms "one embodiment," "some embodiments," "a specific embodiment," and the like mean that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of this application. In the claims, specification, and drawings of this application, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0287] The above are merely preferred embodiments of the present application and are not intended to limit the present application. Those skilled in the art will readily appreciate that various modifications and variations are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A method for determining an emotion recognition model, wherein, The method for determining the emotion recognition model includes: Obtaining model training data; Performing data training on the model training data to obtain an emotion joint network, where the emotion joint network stores the emotion prediction probabilities corresponding to the model training data; Determining a loss function based on the emotion prediction probabilities; Creating an emotion recognition model according to the loss function and the emotion joint network.
2. The method for determining the emotion recognition model according to claim 1, wherein, The emotion prediction probabilities include multiple first probabilities and multiple second probabilities. Determining the loss function based on the emotion prediction probabilities includes: Calculating the maximum value among the multiple first probabilities to obtain a first target probability; Calculating the maximum value among the multiple second probabilities to obtain a second target probability; Calculating the difference between the first target probability and the second target probability to obtain the loss function.
3. The method for determining the emotion recognition model according to claim 1 or 2, wherein, Creating the emotion recognition model according to the loss function and the emotion joint network includes: Creating a data model based on the emotion joint network; Updating the data model according to the loss function to obtain the emotion recognition model.
4. A method for identifying an emotion type, wherein, The method for recognizing the emotion type includes: Obtaining speech data and an emotion recognition model, where the emotion recognition model is the emotion recognition model determined by the method for determining the emotion recognition model according to any one of claims 1 to 3; Determining first feature data corresponding to the speech data based on the emotion joint network of the emotion recognition model; Determining emotion type information corresponding to the speech data according to the first feature data.
5. The method for identifying an emotion type according to claim 4, wherein, The emotion recognition model further includes a speech encoder, a text predictor, and a speech recognition module. Determining the first feature data corresponding to the speech data based on the emotion joint network of the emotion recognition model includes: Inputting the speech data into the speech encoder to obtain speech features output by the speech encoder; Converting the speech data into text data based on the speech recognition module; Inputting the text data into the text predictor to obtain first text features output by the text predictor; Inputting the speech features and the first text features into the emotion joint network to obtain the first feature data output by the emotion joint network.
6. The method for identifying an emotion type according to claim 5, wherein, The emotion recognition model further includes a text joint network. The method for recognizing the emotion type further includes: Inputting the speech features and the first text features into the text joint network to obtain second feature data output by the text joint network; Determining vocabulary data corresponding to the speech data according to the second feature data.
7. The method for identifying the emotional type according to claim 4, wherein, The emotion recognition model further includes a symbol predictor, a speech encoder, and a speech recognition module. Determining the first feature data corresponding to the speech data based on the emotion joint network of the emotion recognition model includes: Inputting the speech data into the speech encoder to obtain speech features output by the speech encoder; Converting the speech data into text data based on the speech recognition module; Inputting the text data into the symbol predictor to obtain second text features output by the symbol predictor; Input the speech feature and the second text feature into the emotion joint network to obtain the first feature data output by the emotion joint network.
8. The method for identifying an emotion type according to claim 7, wherein, The emotion recognition model further includes a text joint network, a symbol joint network, and a text predictor. The method for recognizing the emotion type further includes: Input the text data into the text predictor to obtain the first text feature output by the text predictor. Input the first text feature and the speech feature into the text joint network to obtain the second feature data output by the text joint network. Input the speech feature and the second text feature into the symbol joint network to obtain the third feature data output by the symbol joint network. Determine the lexical data corresponding to the speech data according to the second feature data and the third feature data.
9. The method for identifying an emotion type according to claim 8, wherein, The determining the lexical data corresponding to the speech data according to the second feature data and the third feature data includes: Determine the segmentation symbol information according to the third feature data. Convert the second feature data into the lexical data according to the segmentation symbol information.
10. The method for identifying an emotion type according to any one of claims 4 to 9, wherein, The method for recognizing the emotion type further includes: Obtain the speech of the user. Segment the speech of the user according to a preset duration to obtain the speech data.
11. The method for identifying an emotion type according to any one of claims 4 to 9, wherein, After determining the emotion type information corresponding to the speech data according to the first feature data, the method for recognizing the emotion type further includes: Generate an emotion label corresponding to the speech data according to the emotion type information.
12. An apparatus for determining an emotion recognition model, wherein, The determining device of the emotion recognition model includes: A first processing module, configured to obtain model training data. The first processing module is further configured to perform data training on the model training data to obtain an emotion joint network, and the emotion joint network stores the emotion prediction probability corresponding to the model training data. The first processing module is further configured to determine a loss function based on the emotion prediction probability. The first processing module is further configured to create an emotion recognition model according to the loss function and the emotion joint network.
13. An apparatus for determining an emotion recognition model, wherein, Includes: A processor; A memory, in which a program or instruction is stored, and when the processor executes the program or instruction in the memory, the steps of the method for determining the emotion recognition model according to any one of claims 1 to 3 are implemented.
14. An identification device for an emotion type, wherein, The device for recognizing the emotion type includes: A second processing module, configured to obtain speech data and an emotion recognition model, and the emotion recognition model is an emotion recognition model determined by the method for determining the emotion recognition model according to any one of claims 1 to 3. The second processing module is further configured to determine the first feature data corresponding to the speech data based on the emotion joint network of the emotion recognition model. The second processing module is further configured to determine the emotion type information corresponding to the speech data according to the first feature data.
15. An apparatus for recognizing an emotion type, wherein, Includes: A processor; A memory, in which a program or instruction is stored, and when the processor executes the program or instruction in the memory, the steps of the method for recognizing the emotion type according to any one of claims 4 to 11 are implemented.
16. A readable storage medium, wherein, The program or instructions are stored on the readable storage medium, and when the program or instructions are executed by a processor, the steps of the method for determining an emotion recognition model according to any one of claims 1 to 3 or the method for recognizing an emotion type according to any one of claims 4 to 11 are implemented.
17. A computer program product, wherein, It includes computer instructions, and when the computer instructions are executed by a processor, the steps of the method for determining an emotion recognition model according to any one of claims 1 to 3 or the method for recognizing an emotion type according to any one of claims 4 to 11 are implemented.
Citation Information
Patent Citations
Audio and video multi-mode sentiment classification method and system
CN113408385A
Speech emotion recognition method and device based on multi-modal features and comparative learning
CN115240713A
Method for training emotion recognition model and emotion recognition method and device
CN115713797A
Emotion recognition method and device, computer readable storage medium and electronic equipment
CN116312486A
Speech emotion recognition method based on attention mechanism multi-scale feature extraction
CN116403609A
Cited By
Conversation method and device and related equipment
CN120910207A