A method and device for voice emotion recognition
By comprehensively identifying speech emotions by combining speech and text classification results, the problem of dependence on a large number of labeled data in the prior art is solved, and high-accuracy emotions recognition in the case of scarce data is achieved.
Patent Information
- Application Number
- CN202011583766.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-28
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2040-12-28
AI Technical Summary
The prior art relies on a large amount of speech and labeled data for training in speech emotion recognition, and the lack of quantitative standards leads to low accuracy of labeled data, thereby reducing the accuracy of emotion recognition results.
By comprehensively combining the speech classification results and text classification results, the emotions that are used to recognize speech are recognized are reduced, and the labeled data needs to rely solely on voice data for emotional classification are used. In the case of scarce training data, different levels of speech recognition methods are used to improve the accuracy of emotion recognition.
It effectively reduces the need for labeling data and improves the accuracy of emotion recognition. Especially when training data is scarce, it can still significantly improve the recognition effect.
Smart Images

Figure CN114694686B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular, to a method and device for speech emotion recognition. Background Art
[0002] In the mobile Internet era, users can communicate remotely through speech. During the process of remote communication, enhancing natural language processing (NLP) algorithms by recognizing and appropriately responding to speech content and emotions has become an important development direction for artificial intelligence (AI) systems.
[0003] Speech emotion recognition is a subfield within speech recognition, aiming to retrieve the emotion information lost during the speech-to-text conversion process. Currently, features can be constructed from speech, such as prosodic features or spectrum-based correlation features, etc., and then a classifier is trained using labeled training data. Here, the labeled data requires a human to listen to a piece of speech and then give the corresponding emotion type for that speech.
[0004] However, for classifying the emotions of speakers based on speech feature signals, a large amount of speech and labeled data are required to train the classifier. During the data annotation process, since there is no quantitative standard to distinguish whether it is "happy" or "sad", the accuracy of the labeled data is not high, resulting in a relatively low accuracy of the emotion recognition results output by the classifier. Summary of the Invention
[0005] Embodiments of this application provide a method and device for speech emotion recognition. By integrating the speech classification result and the text classification result, the emotion of the speech to be recognized is identified, which can not only reduce the labeled data for emotion classification relying solely on speech data, but also adopt different levels of speech recognition methods to improve the accuracy of emotion recognition even in the case of scarce training data.
[0006] In view of this, on the one hand, this application provides a method for speech emotion recognition, including:
[0007] Obtain the speech feature signal corresponding to the speech to be recognized;
[0008] Obtain the text to be recognized according to the speech feature signal;
[0009] Based on the speech feature signal, obtain a speech classification result through a speech classification model, where the speech classification result represents the degree of fluctuation of the speech to be recognized, the speech classification result is an excited type or a low type, and the degree of fluctuation of the low type is lower than that of the excited type;
[0010] Based on the text to be recognized, obtain the text classification result through a text classification model, where the text classification result represents the emotion type of the speech to be recognized;
[0011] Determine the emotion recognition result corresponding to the speech to be recognized according to the speech classification result and the text classification result.
[0012] On the other hand, this application provides a method for recognizing speech emotions, including:
[0013] Obtain an instant voice communication message;
[0014] In response to an operation of converting the message content of the instant voice communication message, display a text message containing an emoji corresponding to the instant voice communication message, where the emoji is determined by performing emotion recognition on the voice communication message.
[0015] On the other hand, this application provides a speech emotion recognition device, including:
[0016] An acquisition module, configured to acquire a speech feature signal corresponding to the speech to be recognized;
[0017] The acquisition module is further configured to acquire the text to be recognized according to the speech feature signal;
[0018] The acquisition module is further configured to, based on the speech feature signal, obtain a speech classification result through a speech classification model, where the speech classification result represents the fluctuation degree of the speech to be recognized, the speech classification result is an excited type or a low type, and the fluctuation degree of the low type is lower than that of the excited type;
[0019] The acquisition module is further configured to, based on the text to be recognized, obtain a text classification result through a text classification model, where the text classification result represents the emotion type of the speech to be recognized;
[0020] A determination module, configured to determine the emotion recognition result corresponding to the speech to be recognized according to the speech classification result and the text classification result.
[0021] In a possible design, in another implementation manner of the other aspect of the embodiments of this application,
[0022] The acquisition module is specifically configured to receive the speech to be recognized sent by a terminal device, where the speech to be recognized includes N frames of speech data, and N is an integer greater than or equal to 1;
[0023] Perform feature extraction processing on the speech to be recognized to obtain a speech feature signal, where the speech feature signal includes N signal features, and each signal feature in the speech feature signal corresponds to one frame of speech data.
[0024] In a possible design, in another implementation of another aspect of the embodiments of the present application, the speech to be recognized includes N frames of speech data, and the speech feature signal includes N signal features, each signal feature corresponding to one frame of speech data, where N is an integer greater than or equal to 1;
[0025] The obtaining module is specifically configured to obtain a speech classification result through a speech classification model based on the speech feature signal, including:
[0026] Based on the speech feature signal, obtain a target feature vector through the convolutional neural network included in the speech classification model, where the convolutional neural network includes a convolutional layer, a pooling layer, and a hidden layer;
[0027] Based on the target feature vector, obtain a target score through the temporal neural network included in the speech classification model;
[0028] Determine the speech classification result according to the target score.
[0029] In a possible design, in another implementation of another aspect of the embodiments of the present application,
[0030] The obtaining module is further configured to obtain a historical speech feature signal corresponding to historical speech, where the historical speech is a speech adjacent to the speech to be recognized before, the historical speech includes M frames of speech data, the historical speech feature signal includes M signal features, and each signal feature corresponds to one frame of speech data, where M is an integer greater than or equal to 1;
[0031] The obtaining module is further configured to obtain an intermediate feature vector through the convolutional neural network included in the speech classification model based on the historical speech feature signal, where the convolutional neural network includes a convolutional layer, a pooling layer, and a hidden layer;
[0032] The obtaining module is further configured to obtain a historical score through the temporal neural network included in the speech classification model based on the intermediate feature vector;
[0033] The determining module is specifically configured to determine the speech classification result according to the historical score and the target score.
[0034] In a possible design, in another implementation of another aspect of the embodiments of the present application, the speech emotion recognition device further includes a generating module;
[0035] The obtaining module is further configured to obtain P emoji, where the P emoji are emoji adjacent to the speech to be recognized before, or the P emoji are emoji adjacent to the speech to be recognized after, and P is an integer greater than or equal to 1;
[0036] The generating module is configured to generate a gain score according to the number of the P emoji;
[0037] An acquisition module, specifically configured to determine a voice classification result according to a gain score and a target score.
[0038] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,
[0039] The acquisition module is specifically configured to, if the target score is within the first score interval, determine that the voice classification result is an excited type;
[0040] If the target score is within the second score interval, determine that the voice classification result is a low type, where the fluctuation degree of the low type is lower than that of the excited type.
[0041] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,
[0042] The acquisition module is specifically configured to, based on the text to be recognized, obtain a text distribution probability through a text classification model, where the text distribution probability includes K first probability values, and each first probability value corresponds to a text type, and K is an integer greater than 1;
[0043] Determine a target probability value according to the text distribution probability;
[0044] Determine the text type corresponding to the target probability value as the text classification result.
[0045] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,
[0046] The acquisition module is further configured to obtain a historical voice feature signal corresponding to a historical voice, where the historical voice is a voice adjacent to and before the voice to be recognized, the historical voice includes M frames of voice data, the historical voice feature signal includes M signal features, and each signal feature corresponds to one frame of voice data, and M is an integer greater than or equal to 1;
[0047] The acquisition module is further configured to obtain a historical text to be recognized according to the historical voice feature signal;
[0048] The acquisition module is further configured to, based on the historical text to be recognized, obtain a historical text distribution probability through a text classification model, where the historical text distribution probability includes K second probability values, and each second probability value corresponds to a text type;
[0049] The acquisition module is specifically configured to generate an updated text distribution probability according to the text distribution probability and the historical text distribution probability;
[0050] Determine a target probability value according to the updated text distribution probability.
[0051] In a possible design, in another implementation of another aspect of the embodiments of the present application, the voice emotion recognition device further includes a generation module;
[0052] The acquisition module is further configured to acquire P emoji, where the P emoji are adjacent emoji that appear before the voice to be recognized, or the P emoji are adjacent emoji that appear after the voice to be recognized, and P is an integer greater than or equal to 1;
[0053] The generation module is configured to generate a gain text distribution probability according to the types of the P emoji;
[0054] The acquisition module is specifically configured to generate an updated text distribution probability according to the text distribution probability and the gain text distribution probability;
[0055] Determine a target probability value according to the updated text distribution probability.
[0056] In a possible design, in another implementation of another aspect of the embodiments of the present application,
[0057] The determination module is specifically configured to: if the voice classification result is an excited type and the text classification result is a happy text type, determine that the emotion recognition result corresponding to the voice to be recognized is a happy emotion type;
[0058] If the voice classification result is a low type and the text classification result is a happy text type, determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type.
[0059] In a possible design, in another implementation of another aspect of the embodiments of the present application,
[0060] The determination module is specifically configured to: if the voice classification result is an excited type and the text classification result is an angry text type, determine that the emotion recognition result corresponding to the voice to be recognized is an angry emotion type;
[0061] If the voice classification result is a low type and the text classification result is an angry text type, determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type.
[0062] In a possible design, in another implementation of another aspect of the embodiments of the present application,
[0063] The determination module is specifically configured to: if the voice classification result is an excited type and the text classification result is a sad text type, determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type;
[0064] If the voice classification result is a low type and the text classification result is a sad text type, then determine that the emotion recognition result corresponding to the voice to be recognized is a sad emotion type.
[0065] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,
[0066] The determination module is specifically configured to, if the voice classification result is an excited type and the text classification result is a neutral text type, determine that the emotion recognition result corresponding to the voice to be recognized is a non-emotion type;
[0067] If the voice classification result is a low type and the text classification result is a neutral text type, then determine that the emotion recognition result corresponding to the voice to be recognized is a non-emotion type.
[0068] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the voice emotion recognition device further includes a sending module;
[0069] The sending module is configured to, after the determination module determines the emotion recognition result corresponding to the voice to be recognized according to the voice classification result and the text classification result, if the emotion recognition result is a happy emotion type, send a first emoticon or a first prompt text to the terminal device, so that the terminal device displays the first emoticon or the first prompt text;
[0070] The sending module is further configured to, if the emotion recognition result is an angry emotion type, send a second emoticon or a second prompt text to the terminal device, so that the terminal device displays the second emoticon or the second prompt text;
[0071] The sending module is further configured to, if the emotion recognition result is a sad emotion type, send a third emoticon or a third prompt text to the terminal device, so that the terminal device displays the third emoticon or the third prompt text.
[0072] Another aspect of the present application provides a voice emotion recognition device, including:
[0073] An obtaining module, configured to obtain an instant voice communication message;
[0074] A display module, configured to display a text message including an emoticon corresponding to the instant voice communication message in response to a message content conversion operation on the instant voice communication message, where the emoticon is determined by performing emotion recognition on the voice communication message.
[0075] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,
[0076] A display module, specifically configured to obtain a voice feature signal corresponding to an instant voice communication message in response to a message content conversion operation on the instant voice communication message;
[0077] Obtain the text to be recognized according to the voice feature signal;
[0078] Based on the voice feature signal, obtain a voice classification result through a voice classification model, where the voice classification result represents the undulation degree of the instant voice communication message, the voice classification result is an excited type or a low type, and the undulation degree of the low type is lower than that of the excited type;
[0079] Based on the text to be recognized, obtain a text classification result through a text classification model, where the text classification result represents the emotion type of the instant voice communication message;
[0080] Determine the emotion recognition result corresponding to the instant voice communication message according to the voice classification result and the text classification result;
[0081] Generate a text message including an emoji corresponding to the instant voice communication message according to the emotion recognition result corresponding to the instant voice communication message;
[0082] Display the text message including the emoji corresponding to the instant voice communication message.
[0083] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the instant voice communication message includes N frames of voice data, the voice feature signal includes N signal features, each signal feature corresponds to one frame of voice data, and N is an integer greater than or equal to 1;
[0084] The display module is specifically configured to obtain a target feature vector through a convolutional neural network included in the voice classification model based on the voice feature signal, where the convolutional neural network includes a convolutional layer, a pooling layer, and a hidden layer;
[0085] Obtain a target score through a temporal neural network included in the voice classification model based on the target feature vector;
[0086] Determine the voice classification result according to the target score.
[0087] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,
[0088] The acquisition module is further configured to acquire P emojis, where the P emojis are adjacent emojis that appear before the instant voice communication message, or the P emojis are adjacent emojis that appear after the instant voice communication message, and P is an integer greater than or equal to 1;
[0089] The acquisition module is further configured to generate a gain score according to the number of P emojis;
[0090] The display module is specifically configured to determine a voice classification result according to the gain score and the target score.
[0091] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,
[0092] The display module is specifically configured to, if the voice classification result is an excited type and the text classification result is a happy text type, determine that the emotion recognition result corresponding to the voice to be recognized is a happy emotion type;
[0093] If the voice classification result is a low type and the text classification result is a happy text type, determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type;
[0094] If the voice classification result is an excited type and the text classification result is an angry text type, determine that the emotion recognition result corresponding to the voice to be recognized is an angry emotion type;
[0095] If the voice classification result is a low type and the text classification result is an angry text type, determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type;
[0096] If the voice classification result is an excited type and the text classification result is a sad text type, determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type;
[0097] If the voice classification result is a low type and the text classification result is a sad text type, determine that the emotion recognition result corresponding to the voice to be recognized is a sad emotion type;
[0098] If the voice classification result is an excited type and the text classification result is a neutral text type, determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type;
[0099] If the voice classification result is a low type and the text classification result is a neutral text type, determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type.
[0100] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,
[0101] The display module is specifically configured to, if the emotion recognition result is a happy emotion type, display a first emoji;
[0102] If the emotion recognition result is an angry emotion type, display a second emoji;
[0103] If the emotion recognition result is the sad emotion type, display the third emoji.
[0104] In a possible design, in another implementation of another aspect of the present application,
[0105] The obtaining module is further configured to, after the display module displays the text message including the emoji corresponding to the instant voice communication message in response to the message content conversion operation on the instant voice communication message, obtain the setting operation for the emoji;
[0106] The display module is further configured to display at least two selectable emojis in response to the setting operation for the emoji;
[0107] The obtaining module is further configured to obtain the selection operation for the target emoji;
[0108] The display module is further configured to display the text message including the target emoji corresponding to the instant voice communication message in response to the selection operation for the target emoji.
[0109] Another aspect of the present application provides a computer device, including: a memory, a processor, and a bus system;
[0110] Wherein, the memory is used to store programs;
[0111] The processor is used to execute the programs in the memory, and the processor is used to execute the methods in the above aspects according to the instructions in the program code;
[0112] The bus system is used to connect the memory and the processor to enable the memory and the processor to communicate.
[0113] Another aspect of the present application provides a computer-readable storage medium, in which instructions are stored, and when it runs on a computer, it enables the computer to execute the methods in the above aspects.
[0114] Another aspect of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above aspects.
[0115] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:
[0116] In an embodiment of the present application, a method for speech emotion recognition is provided. First, a speech feature signal corresponding to the speech to be recognized is obtained. Then, based on the speech feature signal, a speech classification result is obtained through a speech classification model, where the speech classification result represents the undulation degree of the speech to be recognized. And based on the text to be recognized, a text classification result is obtained through a text classification model, where the text classification result represents the emotion type of the speech to be recognized. Finally, combining the speech classification result and the text classification result, the emotion recognition result corresponding to the speech to be recognized is determined. By the above method, by comprehensively considering the speech classification result and the text classification result to recognize the emotion of the speech to be recognized, not only can the labeled data for emotion classification relying solely on speech data be reduced, but also different levels of speech recognition methods are adopted, which can still improve the accuracy of emotion recognition in the case of scarce training data. BRIEF DESCRIPTION OF THE DRAWINGS
[0117] Figure 1 It is a schematic architecture diagram of a speech emotion recognition system in an embodiment of the present application;
[0118] Figure 2 It is a schematic flowchart of a speech emotion recognition method in an embodiment of the present application;
[0119] Figure 3 It is a schematic diagram of an embodiment of a speech emotion recognition method in an embodiment of the present application;
[0120] Figure 4 It is a schematic diagram of an interface for user recording in an embodiment of the present application;
[0121] Figure 5 It is a schematic diagram of an interface for displaying emotion recognition results in an embodiment of the present application;
[0122] Figure 6 It is a schematic diagram of a network structure of a speech classification model in an embodiment of the present application;
[0123] Figure 7 It is a schematic diagram of an interface for displaying emotion recognition results based on historical speech in an embodiment of the present application;
[0124] Figure 8 It is a schematic diagram of an interface for displaying emotion recognition results based on emojis in an embodiment of the present application;
[0125] Figure 9 It is another schematic diagram of an interface for displaying emotion recognition results based on emojis in an embodiment of the present application;
[0126] Figure 10 It is a schematic diagram of a network structure of a text classification model in an embodiment of the present application;
[0127] Figure 11Schematic diagram of an interface for displaying emojis and prompt texts based on a happy mood type in an embodiment of the present application;
[0128] Figure 12 Schematic diagram of an interface for displaying emojis and prompt texts based on an angry mood type in an embodiment of the present application;
[0129] Figure 13 Schematic diagram of an interface for displaying emojis and prompt texts based on a sad mood type in an embodiment of the present application;
[0130] Figure 14 Schematic diagram of an interface for displaying speech recognition content based on a non-emotional mood type in an embodiment of the present application;
[0131] Figure 15 Schematic diagram of an embodiment of a speech emotion recognition method in an embodiment of the present application;
[0132] Figure 16 Schematic diagram of an embodiment of a speech emotion recognition device in an embodiment of the present application;
[0133] Figure 17 Schematic diagram of another embodiment of a speech emotion recognition device in an embodiment of the present application;
[0134] Figure 18 Schematic diagram of a structure of a server in an embodiment of the present application;
[0135] Figure 19 Schematic diagram of a structure of a terminal device in an embodiment of the present application. Detailed implementation manners
[0136] The embodiments of the present application provide a method and a device for speech emotion recognition, which comprehensively use the speech classification result and the text classification result to recognize the emotion of the speech to be recognized. This can not only reduce the labeled data for emotion classification relying solely on speech data, but also adopt different levels of speech recognition methods, so as to still improve the accuracy of emotion recognition in the case of scarce training data.
[0137] In the description and claims of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0138] With the research and progress of artificial intelligence (AI) technology, AI technology has been applied in more and more fields and plays an increasingly important role. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics.
[0139] Among them, speech technology is an important branch of artificial intelligence technology. The key technologies of speech technology include automatic speech recognition technology (ASR), text-to-speech technology, and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and speech has become one of the most promising human-computer interaction methods in the future.
[0140] Human speech contains a lot of information, including semantic information that a person wants to convey through speech, speaker identity information to which the speech belongs, speech information used by the speaker, and the speaker's emotional information. Speech emotion recognition refers to automatically identifying the emotions contained in the speech spoken by the speaker through a computer. The emotional information in speech is a very important behavioral signal reflecting human emotions. For the same speech content, when spoken with different emotions, the semantics carried may vary greatly. Therefore, accurately understanding the speaker's emotions can make human-computer interaction more natural and fluent.
[0141] In an application scenario, user A sends a piece of speech to user B. However, user B is not convenient to directly listen to this piece of speech. Therefore, the speech-to-text conversion function can be activated to convert the speech into text to be recognized. In order to enable user B to better understand the emotions of user A when speaking this piece of speech in the above scenario, this application provides a method for speech emotion recognition, which can display the speaker's speech emotions in the form of expressions. In this way, even if user B does not listen to the speech, they can understand the emotions of user A when speaking this piece of speech.
[0142] Based on this, the speech emotion recognition method provided in this application can be applied to Figure 1 the speech emotion recognition system shown in Figure 1 the speech emotion recognition system shown in the figure. As shown in the figure, the speech emotion recognition system includes a server and a terminal device, and the client is deployed on the terminal device. The server involved in this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device can be a smart phone, a tablet computer, a laptop computer, a handheld computer, a personal computer, a smart TV, a smart watch, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here. The number of servers and terminal devices is also not restricted.
[0143] Combined with Figure 1 the architecture shown in Figure 2 , Figure 2 is a schematic flowchart of the speech emotion recognition method in an embodiment of this application. As shown in the figure, specifically:
[0144] In step S1, user A sends a piece of speech through terminal device A. Terminal device A sends the speech to the server, and the server extracts the speech feature signal of the speech.
[0145] In step S2, the server inputs the extracted speech feature signal into the trained speech classification model.
[0146] In step S3, the speech classification result is output through the speech classification model, where the speech classification result represents the fluctuation degree of the speech to be recognized.
[0147] In step S4, the server performs speech recognition processing on the extracted speech feature signal, thereby generating the text to be recognized corresponding to the speech segment.
[0148] In step S5, the server inputs the text to be recognized into the trained text classification model.
[0149] In step S6, the text classification result is output through the text classification model, where the text classification result represents the emotion type of the speech to be recognized.
[0150] In step S7, based on the speech classification result and the text classification result, the emotion recognition result corresponding to the speech to be recognized is determined. Further, the server generates corresponding emojis or prompt texts according to the emotion recognition result, and then sends the speech recognition content to terminal device B. The speech recognition content includes the text to be recognized corresponding to the speech, and at least one of the emojis or prompt texts.
[0151] In view of the fact that this application involves some professional terms, therefore, these professional terms will be introduced separately below.
[0152] 1. Speech recognition: It is the technology that the machine transforms the speech signal into the corresponding text through the recognition and understanding process. Generally speaking, the recognized result is a piece of pure text.
[0153] 2. Speech feature: Through acoustic processing technology, the acoustic signal represented by continuous binary is converted into the feature represented by the feature vector.
[0154] 3. Speech emotion recognition (SER): The machine recognizes the emotion information contained in the speech signal through the speech signal, so that users can obtain the information contained in the speech more completely. Generally speaking, emotions can be divided into six major categories, namely "happy", "sad", "angry", "fear", "scared" and "disgust".
[0155] 4. Text classification: Text classification is similar to other classification tasks. After extracting features from a text segment, a classification category that best matches the features is selected. In this application, text classification refers to text sentiment classification, that is, selecting an emotion that best expresses a text segment.
[0156] 5. Convolutional neural network (CNN): It is a feedforward neural network. A convolutional neural network consists of one or more convolutional layers and a fully connected layer at the top (corresponding to a classic neural network), and can give better results in image and speech recognition.
[0157] 6. Long short-term memory (LSTM) artificial neural network: It is a variant of the Recurrent Neural Network (RNN). When processing a time series, RNN is prone to problems such as gradient explosion or gradient disappearance. LSTM solves this problem by adding a cell, and there are an input gate, a forget gate, and an output gate in the cell to determine which information from the previous moment can be input, forgotten, and output to the next moment.
[0158] 7. FastText: It is a neural network algorithm model for text classification. The model takes continuous text as input, can be trained unsupervised to obtain the embedding representation of the text, or can accept labeled training data for supervised training to classify the text.
[0159] Combined with the above introduction, the method for speech emotion recognition in this application will be introduced below. Please refer to Figure 3 , and an embodiment of the speech emotion recognition method in the embodiment of this application includes:
[0160] 101. Obtain the speech feature signal corresponding to the speech to be recognized;
[0161] In this embodiment, the speech emotion recognition device obtains the speech to be recognized sent by the user through the terminal device, and extracts the corresponding speech feature signal based on the speech to be recognized. Among them, the speech feature signal can be Mel Frequency Cepstrum Coefficient (MFCC) feature, Filter Bank (FBank) feature, Log Filter Bank (logfbank) feature, or Subband Spectrum Centroid (SSC) feature, etc.
[0162] For ease of understanding, please refer to Figure 4 , Figure 4 FIG. 1 is a schematic diagram of an interface for user recording in an embodiment of the present application. As shown in the figure, taking the voice sender as "User A" and the voice receiver as "User B" as an example, User A clicks on the interface to enter the conversation with "User B" in the instant messaging application. In this interface, a "Press and hold to talk" module is provided. When User A presses and holds this module, they can speak through the microphone, thereby inputting the voice to be recognized.
[0163] It should be noted that the voice emotion recognition device can be deployed on the server, or on the terminal device, or on a voice emotion recognition system composed of the server and the terminal device. In this application, it is introduced by taking the deployment on the server as an example, but this should not be construed as a limitation to this application.
[0164] 102. Obtain the text to be recognized according to the voice feature signal;
[0165] In this embodiment, the voice emotion recognition device can convert the voice feature signal into the corresponding text to be recognized. The text to be recognized is plain text. For example, "Let me tell you something. I'm very angry today. Really, very, very angry."
[0166] Specifically, voice recognition can also be called ASR, which is the process of converting sound into text. Voice recognition can use the Hidden Markov model (HMM) or Deep Neural Networks (DNN) to output the corresponding text to be recognized.
[0167] 103. Based on the voice feature signal, obtain the voice classification result through the voice classification model, where the voice classification result represents the fluctuation degree of the voice to be recognized. The voice classification result is an excited type or a low type, and the fluctuation degree of the low type is lower than that of the excited type;
[0168] In this embodiment, the voice emotion recognition device inputs the voice feature signal into the trained voice classification model, and outputs the voice classification result through the voice classification model. The voice classification result includes an excited type and a low type. Based on this, the voice classification result can represent the fluctuation degree of the voice to be recognized. If the fluctuation degree is large, it is the excited type; if the fluctuation degree is small, it is the low type.
[0169] Specifically, the voice feature signal is input into the voice classification model, and the target score is output by the voice classification model. Based on the target score, the voice classification result is determined. Among them, the target score can be distributed in a continuous interval. For example, in an interval from -1 to 1, the larger the target score, the greater the emotional fluctuation. If the target score is within the first score interval (such as [-1, 0]), the voice classification result is determined to be the excited type. If the target score is within the second score interval (such as [0, 1]), the voice classification result is determined to be the low type. When the target score is 0, it can be considered as the low type or the excited type, and no limit is set here.
[0170] Alternatively, the target score can be distributed in a discrete interval. For example, 1 or 0. If the target score is 1, the voice classification result is determined to be the excited type. If the target score is 0, the voice classification result is determined to be the low type.
[0171] 104. Based on the text to be recognized, the text classification result is obtained through the text classification model, where the text classification result represents the emotional type of the voice to be recognized.
[0172] In this embodiment, the voice emotion recognition device inputs the text to be recognized into the trained text classification model, and the text classification result is output through the text classification model. The text classification result includes the happy text type, the angry text type, the sad text type, and the neutral text type. Based on this, the text classification result can represent the emotional type of the voice to be recognized.
[0173] 105. According to the voice classification result and the text classification result, the emotion recognition result corresponding to the voice to be recognized is determined.
[0174] In this embodiment, the voice emotion recognition device can determine the emotion recognition result corresponding to the voice to be recognized according to the voice classification result and the text classification result. Combining the descriptions of the voice classification result and the text classification result in the foregoing steps, please refer to Table 1. Table 1 is a schematic diagram of the relationship between the emotion recognition result, the voice classification result, and the text classification result.
[0175] Table 1
[0176] Excited type Low type Happy text type Happy emotion type No emotion type Angry text type Angry emotion type No emotion type Sad text type No emotion type Sad emotion type Neutral text type No emotion type No emotion type
[0177] As can be seen from Table 1, based on the voice classification result and the text classification result, the corresponding emotion recognition result can be matched. Further, corresponding emoji or prompt text can be generated in combination with the emotion recognition result. For the convenience of introduction, please refer to Figure 5 , Figure 5 which is a schematic diagram of an interface for displaying the emotion recognition result in an embodiment of the present application. As shown in Figure 5As shown in Figure (A), after the server recognizes the voice to be recognized sent by User A, the text to be recognized obtained is "Let me tell you something. I'm very angry today. Really, very, very angry". Assuming that the voice classification result is the excited type and the text classification result is the angry text type, then the emotion recognition result is the angry emotion type. Thus, the voice recognition content obtained not only includes the text to be recognized, but also an emoji of "angry".
[0178] In the embodiments of the present application, a method for voice emotion recognition is provided. First, a voice feature signal corresponding to the voice to be recognized is obtained, and then based on the voice feature signal, a voice classification result is obtained through a voice classification model. Among them, the voice classification result represents the fluctuation degree of the voice to be recognized, and based on the text to be recognized, a text classification result is obtained through a text classification model. Among them, the text classification result represents the emotion type of the voice to be recognized. Finally, combining the voice classification result and the text classification result, the emotion recognition result corresponding to the voice to be recognized is determined. Through the above method, by comprehensively considering the voice classification result and the text classification result, the emotion of the voice to be recognized is recognized, which can not only reduce the labeled data for emotion classification relying only on voice data, but also adopt different levels of voice recognition methods, and can still improve the accuracy of emotion recognition in the case of scarce training data.
[0179] Optionally, on the basis of the above Figure 3 corresponding embodiment, in another optional embodiment provided by the embodiments of the present application, obtaining the voice feature signal corresponding to the voice to be recognized specifically includes:
[0180] Receiving the voice to be recognized sent by the terminal device, where the voice to be recognized includes N frames of voice data, and N is an integer greater than or equal to 1;
[0181] Performing feature extraction processing on the voice to be recognized to obtain a voice feature signal, where the voice feature signal includes N signal features, and each signal feature in the voice feature signal corresponds to one frame of voice data.
[0182] In this embodiment, a method for extracting a voice feature signal is introduced. After the server receives the voice to be recognized sent by the terminal device, frame division processing can be performed, that is, N frames of voice data are obtained. Each frame of voice data can be 20 milliseconds or 30 milliseconds, etc., which is not limited here. Feature extraction processing is performed on each frame of voice data to obtain signal features, and the N frame signal features corresponding to the N frames of voice data constitute the voice feature signal.
[0183] Specifically, taking the extraction of MFCC features of the speech to be recognized as an example, the speech to be recognized is continuous speech. First, the speech to be recognized can be pre-emphasized. Pre-emphasis can, to a certain extent, compensate for the loss of the high-frequency part and protect the integrity of the vocal tract information. Next, the pre-emphasized speech to be recognized is framed. Processing each frame of speech data after framing is equivalent to processing a continuous signal with fixed features, which can reduce the influence of non-steady time-varying. Discontinuities will occur at the start and end segments of each frame after framing, resulting in an increasing error from the original signal. Windowing can make the framed signal become relatively continuous, and generally, a Hamming window is selected.
[0184] After windowing, it is transformed to the frequency domain using the fast Fourier transform (FFT), and a spectrogram can be obtained after the transformation. Next, the absolute value or square value is taken, and then it is filtered using a Mel filter bank. Each filter of the Mel filter bank has a triangular filtering characteristic, and these filters are of equal bandwidth. The logarithm of the filtered signal is taken, and then a discrete cosine transform (DCT) is performed. A dimensionality reduction on the output after the DCT transformation can obtain the final MFCC features.
[0185] Secondly, in the embodiments of the present application, a method for extracting speech feature signals is provided. Through the above method, feature extraction is performed on the speech to be recognized, thereby enabling subsequent speech processing and improving the feasibility of the solution.
[0186] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiments of the present application based on the corresponding embodiment, the speech to be recognized includes N frames of speech data, the speech feature signal includes N signal features, each signal feature corresponds to one frame of speech data, and N is an integer greater than or equal to 1;
[0187] Based on the speech feature signal, a speech classification result is obtained through a speech classification model, specifically including:
[0188] Based on the speech feature signal, a target feature vector is obtained through the convolutional neural network included in the speech classification model, where the convolutional neural network includes a convolutional layer, a pooling layer, and a hidden layer;
[0189] Based on the target feature vector, a target score is obtained through the temporal neural network included in the speech classification model;
[0190] The speech classification result is determined according to the target score.
[0191] In this embodiment, a method for outputting a target score based on a voice classification model is introduced. As can be seen from the foregoing embodiments, the voice to be recognized includes N frames of voice data. After feature extraction is performed on each frame of voice data, N signal features can be obtained. The N signal features are input into a convolutional neural network, and the convolutional neural network outputs a target feature vector. The target feature vector is input into a temporal neural network, and the temporal neural network outputs a target score.
[0192] Specifically, for the sake of easy understanding, please refer to Figure 6 , Figure 6 which is a schematic diagram of a network structure of the voice classification model in the embodiment of the present application. As shown in the figure, the voice classification model includes two parts: a convolutional neural network and a temporal neural network. Among them, the convolutional neural network is a feedforward neural network, and the convolutional neural network includes a convolutional layer, a pooling layer, and a hidden layer. Among them, the convolutional layer can be used to extract features, the pooling layer can be used to select features, and the hidden layer can be used to output a feature vector. It is assumed that the convolutional neural network includes at least one convolutional network, and each convolutional network includes a convolutional layer and a pooling layer. After obtaining the N signal features, a central frame is taken out from them, that is, the signal feature Xt at the t-th moment. Considering the content relevance, L frames are extended to the left, and for example, R frames are extended to the right. Then the input feature sequence is [Xt-L, …, Xt, Xt+R], that is, each input is several signal features among the N signal features. Finally, a target feature vector is output through the hidden layer. Then, in chronological order, the target feature vector is sequentially input into the temporal neural network, and then the output of the feature vector corresponding to each frame of voice data is used as the input of the next frame. Finally, the output of the feature vector corresponding to the last frame of voice data is used as the input of the fully connected layer, and after passing through softmax, a target score in the range of [-1, 1] is output.
[0193] It can be understood that the larger the target score, the greater the emotional fluctuation. When the target score is "1", the voice classification result is "excited type", and when the target score is "0", the voice classification result is "low type". Optionally, the case where the target score is greater than 0 and less than or equal to 1 can also be determined as the "excited type", and the case where the target score is greater than or equal to -1 and less than 0 can be determined as the "low type". Optionally, other methods can also be used to determine whether the voice classification result belongs to the "excited type" or the "low type".
[0194] It can be understood that in addition to adopting the network structure combining CNN and LSTM, the voice classification model can also only adopt the CNN network structure or the LSTM network structure, or use a support vector machine (Support Vector Machine, SVM), etc., which is not limited here.
[0195] Secondly, in the embodiments of the present application, a method for outputting a target score based on a voice classification model is provided. Through the above method, the target feature vector of the voice feature signal can be extracted by using the CNN network included in the voice classification model, and the LSTM network included in the voice classification model can further perform temporal encoding on the target feature vector to incorporate the influence of time series on the predicted score, thereby improving the accuracy of score prediction.
[0196] Optionally, on the basis of the above Figure 3 corresponding embodiment, in another optional embodiment provided by the embodiments of the present application, it may further include:
[0197] Obtain the historical voice feature signal corresponding to the historical voice, where the historical voice is a voice adjacent to the voice to be recognized before, the historical voice includes M frames of voice data, the historical voice feature signal includes M signal features, and each signal feature corresponds to one frame of voice data, and M is an integer greater than or equal to 1;
[0198] Based on the historical voice feature signal, obtain an intermediate feature vector through the convolutional neural network included in the voice classification model, where the convolutional neural network includes a convolutional layer, a pooling layer, and a hidden layer;
[0199] Based on the intermediate feature vector, obtain the historical score through the temporal neural network included in the voice classification model;
[0200] Determining the voice classification result according to the target score may include:
[0201] Determine the voice classification result according to the historical score and the target score.
[0202] In this embodiment, a method for obtaining a voice classification result based on multiple voices is introduced. If the voice receiver also receives the historical voice sent by the same voice sender before the voice to be recognized, then in a similar way, the features of the historical voice are extracted, that is, the historical voice feature signal is obtained, where the historical voice feature signal includes M signal features, and each signal feature corresponds to one frame of voice data in the historical voice. Next, the historical voice feature signal is input into the trained voice classification model, and the historical score corresponding to the historical voice is output through the voice classification model. It should be noted that the method for the voice classification model to predict the historical score based on the historical voice feature signal is similar to the method for the voice classification model to predict the target score based on the voice feature signal, and will not be elaborated here.
[0203] For ease of understanding, the following will take one historical voice as an example for introduction. In practical applications, the corresponding historical scores can also be calculated for multiple historical voices respectively, which is not limited here. Please refer to Figure 7 , Figure 7This is a schematic diagram of an interface for displaying emotion recognition results based on historical speech in an embodiment of the present application. As shown in the figure, the 2-second speech in the figure is historical speech, and the 5-second speech is the speech to be recognized. First, obtain the historical speech feature signal corresponding to the historical speech, then input the historical speech feature signal into the speech classification model, and finally, the speech classification model outputs the historical score. Similarly, first, obtain the speech feature signal corresponding to the speech to be recognized, then input the speech feature signal into the speech classification model, and finally, the speech classification model outputs the target score.
[0204] After obtaining the historical score and the target score, the following three methods can be used to determine the speech classification result, which will be introduced below.
[0205] I. Determine the speech classification result based on the maximum and minimum values;
[0206] Taking the score of 0 as the dividing line, when it is greater than 0, determine the maximum value as the maximum and minimum value from the historical score and the target score. When it is less than 0, determine the minimum value as the maximum and minimum value from the historical score and the target score. Suppose the historical score is 0.8 and the target score is 1, then the maximum and minimum value is 1. At this time, the speech classification result is the "excited type". Suppose the historical score is -1 and the target score is 0.8, then the maximum and minimum value is -1. At this time, the speech classification result is the "low type".
[0207] II. Determine the speech classification result based on the average value;
[0208] Calculate the average value according to the historical score and the target score. Suppose the historical score is 0.8 and the target score is 1, then the maximum and minimum value is 0.9. It can be considered that the speech classification result is the "excited type".
[0209] III. Determine the speech classification result based on weight allocation;
[0210] Allocate the proportion of the historical score and the target score according to a certain weight ratio. Suppose the weight of the historical score is 0.2 and the weight of the target score is 0.8, and suppose the historical score is 0.8 and the target score is 1. From this, the calculated score is 0.2 * 0.8 + 0.8 * 1 = 0.96. It can be considered that the speech classification result is the "excited type".
[0211] Again, in the embodiment of the present application, a method for obtaining a speech classification result based on multiple speeches is provided. Through the above method, combined with the historical speeches sent by the user in the past period of time, the cumulative emotion information of the user in the past period of time can be obtained, that is, the historical score is obtained. The historical score and the target score are jointly used as the basis for determining the speech classification result. Therefore, it is beneficial to improve the accuracy of the speech classification result.
[0212] Optionally, in the above Figure 3Based on the corresponding embodiments, in another alternative embodiment provided by the embodiments of the present application, it may further include:
[0213] Obtain P emoji, where the P emoji are adjacent emoji that appear before the speech to be recognized, or the P emoji are adjacent emoji that appear after the speech to be recognized, and P is an integer greater than or equal to 1;
[0214] Generate a gain score according to the number of the P emoji;
[0215] Determining the speech classification result according to the target score may include:
[0216] Determine the speech classification result according to the gain score and the target score.
[0217] In this embodiment, a method for obtaining a speech classification result based on emoji is introduced. If the speech receiver also receives P emoji sent by the same speech sender before the speech to be recognized, the number of the P emoji is further obtained, that is, the value of P is determined. Among them, the types of the P emoji are all excited-type emoji, for example, the emoji of "laughing", the emoji of "crying", the emoji of "angry", etc. Therefore, the larger the value of P, the larger the gain score.
[0218] Exemplarily, please refer to Figure 8 , Figure 8 is a schematic diagram of an interface for displaying an emotion recognition result based on emoji in the embodiments of the present application. As shown in the figure, before user A sends the speech to be recognized, user A also sends 1 emoji of "angry". When the emoji of "angry" is recognized, a gain score is generated according to the number of 1 emoji. For example, the gain score is 0.1. In addition, obtain the speech feature signal corresponding to the speech to be recognized, then input the speech feature signal into the speech classification model, and finally, the speech classification model outputs the target score.
[0219] Exemplarily, please refer to Figure 9 , Figure 9 is another schematic diagram of an interface for displaying an emotion recognition result based on emoji in the embodiments of the present application. As shown in the figure, after user A sends the speech to be recognized, user A also sends 2 emojis of "angry". When the emoji of "angry" is recognized, a gain score is generated according to the number of 2 emojis. For example, the gain score is 0.2. In addition, obtain the speech feature signal corresponding to the speech to be recognized, then input the speech feature signal into the speech classification model, and finally, the speech classification model outputs the target score.
[0220] After obtaining the gain score and the target score, add the two together. If the sum of the scores is greater than 1, it is also recognized as a score of 1, that is, determine the voice classification result as the "excited type".
[0221] Again, in the embodiments of the present application, a method for obtaining the voice classification result based on emojis is provided. Through the above method, by combining the emojis sent by the user in the past period of time, the cumulative emotional information of the user in the past period of time can be obtained, that is, the gain score is obtained. The gain score and the target score are jointly used as the basis for determining the voice classification result. Therefore, it is beneficial to improve the accuracy of the voice classification result.
[0222] Optionally, on the basis of the above Figure 3 In another optional embodiment provided by the embodiments of the present application corresponding to the above embodiment, determining the voice classification result according to the target score may include:
[0223] If the target score is within the first score interval, determine the voice classification result as the excited type;
[0224] If the target score is within the second score interval, determine the voice classification result as the low type, where the fluctuation degree of the low type is lower than that of the excited type.
[0225] In this embodiment, a method for determining the voice classification result according to the target score is introduced. Taking the score interval of [-1, 1] as an example, the first score interval and the second score interval are set according to the score interval. For example, the first score interval is the interval greater than 0 and less than or equal to 1, and the second score interval is the interval greater than or equal to -1 and less than or equal to 0. Based on this, if the target score is within the first score interval, determine the voice classification result as the excited type, and if the target score is within the second score interval, determine the voice classification result as the low type.
[0226] It should be noted that the ranges of the first score interval and the second score interval can be adjusted according to the actual situation. The above example is only for illustration and should not be construed as a limitation of the present application.
[0227] Again, in the embodiments of the present application, a method for determining the voice classification result according to the target score is provided. Through the above method, according to the score interval where the target score is located, the voice classification result can be further determined based on the score interval where it is located, thereby improving the feasibility and operability of the solution.
[0228] Optionally, on the basis of the above Figure 3 In another optional embodiment provided by the embodiments of the present application corresponding to the above embodiment, obtaining the text classification result through the text classification model based on the text to be recognized may include:
[0229] Based on the text to be recognized, obtain the text distribution probability through a text classification model, where the text distribution probability includes K first probability values, and each first probability value corresponds to a text type, and K is an integer greater than 1;
[0230] Determine the target probability value according to the text distribution probability;
[0231] Determine the text type corresponding to the target probability value as the text classification result.
[0232] In this embodiment, a method for obtaining a text classification result based on a text classification model is introduced. After converting the speech to be recognized into text to be recognized, the text to be recognized can be input into the trained text classification model, and the text distribution probability is output through the text classification model, where the text distribution probability includes K first probability values. In this application, K can be set to 4, that is, the text distribution probability can be expressed as (a, b, c, d), and a + b + c + d = 1. Among them, a represents the first probability value corresponding to the happy text type, b represents the first probability value corresponding to the angry text type, c represents the first probability value corresponding to the sad text type, and d represents the first probability value corresponding to the neutral text type.
[0233] According to the text distribution probability (a, b, c, d), the maximum value can be selected as the target probability value. For example, if the text distribution probability is (0.8, 0.1, 0.05, 0.05), then the target probability value is 0.8. Therefore, the text type corresponding to the target probability value is determined as the text classification result, that is, the happy text type corresponding to the first probability value 0.8 is used as the text classification result.
[0234] Specifically, the text classification model can be a fasttext model. For ease of understanding, please refer to Figure 10 , Figure 10 is a schematic diagram of a network structure of the text classification model in the embodiment of the present application. As shown in the figure, x1, x2,... xT represent the n-gram vectors in the text to be recognized, and each feature is the average value of the word vectors. Here, all n-gram vectors are used to predict the specified category, that is, the text distribution probability is output.
[0235] It can be understood that in addition to the fasttext model, the text classification model can also adopt a CNN model or an LSTM model, or use a model from Bidirectional Encoder Representations from Transformers (BERT) model, etc., which is not limited here.
[0236] Secondly, in the embodiments of the present application, a method for obtaining a text classification result based on a text classification model is provided. Through the above method, the trained text classification model can be used to perform text classification on the text to be recognized, thereby improving the feasibility of the solution and being able to output a more accurate text classification result.
[0237] Optionally, based on the above Figure 3 corresponding embodiment, in another optional embodiment provided by the embodiments of the present application, it may further include:
[0238] Obtain a historical speech feature signal corresponding to a historical speech, where the historical speech is a speech adjacent to and before the speech to be recognized, the historical speech includes M frames of speech data, the historical speech feature signal includes M signal features, and each signal feature corresponds to one frame of speech data, and M is an integer greater than or equal to 1;
[0239] Obtain a historical text to be recognized according to the historical speech feature signal;
[0240] Based on the historical text to be recognized, obtain a historical text distribution probability through the text classification model, where the historical text distribution probability includes K second probability values, and each second probability value corresponds to a text type;
[0241] Determining the target probability value according to the text distribution probability may include:
[0242] Generate an updated text distribution probability according to the text distribution probability and the historical text distribution probability;
[0243] Determine the target probability value according to the updated text distribution probability.
[0244] In this embodiment, a method for obtaining a text classification result based on multiple speeches is introduced. If the speech receiver also receives a historical speech sent by the same speech sender before the speech to be recognized, then in a similar manner, the features of the historical speech are extracted, that is, the historical speech feature signal is obtained, where the historical speech feature signal includes M signal features, and each signal feature corresponds to one frame of speech data in the historical speech. Next, a historical text to be recognized is obtained according to the historical speech feature signal, and then the historical text to be recognized is input into the trained text classification model, and the historical text distribution probability is output through the text classification model. It should be noted that the prediction method of the text classification model for the historical text to be recognized is similar to the prediction method of the text classification model for the historical text to be recognized, and will not be elaborated here.
[0245] For ease of understanding, the following will be introduced by taking a historical voice as an example. In actual applications, the corresponding historical text distribution probabilities can also be calculated for multiple historical voices respectively, which is not limited here. First, obtain the historical voice feature signal corresponding to the historical voice, then obtain the historical text to be recognized according to the historical voice feature signal, and then input the historical text to be recognized into the text classification model. Finally, the text classification model outputs the historical text distribution probability, where the historical text distribution probability includes K second probability values. In this application, K can be set to 4, that is, the text distribution probability can be expressed as (x, y, z, r), and x + y + z + r = 1. Among them, x represents the second probability value corresponding to the happy text type, y represents the second probability value corresponding to the angry text type, z represents the second probability value corresponding to the sad text type, and r represents the second probability value corresponding to the neutral text type. Similarly, first, obtain the voice feature signal corresponding to the voice, then obtain the text to be recognized according to the voice feature signal, and then input the text to be recognized into the text classification model. Finally, the text classification model outputs the text distribution probability.
[0246] After obtaining the historical text distribution probability and the text distribution probability, the following two methods can be used to determine the text classification result, which will be introduced below.
[0247] I. Determine the text classification result based on the maximum value;
[0248] Suppose the historical text distribution probability is (0.7, 0.1, 0.1, 0.1), and the text distribution probability is (0.1, 0.8, 0.1, 0). Take the maximum value at the corresponding position in the historical text distribution probability and the text distribution probability as the updated element. Therefore, the updated text distribution probability is (0.7, 0.8, 0.1, 0.1). Then determine the target probability value as 0.8. Therefore, the text classification result is the angry text type.
[0249] Optionally, the updated text distribution probability can also be normalized so that the sum of all elements in the updated text distribution probability is 1.
[0250] II. Determine the text classification result based on the average value;
[0251] Suppose the historical text distribution probability is (0.7, 0.1, 0.1, 0.1), and the text distribution probability is (0.1, 0.8, 0.1, 0). Take the average value at the corresponding position in the historical text distribution probability and the text distribution probability as the updated element. Therefore, the updated text distribution probability is (0.4, 0.45, 0.1, 0.05). Then determine the target probability value as 0.45. Therefore, the text classification result is the angry text type.
[0252] Again, in the embodiments of the present application, a method for obtaining a text classification result based on multiple voices is provided. Through the above method, by combining the historical voices sent by the user in the past period of time, the emotional information accumulated by the user in the past period of time can be obtained, that is, the historical text distribution probability is obtained. The historical text distribution probability and the text distribution probability are jointly used as the basis for determining the text classification result. Therefore, it is beneficial to improve the accuracy of the text classification result.
[0253] Optionally, based on the above Figure 3 corresponding embodiment, in another optional embodiment provided by the embodiments of the present application, it may further include:
[0254] Obtain P emoji, where the P emoji are the adjacent emoji before the voice to be recognized, or the P emoji are the adjacent emoji after the voice to be recognized, and P is an integer greater than or equal to 1;
[0255] Generate a gain text distribution probability according to the types of the P emoji;
[0256] Determining the target probability value according to the text distribution probability may include:
[0257] Generate an updated text distribution probability according to the text distribution probability and the gain text distribution probability;
[0258] Determine the target probability value according to the updated text distribution probability.
[0259] In this embodiment, a method for obtaining a text classification result based on multiple emoji is introduced. If the voice receiver also receives P emoji sent by the same voice sender before the voice to be recognized, the types of the P emoji are further obtained, and the gain text distribution probability is determined according to the types of the P emoji. Among them, the types of the P emoji include a happy type (for example, the emoji of "laughing"), an angry type (for example, the emoji of "angry"), a sad type (for example, the emoji of "crying"), and a neutral type (for example, the emoji of "coffee" or the emoji of "computer").
[0260] Specifically, the gain text distribution probability includes K third probability values. In the present application, K can be set to 4, that is, the gain text distribution probability can be expressed as (e, f, g, h), and e + f + g + h = 1. Where e represents the third probability value corresponding to the happy text type, f represents the third probability value corresponding to the angry text type, g represents the third probability value corresponding to the sad text type, and h represents the third probability value corresponding to the neutral text type.
[0261] For P emoji, a certain probability value can be added to each corresponding type of emoji. For example, if 1 emoji of "laughing" is detected, the corresponding probability value is increased by 0.1, that is, the probability distribution of the augmented text is (0.1, 0, 0, 0). Another example, if 2 emojis of "crying" are detected, the corresponding probability value is increased by 0.3. That is, the probability distribution of the augmented text is (0, 0, 0.3, 0).
[0262] After obtaining the probability distribution of the augmented text and the probability distribution of the text, the elements at each corresponding position can be directly added. For example, the probability distribution of the text is (0.1, 0.8, 0.1, 0), and the probability distribution of the augmented text is (0.1, 0, 0, 0). Based on this, the updated probability distribution of the text is (0.2, 0.8, 0.1, 0). Then, the target probability value is determined to be 0.8. Therefore, the text classification result is the angry text type.
[0263] Again, in the embodiments of the present application, a method for obtaining a text classification result based on multiple emojis is provided. Through the above method, by combining the emojis sent by the user in the past period of time, the cumulative emotion information of the user in the past period of time can be obtained, that is, the probability distribution of the augmented text is obtained. The probability distribution of the augmented text and the probability distribution of the text are jointly used as the basis for determining the text classification result. Thus, it is beneficial to improve the accuracy of the text classification result.
[0264] Optionally, on the basis of the above Figure 3 In another optional embodiment provided by the embodiments of the present application corresponding to the embodiment, according to the voice classification result and the text classification result, determining the emotion recognition result corresponding to the voice to be recognized may include:
[0265] If the voice classification result is the excited type and the text classification result is the happy text type, then determine that the emotion recognition result corresponding to the voice to be recognized is the happy emotion type;
[0266] If the voice classification result is the low type and the text classification result is the happy text type, then determine that the emotion recognition result corresponding to the voice to be recognized is the emotionless type.
[0267] In this embodiment, a method for determining the emotion type by integrating the voice classification result and the text classification result is introduced. As can be seen from the foregoing embodiments, the voice classification result can be divided into the excited type and the low type. Based on this, if the text classification result is the happy text type and the voice classification result is the excited type, then by superimposing the happy text type and the excited type, it can be determined that the emotion recognition result corresponding to the voice to be recognized is the happy emotion type. If the text classification result is the happy text type and the voice classification result is the low type, then by superimposing the happy text type and the low type, it can be determined that the emotion recognition result corresponding to the voice to be recognized is the emotionless type.
[0268] Specifically, assume that after the speech to be recognized is recognized, the text to be recognized obtained is "It's really great to have rice noodle rolls when getting up in the morning". If the user says this sentence in an excited tone, then it is determined that the emotion recognition result corresponding to the speech to be recognized is the happy emotion type. If the user says this sentence in a low tone, then it is determined that the emotion recognition result corresponding to the speech to be recognized is the emotionless type.
[0269] Furthermore, in the embodiments of the present application, a method for determining the emotion type by integrating the speech classification result and the text classification result is provided. Through the above method, for the happy text type, it is also necessary to consider whether it belongs to the excited type. Only when both are met is the emotion recognition result considered to be the happy emotion type. Otherwise, it is not determined to be the happy emotion type. Adopting the "dual" determination can improve the accuracy of emotion recognition, thereby improving the reliability of the solution.
[0270] Optionally, on the basis of the above Figure 3 corresponding embodiment, in another optional embodiment provided by the embodiments of the present application, according to the speech classification result and the text classification result, determining the emotion recognition result corresponding to the speech to be recognized may include:
[0271] If the speech classification result is the excited type and the text classification result is the angry text type, then it is determined that the emotion recognition result corresponding to the speech to be recognized is the angry emotion type;
[0272] If the speech classification result is the low type and the text classification result is the angry text type, then it is determined that the emotion recognition result corresponding to the speech to be recognized is the emotionless type.
[0273] In this embodiment, a method for determining the emotion type by integrating the speech classification result and the text classification result is introduced. As can be seen from the foregoing embodiments, the speech classification result can be divided into the excited type and the low type. Based on this, if the text classification result is the angry text type and the speech classification result is the excited type, then by superimposing the angry text type and the excited type, it can be determined that the emotion recognition result corresponding to the speech to be recognized is the angry emotion type. If the text classification result is the angry text type and the speech classification result is the low type, then by superimposing the angry text type and the low type, it can be determined that the emotion recognition result corresponding to the speech to be recognized is the emotionless type.
[0274] Specifically, assume that after the speech to be recognized is recognized, the text to be recognized obtained is "Why don't you always reply to me?". If the user says this sentence in an excited tone, then it is determined that the emotion recognition result corresponding to the speech to be recognized is the angry emotion type. If the user says this sentence in a low tone, then it is determined that the emotion recognition result corresponding to the speech to be recognized is the emotionless type.
[0275] Furthermore, in the embodiments of the present application, a method for determining an emotion type by integrating a voice classification result and a text classification result is provided. Through the above method, for an angry text type, it is also necessary to consider whether it belongs to an excited type. Only when both conditions are met is the emotion recognition result considered to be the angry emotion type; otherwise, it is not determined to be the angry emotion type. Adopting a "dual" determination can improve the accuracy of emotion recognition, thereby enhancing the reliability of the solution.
[0276] Optionally, based on the corresponding embodiments above, in another optional embodiment provided by the embodiments of the present application, determining an emotion recognition result corresponding to the voice to be recognized according to the voice classification result and the text classification result may include: Figure 3 If the voice classification result is the excited type and the text classification result is the sad text type, then determine that the emotion recognition result corresponding to the voice to be recognized is the no-emotion type;
[0277] If the voice classification result is the low type and the text classification result is the sad text type, then determine that the emotion recognition result corresponding to the voice to be recognized is the sad emotion type.
[0278] In this embodiment, a method for determining an emotion type by integrating a voice classification result and a text classification result is introduced. As can be seen from the foregoing embodiments, the voice classification result can be divided into the excited type and the low type. Based on this, if the text classification result is the sad text type and the voice classification result is the excited type, then by superimposing the sad text type and the excited type, it can be determined that the emotion recognition result corresponding to the voice to be recognized is the no-emotion type. If the text classification result is the sad text type and the voice classification result is the low type, then by superimposing the sad text type and the low type, it can be determined that the emotion recognition result corresponding to the voice to be recognized is the sad emotion type.
[0279] Specifically, assume that after the voice to be recognized is recognized, the text to be recognized obtained is "I've been in a really bad mood lately". If the user says this sentence in an excited tone, then determine that the emotion recognition result corresponding to the voice to be recognized is the no-emotion type. If the user says this sentence in a low tone, then determine that the emotion recognition result corresponding to the voice to be recognized is the sad emotion type.
[0280] Furthermore, in the embodiments of the present application, a method for determining an emotion type by integrating a voice classification result and a text classification result is provided. Through the above method, for a sad text type, it is also necessary to consider whether it belongs to an excited type. Only when both conditions are met is the emotion recognition result considered to be the sad emotion type; otherwise, it is not determined to be the sad emotion type. Adopting a "dual" determination can improve the accuracy of emotion recognition, thereby enhancing the reliability of the solution.
[0281] Furthermore, in the embodiments of the present application, a method for determining an emotion type by integrating a voice classification result and a text classification result is provided. Through the above method, for a sad text type, it is also necessary to consider whether it belongs to an excited type. Only when both conditions are met is the emotion recognition result considered to be the sad emotion type; otherwise, it is not determined to be the sad emotion type. Adopting a "dual" determination can improve the accuracy of emotion recognition, thereby enhancing the reliability of the solution.
[0282] Optionally, based on the above Figure 3 corresponding embodiment, in another optional embodiment provided by the embodiments of the present application, according to the voice classification result and the text classification result, determining the emotion recognition result corresponding to the voice to be recognized may include:
[0283] If the voice classification result is an excited type and the text classification result is a neutral text type, then determine that the emotion recognition result corresponding to the voice to be recognized is a non-emotional type;
[0284] If the voice classification result is a low type and the text classification result is a neutral text type, then determine that the emotion recognition result corresponding to the voice to be recognized is a non-emotional type.
[0285] In this embodiment, a method for determining the emotion type by combining the voice classification result and the text classification result is introduced. As can be seen from the foregoing embodiments, the voice classification result can be divided into an excited type and a low type. Based on this, if the text classification result is a neutral text type, then regardless of whether the voice classification result is an excited type or a low type, it is determined that the emotion recognition result corresponding to the voice to be recognized is a non-emotional type.
[0286] Specifically, assume that after the voice to be recognized is recognized, the text to be recognized obtained is "I'm going to the art museum to see an exhibition this morning". Regardless of whether the user says this sentence in an excited tone or a low tone, it is determined that the emotion recognition result corresponding to the voice to be recognized is a non-emotional type.
[0287] Furthermore, in the embodiments of the present application, a method for determining the emotion type by combining the voice classification result and the text classification result is provided. Through the above method, for a neutral text type, regardless of whether it belongs to an excited type or a low type, it is determined as a non-emotional type. Using "dual" determination can improve the accuracy of emotion recognition, thereby improving the reliability of the solution.
[0288] Optionally, based on the above Figure 3 corresponding embodiment, in another optional embodiment provided by the embodiments of the present application, after determining the emotion recognition result corresponding to the voice to be recognized according to the voice classification result and the text classification result, it may further include:
[0289] If the emotion recognition result is a happy emotion type, send a first emoji or a first prompt text to the terminal device so that the terminal device displays the first emoji or the first prompt text;
[0290] If the emotion recognition result is an angry emotion type, send a second emoji or a second prompt text to the terminal device so that the terminal device displays the second emoji or the second prompt text;
[0291] If the emotion recognition result is the sad emotion type, send a third emoji or a third prompt text to the terminal device so that the terminal device displays the third emoji or the third prompt text.
[0292] In this embodiment, a method for generating corresponding information based on the emotion recognition result is introduced. As can be seen from the foregoing embodiments, if the voice emotion recognition device is deployed on the server, the server sends an emoji or a prompt text to the terminal device. If the voice emotion recognition device is deployed on the terminal device, the terminal device directly generates an emoji or a prompt text and displays the emoji or the prompt text.
[0293] For ease of understanding, please refer to Figure 11 , Figure 11 FIG. Figure 11 is a schematic diagram of an interface for displaying an emoji and a prompt text based on the happy emotion type in an embodiment of the present application. Assuming that the emotion recognition result is the happy emotion type, when the user triggers the voice-to-text function, the voice recognition content can be displayed. Exemplarily, as shown in FIG. Figure 11 (A), the voice recognition content includes the text to be recognized and a first emoji. The text to be recognized is "It's great to have delicious barbecue in the morning", and the first emoji is the emoji of "showing teeth". Exemplarily, as shown in FIG.
[0294] (B), the voice recognition content includes the text to be recognized and a first prompt text. The text to be recognized is "It's great to have delicious barbecue in the morning", and the first prompt text is "happy". Thus, the user can know that the emotion of the voice sender "User A" when saying this voice is happy. Figure 12 For ease of understanding, please refer to Figure 12 FIG. Figure 12 is a schematic diagram of an interface for displaying an emoji and a prompt text based on the angry emotion type in an embodiment of the present application. Assuming that the emotion recognition result is the angry emotion type, when the user triggers the voice-to-text function, the voice recognition content can be displayed. Exemplarily, as shown in FIG. Figure 12 (A), the voice recognition content includes the text to be recognized and a second emoji. The text to be recognized is "You don't reply to me when I talk to you. I'm so angry, hum", and the second emoji is the emoji of "angry". Exemplarily, as shown in FIG.
[0295] (B), the voice recognition content includes the text to be recognized and a second prompt text. The text to be recognized is "You don't reply to me when I talk to you. I'm so angry, hum", and the second prompt text is "angry". Thus, the user can know that the emotion of the voice sender "User A" when saying this voice is angry. Figure 13 For ease of understanding, please refer to Figure 13Schematic diagram of an interface for displaying emojis and prompt texts based on the sad emotion type in an embodiment of the present application. Assuming that the emotion recognition result is the sad emotion type, when the user triggers the voice-to-text function, the voice recognition content can be displayed. Exemplarily, as Figure 13 shown in Figure (A) in Figure 13 , the voice recognition content includes the text to be recognized and the third emoji. The text to be recognized is "Alas, for some reason, suddenly I feel a sense of loneliness", and the third emoji is the "sad" emoji. Exemplarily, as
[0296] shown in Figure (B) in Figure 14 , Figure 14 the voice recognition content includes the text to be recognized and the third prompt text. The text to be recognized is "Alas, for some reason, suddenly I feel a sense of loneliness", and the third prompt text is "sadness". Thus, the user can understand that the emotion of the voice sender "User A" when saying this voice is sadness.
[0297] Secondly, in an embodiment of the present application, a method for generating corresponding information based on the emotion recognition result is provided. Through the above method, corresponding feedback can be automatically generated for different emotion recognition results, such as generating emojis or generating prompt texts, etc. Thus, even if the voice receiver does not listen to the voice, they can understand the text content corresponding to the voice and the emotional state of the speaker, thereby improving the practicality and flexibility of the solution.
[0298] Combined with the above introduction, the voice emotion recognition application method in the present application will be introduced below. Please refer to Figure 15 , an embodiment of the emoji display method in an embodiment of the present application includes:
[0299] 201. The terminal device obtains an instant voice communication message;
[0300] In this embodiment, the terminal device obtains an instant voice communication message through an instant messaging application, where the instant voice communication message is presented as a voice message.
[0301] 202. The terminal device responds to the message content conversion operation of the instant voice communication message and displays a text message containing emojis corresponding to the instant voice communication message, where the emojis are determined by performing emotion recognition on the voice communication message.
[0302] In this embodiment, the terminal device receives a content conversion operation triggered by the user for an instant voice communication message. For example, the user clicks on the "Convert Voice to Text" module. Thus, the instant voice communication message is converted into text to be recognized, and a corresponding voice feature signal is extracted based on the instant voice communication message.
[0303] Specifically, the voice feature signal is input into a voice classification model, and thus a voice classification result is obtained. The voice classification result represents the degree of fluctuation of the voice to be recognized. The voice classification result is of an excited type or a low type, and the degree of fluctuation of the low type is lower than that of the excited type. The text to be recognized is input into a text classification model, and thus a text classification result is obtained. The text classification result represents the emotional type of the voice to be recognized. Finally, based on the voice classification result and the text classification result, an emotion recognition result corresponding to the voice to be recognized is determined. According to the emotion recognition result, a corresponding emoji is determined, and a text message containing the emoji is generated in combination with the text to be recognized.
[0304] It should be noted that the method of emotion recognition can refer to Figure 3 the corresponding embodiments, which will not be elaborated here.
[0305] In the embodiment of the present application, a method for applying voice emotion recognition is provided. Through the above method, by integrating the voice classification result and the text classification result, the emotion of the voice to be recognized is recognized. This can not only reduce the labeled data for emotion classification relying solely on voice data, but also adopt different levels of voice recognition methods, which can still improve the accuracy of emotion recognition in the case of scarce training data.
[0306] Optionally, on the basis of the above Figure 15 corresponding embodiment, in another optional embodiment provided by the embodiment of the present application, when the terminal device responds to a message content conversion operation on an instant voice communication message, it displays a text message containing an emoji corresponding thereto, which specifically includes:
[0307] When the terminal device responds to a message content conversion operation on an instant voice communication message, it obtains the voice feature signal corresponding to the instant voice communication message;
[0308] The terminal device obtains the text to be recognized according to the voice feature signal;
[0309] Based on the voice feature signal, the terminal device obtains a voice classification result through a voice classification model. The voice classification result represents the degree of fluctuation of the instant voice communication message. The voice classification result is of an excited type or a low type, and the degree of fluctuation of the low type is lower than that of the excited type;
[0310] The terminal device obtains a text classification result based on the text to be recognized through a text classification model, where the text classification result represents the emotion type of an instant voice communication message;
[0311] The terminal device determines an emotion recognition result corresponding to the instant voice communication message according to the voice classification result and the text classification result;
[0312] The terminal device generates a text message containing an emoji corresponding to the instant voice communication message according to the emotion recognition result corresponding to the instant voice communication message;
[0313] The terminal device displays the text message containing an emoji corresponding to the instant voice communication message.
[0314] Optionally, based on the above Figure 15 In another optional embodiment provided by the embodiment of the present application based on the corresponding embodiment, the instant voice communication message includes N frames of voice data, the voice feature signal includes N signal features, each signal feature corresponds to one frame of voice data, and N is an integer greater than or equal to 1;
[0315] The terminal device obtains a voice classification result based on the voice feature signal through a voice classification model, which may include:
[0316] The terminal device obtains a target feature vector based on the voice feature signal through the convolutional neural network included in the voice classification model, where the convolutional neural network includes a convolutional layer, a pooling layer, and a hidden layer;
[0317] The terminal device obtains a target score based on the target feature vector through the temporal neural network included in the voice classification model;
[0318] The terminal device determines the voice classification result according to the target score.
[0319] Optionally, based on the above Figure 15 In another optional embodiment provided by the embodiment of the present application based on the corresponding embodiment, it may further include:
[0320] The terminal device obtains P emojis, where the P emojis are the adjacent emojis before the instant voice communication message, or the P emojis are the adjacent emojis after the instant voice communication message, and P is an integer greater than or equal to 1;
[0321] The terminal device generates a gain score according to the number of the P emojis;
[0322] The terminal device determines the voice classification result according to the target score, including:
[0323] The terminal device determines the voice classification result according to the gain score and the target score.
[0324] Optionally, based on the above Figure 15 In another alternative embodiment provided by the embodiments of the present application based on the corresponding embodiment, the terminal device determines the emotion recognition result corresponding to the instant voice communication message according to the voice classification result and the text classification result, which may include:
[0325] If the voice classification result is an excited type and the text classification result is a happy text type, the terminal device determines that the emotion recognition result corresponding to the voice to be recognized is a happy emotion type;
[0326] If the voice classification result is a low type and the text classification result is a happy text type, the terminal device determines that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type;
[0327] If the voice classification result is an excited type and the text classification result is an angry text type, the terminal device determines that the emotion recognition result corresponding to the voice to be recognized is an angry emotion type;
[0328] If the voice classification result is a low type and the text classification result is an angry text type, the terminal device determines that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type;
[0329] If the voice classification result is an excited type and the text classification result is a sad text type, the terminal device determines that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type;
[0330] If the voice classification result is a low type and the text classification result is a sad text type, the terminal device determines that the emotion recognition result corresponding to the voice to be recognized is a sad emotion type;
[0331] If the voice classification result is an excited type and the text classification result is a neutral text type, the terminal device determines that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type;
[0332] If the voice classification result is a low type and the text classification result is a neutral text type, the terminal device determines that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type.
[0333] It should be noted that the manner of processing the instant voice communication message is similar to the manner of processing the voice to be recognized in the foregoing embodiment, so it will not be elaborated here.
[0334] Optionally, based on the above Figure 15 In another alternative embodiment provided by the embodiments of the present application based on the corresponding embodiment, the terminal device displays a text message containing an emoticon corresponding to the instant voice communication message, specifically including:
[0335] If the emotion recognition result is the happy emotion type, the terminal device displays the first emoji;
[0336] If the emotion recognition result is the angry emotion type, the terminal device displays the second emoji;
[0337] If the emotion recognition result is the sad emotion type, the terminal device displays the third emoji.
[0338] In this embodiment, a method of displaying corresponding emojis based on emotion recognition results is introduced. The terminal device can generate corresponding emojis based on the specific emotion type corresponding to the emotion recognition result. For ease of understanding, please refer to Figure 11 Figure (A) in. Suppose the instant voice communication message is "It's great to have a delicious barbecue when waking up in the morning". After recognition, a text message containing the first emoji corresponding to the instant voice communication message is displayed. The first emoji is the "grinning" emoji. Thus, the user can know that the emotion of the voice sender when speaking this voice is happy.
[0339] Please refer to Figure 12 Figure (A) in again. Suppose the instant voice communication message is "You don't reply to me when I talk to you. I'm so angry, hum". After recognition, a text message containing the second emoji corresponding to the instant voice communication message is displayed. The second emoji is the "angry" emoji. Thus, the user can know that the emotion of the voice sender when speaking this voice is angry.
[0340] Please refer to Figure 13 Figure (A) in again. Suppose the instant voice communication message is "Well, for some reason, suddenly I feel a sense of loneliness". After recognition, a text message containing the third emoji corresponding to the instant voice communication message is displayed. The third emoji is the "sad" emoji. Thus, the user can know that the emotion of the voice sender when speaking this voice is sad.
[0341] Secondly, in the embodiment of the present application, a method of displaying corresponding emojis based on emotion recognition results is provided. Through the above method, corresponding feedback can be automatically generated for different emotion recognition results. For example, generating emojis or generating prompt texts, etc. Thus, even if the voice receiver doesn't listen to the voice, they can understand the text content corresponding to the voice and the emotional state of the speaker, thereby improving the practicality and flexibility of the solution.
[0342] Optionally, on the basis of the above Figure 15 corresponding embodiment, in another optional embodiment provided by the embodiment of the present application, after the terminal device responds to the message content conversion operation of the instant voice communication message and displays the text message containing the emoji corresponding to the instant voice communication message, it may further include:
[0343] The terminal device obtains a setting operation for the emoji.
[0344] In response to the setting operation for the emoji, the terminal device displays at least two selectable emojis.
[0345] The terminal device obtains a selection operation for the target emoji.
[0346] In response to the selection operation for the target emoji, the terminal device displays a text message corresponding to the instant voice communication message and containing the target emoji.
[0347] In this embodiment, a way for users to customize emojis is provided. Users can also select a target emoji from at least two selectable emojis according to their own preferences or habits, etc. Based on this, the original emoji in the text message is updated to the target emoji.
[0348] Specifically, for example, the emoji is "angry". When the terminal device obtains a setting operation for the emoji, a selection box for emojis pops up, and at least two selectable emojis are displayed in the selection box. For example, angry emoji 1 and angry emoji 2. Assuming that the default "angry" emoji is angry emoji 1, after the user selects angry emoji 2, angry emoji 2 is displayed in the text message.
[0349] Secondly, in the embodiments of the present application, a way for users to customize emojis is provided. Through the above method, users can also select the emojis to be displayed in the text message according to their personal preferences, thereby improving the flexibility of the solution.
[0350] Combined with the introduction of the foregoing embodiments, the voice emotion recognition method provided by the present application can more accurately identify the emotion information carried in the voice. After experiments, the experimental data as Figure 2 shown is obtained. Please refer to Table 2.
[0351] Table 2
[0352] System classification method Pure voice classification Pure text classification Combined voice and text classification Accuracy rate 77% 74% 93%
[0353] As can be seen from Table 2, the accuracy rate of pure voice classification is 77%. This is the highest accuracy rate obtained so far. In theory, the effect can be further improved with more data, but the cost of obtaining labeled data for voice classification is very high. Limited by the cost, it is almost infeasible to achieve an accuracy of 90%. And, as mentioned in the foregoing embodiments, if using a "happy" tone to scold people, the accuracy of classification relying only on voice features is very low. In summary, the technical solution provided by the present application has a high accuracy rate at low cost.
[0354] The voice emotion recognition device in the present application will be described in detail below. Please refer to Figure 16 , Figure 16 FIG. 5 is a schematic diagram of an embodiment of the voice emotion recognition device in an embodiment of the present application. The voice emotion recognition device 30 includes:
[0355] An acquisition module 301, configured to acquire a voice feature signal corresponding to the voice to be recognized;
[0356] The acquisition module 301 is further configured to acquire the text to be recognized according to the voice feature signal;
[0357] The acquisition module 301 is further configured to, based on the voice feature signal, obtain a voice classification result through a voice classification model, where the voice classification result represents the undulation degree of the voice to be recognized, the voice classification result is an excited type or a low type, and the undulation degree of the low type is lower than that of the excited type;
[0358] The acquisition module 301 is further configured to, based on the text to be recognized, obtain a text classification result through a text classification model, where the text classification result represents the emotion type of the voice to be recognized;
[0359] A determination module 302, configured to determine an emotion recognition result corresponding to the voice to be recognized according to the voice classification result and the text classification result.
[0360] Optionally, based on the above Figure 16 corresponding embodiment, in another embodiment of the voice emotion recognition device 30 provided in the embodiment of the present application,
[0361] The acquisition module 301 is specifically configured to receive the voice to be recognized sent by the terminal device, where the voice to be recognized includes N frames of voice data, and N is an integer greater than or equal to 1;
[0362] Perform feature extraction processing on the voice to be recognized to obtain a voice feature signal, where the voice feature signal includes N signal features, and each signal feature in the voice feature signal corresponds to one frame of voice data.
[0363] Optionally, based on the above Figure 16 corresponding embodiment, in another embodiment of the voice emotion recognition device 30 provided in the embodiment of the present application, the voice to be recognized includes N frames of voice data, the voice feature signal includes N signal features, each signal feature corresponds to one frame of voice data, and N is an integer greater than or equal to 1;
[0364] The acquisition module 301 is specifically configured to obtain a voice classification result through a voice classification model based on the voice feature signal, including:
[0365] Based on the voice feature signal, obtain a target feature vector through the convolutional neural network included in the voice classification model, where the convolutional neural network includes a convolutional layer, a pooling layer, and a hidden layer;
[0366] Based on the target feature vector, obtain a target score through the time series neural network included in the voice classification model;
[0367] Determine the voice classification result according to the target score.
[0368] Optionally, based on the above Figure 16 In another embodiment of the voice emotion recognition device 30 provided in the embodiments of the present application, on the basis of the corresponding embodiment,
[0369] The acquisition module 301 is further configured to acquire a historical voice feature signal corresponding to a historical voice, where the historical voice is a voice adjacent to the voice to be recognized before, the historical voice includes M frames of voice data, the historical voice feature signal includes M signal features, and each signal feature corresponds to one frame of voice data, and M is an integer greater than or equal to 1;
[0370] The acquisition module 301 is further configured to, based on the historical voice feature signal, obtain an intermediate feature vector through the convolutional neural network included in the voice classification model, where the convolutional neural network includes a convolutional layer, a pooling layer, and a hidden layer;
[0371] The acquisition module 301 is further configured to, based on the intermediate feature vector, obtain a historical score through the time series neural network included in the voice classification model;
[0372] The determination module 302 is specifically configured to determine the voice classification result according to the historical score and the target score.
[0373] Optionally, based on the above Figure 16 In another embodiment of the voice emotion recognition device 30 provided in the embodiments of the present application, on the basis of the corresponding embodiment, the voice emotion recognition device 30 further includes a generation module 303;
[0374] The acquisition module 301 is further configured to acquire P emoji, where the P emoji are emoji adjacent to the voice to be recognized before, or the P emoji are emoji adjacent to the voice to be recognized after, and P is an integer greater than or equal to 1;
[0375] The generation module 303 is configured to generate a gain score according to the number of the P emoji;
[0376] The acquisition module 301 is specifically configured to determine the voice classification result according to the gain score and the target score.
[0377] Optionally, based on the above Figure 16Based on the corresponding embodiment, in another embodiment of the voice emotion recognition device 30 provided in the embodiments of the present application,
[0378] An acquisition module 301, specifically configured to determine that the voice classification result is an excited type if the target score is within the first score interval;
[0379] If the target score is within the second score interval, it is determined that the voice classification result is a low type, where the fluctuation degree of the low type is lower than that of the excited type.
[0380] Optionally, based on the corresponding embodiment above, in another embodiment of the voice emotion recognition device 30 provided in the embodiments of the present application, Figure 16 Based on the corresponding embodiment above, in another embodiment of the voice emotion recognition device 30 provided in the embodiments of the present application,
[0381] An acquisition module 301, specifically configured to obtain a text distribution probability through a text classification model based on the text to be recognized, where the text distribution probability includes K first probability values, and each first probability value corresponds to a text type, and K is an integer greater than 1;
[0382] Determine a target probability value according to the text distribution probability;
[0383] Determine the text type corresponding to the target probability value as the text classification result.
[0384] Optionally, based on the corresponding embodiment above, in another embodiment of the voice emotion recognition device 30 provided in the embodiments of the present application, Figure 16 Based on the corresponding embodiment above, in another embodiment of the voice emotion recognition device 30 provided in the embodiments of the present application,
[0385] The acquisition module 301 is further configured to obtain a historical voice feature signal corresponding to the historical voice, where the historical voice is a voice adjacent to the voice to be recognized before, the historical voice includes M frames of voice data, the historical voice feature signal includes M signal features, and each signal feature corresponds to one frame of voice data, and M is an integer greater than or equal to 1;
[0386] The acquisition module 301 is further configured to obtain a historical text to be recognized according to the historical voice feature signal;
[0387] The acquisition module 301 is further configured to obtain a historical text distribution probability through a text classification model based on the historical text to be recognized, where the historical text distribution probability includes K second probability values, and each second probability value corresponds to a text type;
[0388] The acquisition module 301 is specifically configured to generate an updated text distribution probability according to the text distribution probability and the historical text distribution probability;
[0389] Determine a target probability value according to the updated text distribution probability.
[0390] Optionally, based on the above Figure 16 In another embodiment of the voice emotion recognition device 30 provided by the embodiment of the present application, on the basis of the corresponding embodiment, the voice emotion recognition device 30 further includes a generation module 303;
[0391] The acquisition module 301 is further configured to acquire P emoji, where the P emoji are adjacent emoji that appear before the voice to be recognized, or the P emoji are adjacent emoji that appear after the voice to be recognized, and P is an integer greater than or equal to 1;
[0392] The generation module 303 is configured to generate a gain text distribution probability according to the types of the P emoji;
[0393] The acquisition module 301 is specifically configured to generate an updated text distribution probability according to the text distribution probability and the gain text distribution probability;
[0394] Determine a target probability value according to the updated text distribution probability.
[0395] Optionally, based on the above Figure 16 In another embodiment of the voice emotion recognition device 30 provided by the embodiment of the present application, on the basis of the corresponding embodiment,
[0396] The determination module 302 is specifically configured to, if the voice classification result is an excited type and the text classification result is a happy text type, determine that the emotion recognition result corresponding to the voice to be recognized is a happy emotion type;
[0397] If the voice classification result is a low type and the text classification result is a happy text type, determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type.
[0398] Optionally, based on the above Figure 16 In another embodiment of the voice emotion recognition device 30 provided by the embodiment of the present application, on the basis of the corresponding embodiment,
[0399] The determination module 302 is specifically configured to, if the voice classification result is an excited type and the text classification result is an angry text type, determine that the emotion recognition result corresponding to the voice to be recognized is an angry emotion type;
[0400] If the voice classification result is a low type and the text classification result is an angry text type, determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type.
[0401] Optionally, based on the above Figure 16 In another embodiment of the voice emotion recognition device 30 provided by the embodiment of the present application, on the basis of the corresponding embodiment,
[0402] A determination module 302, specifically configured to determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type if the voice classification result is an excited type and the text classification result is a sad text type;
[0403] If the voice classification result is a low type and the text classification result is a sad text type, then determine that the emotion recognition result corresponding to the voice to be recognized is a sad emotion type.
[0404] Optionally, based on the corresponding embodiment above, Figure 16 In another embodiment of the voice emotion recognition device 30 provided by the embodiments of the present application,
[0405] A determination module 302, specifically configured to determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type if the voice classification result is an excited type and the text classification result is a neutral text type;
[0406] If the voice classification result is a low type and the text classification result is a neutral text type, then determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type.
[0407] Optionally, based on the corresponding embodiment above, Figure 16 In another embodiment of the voice emotion recognition device 30 provided by the embodiments of the present application, the voice emotion recognition device 30 further includes a sending module 304;
[0408] The sending module 304 is configured to, after the determination module 302 determines the emotion recognition result corresponding to the voice to be recognized according to the voice classification result and the text classification result, if the emotion recognition result is a happy emotion type, send a first emoticon or a first prompt text to the terminal device so that the terminal device displays the first emoticon or the first prompt text;
[0409] The sending module 304 is further configured to, if the emotion recognition result is an angry emotion type, send a second emoticon or a second prompt text to the terminal device so that the terminal device displays the second emoticon or the second prompt text;
[0410] The sending module 304 is further configured to, if the emotion recognition result is a sad emotion type, send a third emoticon or a third prompt text to the terminal device so that the terminal device displays the third emoticon or the third prompt text.
[0411] The following will describe the voice emotion recognition device in the present application in detail. Please refer to Figure 17 , Figure 17 which is a schematic diagram of an embodiment of the voice emotion recognition device in the embodiments of the present application. The voice emotion recognition device 40 includes:
[0412] An acquisition module 401, configured to acquire instant voice communication messages;
[0413] A display module 402, configured to display a text message including an emoji corresponding to the instant voice communication message in response to a message content conversion operation on the instant voice communication message, where the emoji is determined by performing emotion recognition on the voice communication message.
[0414] Optionally, based on the corresponding embodiment above, Figure 17 In another embodiment of the voice emotion recognition device 40 provided in the embodiments of the present application,
[0415] An acquisition module 401, configured to acquire instant voice communication messages;
[0416] A display module 402, configured to display a text message including an emoji corresponding to the instant voice communication message in response to a message content conversion operation on the instant voice communication message, where the emoji is determined by performing emotion recognition on the voice communication message.
[0417] Optionally, based on the corresponding embodiment above, Figure 17 In another embodiment of the voice emotion recognition device 40 provided in the embodiments of the present application,
[0418] The display module 402 is specifically configured to, in response to a message content conversion operation on the instant voice communication message, acquire a voice feature signal corresponding to the instant voice communication message;
[0419] Obtain a text to be recognized according to the voice feature signal;
[0420] Based on the voice feature signal, obtain a voice classification result through a voice classification model, where the voice classification result represents the fluctuation degree of the instant voice communication message, the voice classification result is an excited type or a low type, and the fluctuation degree of the low type is lower than that of the excited type;
[0421] Based on the text to be recognized, obtain a text classification result through a text classification model, where the text classification result represents the emotion type of the instant voice communication message;
[0422] Determine an emotion recognition result corresponding to the instant voice communication message according to the voice classification result and the text classification result;
[0423] Generate a text message including an emoji corresponding to the instant voice communication message according to the emotion recognition result corresponding to the instant voice communication message;
[0424] Display a text message including an emoji corresponding to the instant voice communication message.
[0425] Optionally, based on the aboveFigure 17 Based on the corresponding embodiment, in another embodiment of the voice emotion recognition device 40 provided by the embodiments of the present application, the instant voice communication message includes N frames of voice data, the voice feature signal includes N signal features, each signal feature corresponds to one frame of voice data, and N is an integer greater than or equal to 1;
[0426] A display module 402, specifically configured to obtain a target feature vector through a convolutional neural network included in a voice classification model based on the voice feature signal, where the convolutional neural network includes a convolutional layer, a pooling layer, and a hidden layer;
[0427] Based on the target feature vector, obtain a target score through a temporal neural network included in the voice classification model;
[0428] Determine a voice classification result according to the target score.
[0429] Optionally, based on the corresponding embodiment above, in another embodiment of the voice emotion recognition device 40 provided by the embodiments of the present application, Figure 17 Based on the corresponding embodiment, in another embodiment of the voice emotion recognition device 40 provided by the embodiments of the present application,
[0430] An acquisition module 401 is further configured to acquire P emoji, where the P emoji are adjacent emoji that appear before the instant voice communication message, or the P emoji are adjacent emoji that appear after the instant voice communication message, and P is an integer greater than or equal to 1;
[0431] The acquisition module 401 is further configured to generate a gain score according to the number of the P emoji;
[0432] The display module 402 is specifically configured to determine a voice classification result according to the gain score and the target score.
[0433] Optionally, based on the corresponding embodiment above, in another embodiment of the voice emotion recognition device 40 provided by the embodiments of the present application, Figure 17 Based on the corresponding embodiment, in another embodiment of the voice emotion recognition device 40 provided by the embodiments of the present application,
[0434] The display module 402 is specifically configured to, if the voice classification result is an excited type and the text classification result is a happy text type, determine that the emotion recognition result corresponding to the voice to be recognized is a happy emotion type;
[0435] If the voice classification result is a low type and the text classification result is a happy text type, determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type;
[0436] If the voice classification result is an excited type and the text classification result is an angry text type, determine that the emotion recognition result corresponding to the voice to be recognized is an angry emotion type;
[0437] If the voice classification result is a low type and the text classification result is an angry text type, then determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type;
[0438] If the voice classification result is an excited type and the text classification result is a sad text type, then determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type;
[0439] If the voice classification result is a low type and the text classification result is a sad text type, then determine that the emotion recognition result corresponding to the voice to be recognized is a sad emotion type;
[0440] If the voice classification result is an excited type and the text classification result is a neutral text type, then determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type;
[0441] If the voice classification result is a low type and the text classification result is a neutral text type, then determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type.
[0442] Optionally, based on the above Figure 17 corresponding embodiment, in another embodiment of the voice emotion recognition device 40 provided by the embodiments of the present application,
[0443] The display module 402 is specifically configured to display a first emoticon if the emotion recognition result is a happy emotion type;
[0444] display a second emoticon if the emotion recognition result is an angry emotion type;
[0445] display a third emoticon if the emotion recognition result is a sad emotion type.
[0446] Optionally, based on the above Figure 17 corresponding embodiment, in another embodiment of the voice emotion recognition device 40 provided by the embodiments of the present application,
[0447] The acquisition module 401 is further configured to, after the display module 402 displays a text message containing an emoticon corresponding to an instant voice communication message in response to a message content conversion operation on the instant voice communication message, acquire a setting operation for the emoticon;
[0448] The display module 402 is further configured to display at least two optional emoticons in response to the setting operation for the emoticon;
[0449] The acquisition module 401 is further configured to acquire a selection operation for a target emoticon;
[0450] The display module 402 is further configured to display a text message including the target emoji corresponding to the instant voice communication message in response to a selection operation for the target emoji.
[0451] The voice emotion recognition device provided by this application can be deployed on a server. Please refer to Figure 18 , Figure 18 FIG. is a schematic structural diagram of a server provided by an embodiment of this application. The server 500 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 522 (for example, one or more processors) and a memory 532, and one or more storage media 530 (for example, one or more mass storage devices) storing application programs 542 or data 544. Among them, the memory 532 and the storage media 530 may be transient storage or persistent storage. The program stored in the storage media 530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 522 may be configured to communicate with the storage media 530 and execute a series of instruction operations in the storage media 530 on the server 500.
[0452] The server 500 may further include one or more power supplies 526, one or more wired or wireless network interfaces 550, one or more input / output interfaces 558, and / or one or more operating systems 541, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.
[0453] The steps performed by the server in the above embodiments may be based on the Figure 18 server structure shown.
[0454] The voice emotion recognition device provided by this application can be deployed on a terminal device. Please refer to Figure 19 , for the sake of convenience of description, only the parts related to the embodiments of this application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of this application. The terminal device may be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS) device, an in-vehicle computer, etc. Taking the terminal device as a mobile phone as an example:
[0455] Figure 19The block diagram of a partial structure of a mobile phone related to the terminal device provided by the embodiments of the present application is shown. Refer to Figure 19 , the mobile phone includes components such as a Radio Frequency (RF) circuit 610, a memory 620, an input unit 630, a display unit 640, a sensor 650, an audio circuit 650, a wireless fidelity (WiFi) module 670, a processor 680, and a power supply 690. Those skilled in the art can understand that Figure 19 the structure of the mobile phone shown in
[0456] does not limit the mobile phone, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. Figure 19 The following specifically introduces each component of the mobile phone:
[0457] The RF circuit 610 can be used for receiving and transmitting signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 680 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit 610 includes but is not limited to antennas, at least one amplifier, a transceiver, a coupler, a Low Noise Amplifier (LNA), a duplexer, etc. In addition, the RF circuit 610 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0458] The memory 620 can be used to store software programs and modules. The processor 680 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 620. The memory 620 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 620 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0459] The input unit 630 can be used to receive input digital or character information and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 630 may include a touch panel 631 and other input devices 632. The touch panel 631, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 631), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 631 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 680, and can receive and execute commands sent by the processor 680. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 631. In addition to the touch panel 631, the input unit 630 may further include other input devices 632. Specifically, the other input devices 632 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.
[0460] The display unit 640 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 640 may include a display panel 641. Optionally, the display panel 641 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 631 can cover the display panel 641. When the touch panel 631 detects a touch operation on or near it, it is transmitted to the processor 680 to determine the type of touch event. Subsequently, the processor 680 provides a corresponding visual output on the display panel 641 according to the type of touch event. Although in Figure 19 , the touch panel 631 and the display panel 641 are implemented as two independent components to realize the input and input functions of the mobile phone, but in some embodiments, the touch panel 631 and the display panel 641 can be integrated to realize the input and output functions of the mobile phone.
[0461] The mobile phone may further include at least one sensor 650, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 641 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 641 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscope, barometer, hygrometer, thermometer, infrared sensor that the mobile phone can also be configured with, they will not be elaborated here.
[0462] The audio circuit 650, the speaker 661, and the microphone 662 can provide an audio interface between the user and the mobile phone. The audio circuit 650 can transmit the electrical signal converted from the received audio data to the speaker 661, and the speaker 661 converts it into a sound signal for output; on the other hand, the microphone 662 converts the collected sound signal into an electrical signal, which is received by the audio circuit 650 and then converted into audio data. After the audio data is output to the processor 680 for processing, it is sent through the RF circuit 610 to, for example, another mobile phone, or the audio data is output to the memory 620 for further processing.
[0463] WiFi belongs to short-range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 670. It provides users with wireless broadband Internet access. Although Figure 19The WiFi module 670 is shown, but it can be understood that it does not belong to the essential components of the mobile phone and can be omitted entirely within the scope of not changing the essence of the invention as needed.
[0464] The processor 680 is the control center of the mobile phone, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 620, and by calling data stored in the memory 620, it executes various functions of the mobile phone and processes data, thereby monitoring the mobile phone as a whole. Optionally, the processor 680 may include one or more processing units; optionally, the processor 680 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 680 either.
[0465] The mobile phone also includes a power source 690 (such as a battery) for supplying power to each component. Optionally, the power source can be logically connected to the processor 680 through a power management system, thereby implementing functions such as management of charging, discharging, and power consumption management through the power management system.
[0466] Although not shown, the mobile phone may also include a camera, a Bluetooth module, etc., which will not be elaborated here.
[0467] The steps performed by the terminal device in the above embodiments can be based on the Figure 19 shown terminal device structure.
[0468] In the embodiments of the present application, a computer-readable storage medium is also provided. A computer program is stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the methods described in the foregoing various embodiments.
[0469] In the embodiments of the present application, a computer program product including a program is also provided. When it runs on a computer, it causes the computer to execute the methods described in the foregoing various embodiments.
[0470] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0471] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0472] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0473] In addition, the functional units in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0474] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, and other media that can store program codes.
[0475] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A method for speech emotion recognition, characterized in that, Including: Obtaining a voice feature signal corresponding to the voice to be recognized; Obtaining the text to be recognized according to the voice feature signal; Based on the voice feature signal, obtaining a voice classification result through a voice classification model, wherein the voice classification result represents the undulation degree of the voice to be recognized, the voice classification result is an excited type or a low type, and the undulation degree of the low type is lower than that of the excited type; Based on the text to be recognized, obtaining a text classification result through a text classification model, wherein the text classification result represents the emotion type of the voice to be recognized; Determining an emotion recognition result corresponding to the voice to be recognized according to the voice classification result and the text classification result; The obtaining the text classification result through the text classification model based on the text to be recognized includes: Based on the text to be recognized, obtaining a text distribution probability through the text classification model, wherein the text distribution probability includes K first probability values, and each first probability value corresponds to a text type, and K is an integer greater than 1; Determining a target probability value according to the text distribution probability; Determining the text type corresponding to the target probability value as the text classification result; Obtaining P emoji, wherein the P emoji are adjacent emoji that appear before the voice to be recognized, or the P emoji are adjacent emoji that appear after the voice to be recognized, and P is an integer greater than or equal to 1; Generating a gain text distribution probability according to the types of the P emoji; The determining the target probability value according to the text distribution probability includes: Generating an updated text distribution probability according to the text distribution probability and the gain text distribution probability; Determining the target probability value according to the updated text distribution probability; Obtaining a historical voice feature signal corresponding to a historical voice, wherein the historical voice is a voice adjacent to and before the voice to be recognized, the historical voice includes M frames of voice data, the historical voice feature signal includes M signal features, and each signal feature corresponds to one frame of voice data, and M is an integer greater than or equal to 1; Obtaining a historical text to be recognized according to the historical voice feature signal; Based on the historical text to be recognized, obtaining a historical text distribution probability through the text classification model, wherein the historical text distribution probability includes K second probability values, and each second probability value corresponds to a text type; The determining the target probability value according to the text distribution probability includes: Generating an updated text distribution probability according to the text distribution probability and the historical text distribution probability; Determining the target probability value according to the updated text distribution probability.
2. The method according to claim 1, characterized in that, The voice to be recognized includes N frames of voice data, the voice feature signal includes N signal features, and each signal feature corresponds to one frame of voice data, and N is an integer greater than or equal to 1; The obtaining the voice classification result through the voice classification model based on the voice feature signal includes: Based on the speech feature signal, obtain a target feature vector through the convolutional neural network included in the speech classification model, where the convolutional neural network includes a convolutional layer, a pooling layer, and a hidden layer; Based on the target feature vector, obtain a target score through the recurrent neural network included in the speech classification model; Determine the speech classification result according to the target score.
3. The method according to claim 2, characterized in that, The method further includes: Obtain a historical speech feature signal corresponding to a historical speech, where the historical speech is a speech adjacent to and before the speech to be recognized, the historical speech includes M frames of speech data, the historical speech feature signal includes M signal features, and each signal feature corresponds to one frame of speech data, and M is an integer greater than or equal to 1; Based on the historical speech feature signal, obtain an intermediate feature vector through the convolutional neural network included in the speech classification model, where the convolutional neural network includes a convolutional layer, a pooling layer, and a hidden layer; Based on the intermediate feature vector, obtain a historical score through the recurrent neural network included in the speech classification model; The step of determining the speech classification result according to the target score includes: Determine the speech classification result according to the historical score and the target score.
4. The method according to claim 2, characterized in that, The method further includes: Obtain P emoji, where the P emoji are emoji adjacent to and before the speech to be recognized, or the P emoji are emoji adjacent to and after the speech to be recognized, and P is an integer greater than or equal to 1; Generate a gain score according to the number of the P emoji; The step of determining the speech classification result according to the target score includes: Determine the speech classification result according to the gain score and the target score.
5. The method according to any one of claims 1 to 4, characterized in that, The step of determining the emotion recognition result corresponding to the speech to be recognized according to the speech classification result and the text classification result includes: If the speech classification result is an excited type and the text classification result is a happy text type, determine that the emotion recognition result corresponding to the speech to be recognized is a happy emotion type; If the speech classification result is a low type and the text classification result is a happy text type, determine that the emotion recognition result corresponding to the speech to be recognized is a no-emotion type; If the speech classification result is an excited type and the text classification result is an angry text type, determine that the emotion recognition result corresponding to the speech to be recognized is an angry emotion type; If the speech classification result is a low type and the text classification result is an angry text type, determine that the emotion recognition result corresponding to the speech to be recognized is a no-emotion type; If the speech classification result is an excited type and the text classification result is a sad text type, determine that the emotion recognition result corresponding to the speech to be recognized is a no-emotion type; If the speech classification result is a low type and the text classification result is a sad text type, determine that the emotion recognition result corresponding to the speech to be recognized is a sad emotion type; If the speech classification result is an excited type and the text classification result is a neutral text type, then determine that the emotion recognition result corresponding to the speech to be recognized is a non-emotional type; If the speech classification result is a low type and the text classification result is a neutral text type, then determine that the emotion recognition result corresponding to the speech to be recognized is a non-emotional type.
6. A method for applying voice emotion recognition, characterized in that, Including: Obtain instant voice communication messages; In response to an operation of converting the message content of the instant voice communication message, display a text message including an emoji corresponding to the instant voice communication message, where the emoji is determined by performing emotion recognition on the voice communication message; The step of, in response to an operation of converting the message content of the instant voice communication message, displaying a text message including an emoji corresponding to the instant voice communication message includes: In response to an operation of converting the message content of the instant voice communication message, obtain a voice feature signal corresponding to the instant voice communication message; Obtain a text to be recognized according to the voice feature signal; Based on the voice feature signal, obtain a speech classification result through a speech classification model, where the speech classification result represents the fluctuation degree of the instant voice communication message, the speech classification result is an excited type or a low type, and the fluctuation degree of the low type is lower than that of the excited type; Based on the text to be recognized, obtain a text classification result through a text classification model, where the text classification result represents the emotion type of the instant voice communication message; Determine the emotion recognition result corresponding to the instant voice communication message according to the speech classification result and the text classification result; Generate a text message including an emoji corresponding to the instant voice communication message according to the emotion recognition result corresponding to the instant voice communication message; The step of, based on the text to be recognized, obtaining a text classification result through a text classification model includes: Based on the text to be recognized, obtain a text distribution probability through the text classification model, where the text distribution probability includes K first probability values, and each first probability value corresponds to a text type, and K is an integer greater than 1; Determine a target probability value according to the text distribution probability; Determine the text type corresponding to the target probability value as the text classification result; Obtain P emojis, where the P emojis are the emojis adjacent to the speech to be recognized before, or the P emojis are the emojis adjacent to the speech to be recognized after, and P is an integer greater than or equal to 1; Generate a gain text distribution probability according to the types of the P emojis; The step of determining a target probability value according to the text distribution probability includes: Generate an updated text distribution probability according to the text distribution probability and the gain text distribution probability; Determine the target probability value according to the updated text distribution probability; Obtain the historical speech feature signal corresponding to the historical speech, where the historical speech is an adjacent speech before the speech to be recognized, the historical speech includes M frames of speech data, the historical speech feature signal includes M signal features, and each signal feature corresponds to one frame of speech data, and M is an integer greater than or equal to 1; Obtain the historical text to be recognized according to the historical speech feature signal; Based on the historical text to be recognized, obtain the historical text distribution probability through the text classification model, where the historical text distribution probability includes K second probability values, and each second probability value corresponds to a text type; The determining the target probability value according to the text distribution probability includes: Generate an updated text distribution probability according to the text distribution probability and the historical text distribution probability; Determine the target probability value according to the updated text distribution probability.
7. The voice emotion recognition application method according to claim 6, wherein, The instant voice communication message includes N frames of speech data, the speech feature signal includes N signal features, and each signal feature corresponds to one frame of speech data, and N is an integer greater than or equal to 1; The obtaining the speech classification result based on the speech feature signal through the speech classification model includes: Based on the speech feature signal, obtain the target feature vector through the convolutional neural network included in the speech classification model, where the convolutional neural network includes a convolutional layer, a pooling layer, and a hidden layer; Based on the target feature vector, obtain the target score through the temporal neural network included in the speech classification model; Determine the speech classification result according to the target score.
8. The voice emotion recognition application method according to claim 7, wherein, The method further includes: Obtain P emoji, where the P emoji are adjacent emoji before the instant voice communication message, or the P emoji are adjacent emoji after the instant voice communication message, and P is an integer greater than or equal to 1; Generate a gain score according to the number of the P emoji; The determining the speech classification result according to the target score includes: Determine the speech classification result according to the gain score and the target score.
9. The voice emotion recognition application method according to claim 8, wherein, The determining the emotion recognition result corresponding to the instant voice communication message according to the speech classification result and the text classification result includes: If the speech classification result is an excited type and the text classification result is a happy text type, then determine that the emotion recognition result corresponding to the speech to be recognized is a happy emotion type; If the speech classification result is a low type and the text classification result is a happy text type, then determine that the emotion recognition result corresponding to the speech to be recognized is a no-emotion type; If the speech classification result is an excited type and the text classification result is an angry text type, then determine that the emotion recognition result corresponding to the speech to be recognized is an angry emotion type; If the speech classification result is a low type and the text classification result is an angry text type, then determine that the emotion recognition result corresponding to the speech to be recognized is a no-emotion type; If the speech classification result is the excited type and the text classification result is the sad text type, then determine that the emotion recognition result corresponding to the speech to be recognized is the no-emotion type; If the speech classification result is the low type and the text classification result is the sad text type, then determine that the emotion recognition result corresponding to the speech to be recognized is the sad emotion type; If the speech classification result is the excited type and the text classification result is the neutral text type, then determine that the emotion recognition result corresponding to the speech to be recognized is the no-emotion type; If the speech classification result is the low type and the text classification result is the neutral text type, then determine that the emotion recognition result corresponding to the speech to be recognized is the no-emotion type.
10. The voice emotion recognition application method according to claim 9, wherein, The displaying of the text message including an emoji corresponding to the instant voice communication message includes: If the emotion recognition result is the happy emotion type, then display a first emoji; If the emotion recognition result is the angry emotion type, then display a second emoji; If the emotion recognition result is the sad emotion type, then display a third emoji.
11. The voice emotion recognition application method according to any one of claims 6 to 10, wherein, After the method displays the text message including an emoji corresponding to the instant voice communication message in response to a message content conversion operation on the instant voice communication message, the method further includes: Obtain a setting operation for the emoji; In response to the setting operation for the emoji, display at least two selectable emojis; Obtain a selection operation for a target emoji; In response to the selection operation for the target emoji, display the text message including the target emoji corresponding to the instant voice communication message.
12. A voice emotion recognition device, wherein, including: An obtaining module, configured to obtain a voice feature signal corresponding to the speech to be recognized; The obtaining module is further configured to obtain the text to be recognized according to the voice feature signal; The obtaining module is further configured to, based on the voice feature signal, obtain a speech classification result through a speech classification model, where the speech classification result represents the fluctuation degree of the speech to be recognized, the speech classification result is the excited type or the low type, and the fluctuation degree of the low type is lower than that of the excited type; The obtaining module is further configured to, based on the text to be recognized, obtain a text classification result through a text classification model, where the text classification result represents the emotion type of the speech to be recognized; A determining module, configured to determine the emotion recognition result corresponding to the speech to be recognized according to the speech classification result and the text classification result; The obtaining module is specifically configured to: Based on the text to be recognized, obtain a text distribution probability through the text classification model, where the text distribution probability includes K first probability values, and each first probability value corresponds to a text type, and K is an integer greater than 1; Generate an updated text distribution probability according to the text distribution probability and the gain text distribution probability; Determine a target probability value according to the updated text distribution probability; Determine the text type corresponding to the target probability value as the text classification result; Among them, the process of generating the probability distribution of the gain text includes: Obtain P emoji, where the P emoji are the emoji adjacent to the speech to be recognized before the speech to be recognized, or the P emoji are the emoji adjacent to the speech to be recognized after the speech to be recognized, and P is an integer greater than or equal to 1; Generate the probability distribution of the gain text according to the types of the P emoji; The device further includes: The obtaining module is further configured to obtain a historical speech feature signal corresponding to a historical speech, where the historical speech is a speech adjacent to the speech to be recognized before the speech to be recognized, the historical speech includes M frames of speech data, the historical speech feature signal includes M signal features, and each signal feature corresponds to one frame of speech data, and M is an integer greater than or equal to 1; The obtaining module is further configured to obtain a historical text to be recognized according to the historical speech feature signal; The obtaining module is further configured to obtain a historical text distribution probability based on the historical text to be recognized through the text classification model, where the historical text distribution probability includes K second probability values, and each second probability value corresponds to a text type; The obtaining module is specifically configured to generate an updated text distribution probability according to the text distribution probability and the historical text distribution probability; Determine the target probability value according to the updated text distribution probability.
13. The device according to claim 12, wherein, The speech to be recognized includes N frames of speech data, the speech feature signal includes N signal features, and each signal feature corresponds to one frame of speech data, and N is an integer greater than or equal to 1; The obtaining module is specifically configured to: Based on the speech feature signal, obtain a target feature vector through the convolutional neural network included in the speech classification model, where the convolutional neural network includes a convolutional layer, a pooling layer, and a hidden layer; Based on the target feature vector, obtain a target score through the time series neural network included in the speech classification model; Determine the speech classification result according to the target score.
14. The device according to claim 13, wherein, The obtaining module is further configured to obtain a historical speech feature signal corresponding to a historical speech, where the historical speech is a speech adjacent to the speech to be recognized before the speech to be recognized, the historical speech includes M frames of speech data, the historical speech feature signal includes M signal features, and each signal feature corresponds to one frame of speech data, and M is an integer greater than or equal to 1; The obtaining module is further configured to obtain an intermediate feature vector based on the historical speech feature signal through the convolutional neural network included in the speech classification model, where the convolutional neural network includes a convolutional layer, a pooling layer, and a hidden layer; The obtaining module is further configured to obtain a historical score based on the intermediate feature vector through the time series neural network included in the speech classification model; The determining module is specifically configured to determine the speech classification result according to the historical score and the target score.
15. The device according to claim 13, wherein, The device further includes a generating module; The obtaining module is further configured to obtain P emoji, where the P emoji are adjacent emoji that appear before the speech to be recognized, or the P emoji are adjacent emoji that appear after the speech to be recognized, and P is an integer greater than or equal to 1; The generating module is configured to generate a gain score according to the number of the P emoji; The obtaining module is specifically configured to determine the speech classification result according to the gain score and the target score.
16. The device according to any one of claims 12 - 15, wherein, The determining module is specifically configured to: If the speech classification result is an excited type and the text classification result is a happy text type, determine that the emotion recognition result corresponding to the speech to be recognized is a happy emotion type; If the speech classification result is a low type and the text classification result is a happy text type, determine that the emotion recognition result corresponding to the speech to be recognized is a no-emotion type; If the speech classification result is an excited type and the text classification result is an angry text type, determine that the emotion recognition result corresponding to the speech to be recognized is an angry emotion type; If the speech classification result is a low type and the text classification result is an angry text type, determine that the emotion recognition result corresponding to the speech to be recognized is a no-emotion type; If the speech classification result is an excited type and the text classification result is a sad text type, determine that the emotion recognition result corresponding to the speech to be recognized is a no-emotion type; If the speech classification result is a low type and the text classification result is a sad text type, determine that the emotion recognition result corresponding to the speech to be recognized is a sad emotion type; If the speech classification result is an excited type and the text classification result is a neutral text type, determine that the emotion recognition result corresponding to the speech to be recognized is a no-emotion type; If the speech classification result is a low type and the text classification result is a neutral text type, determine that the emotion recognition result corresponding to the speech to be recognized is a no-emotion type.
17. A voice emotion recognition device, characterized in that, Including: An obtaining module, configured to obtain an instant voice communication message; A display module, configured to display a text message including emoji corresponding to the instant voice communication message in response to a message content conversion operation on the instant voice communication message, where the emoji are determined by performing emotion recognition on the voice communication message; The display module is specifically configured to: In response to a message content conversion operation on the instant voice communication message, obtain a voice feature signal corresponding to the instant voice communication message; Obtain a text to be recognized according to the voice feature signal; Based on the voice feature signal, obtain a speech classification result through a speech classification model, where the speech classification result represents the fluctuation degree of the instant voice communication message, the speech classification result is an excited type or a low type, and the fluctuation degree of the low type is lower than that of the excited type; Based on the text to be recognized, obtain a text classification result through a text classification model, where the text classification result represents the emotion type of the instant voice communication message; Determine the emotion recognition result corresponding to the instant voice communication message according to the voice classification result and the text classification result; Generate a text message including emoji corresponding to the instant voice communication message according to the emotion recognition result corresponding to the instant voice communication message; The obtaining the text classification result by means of the text classification model based on the text to be recognized includes: Obtain the text distribution probability based on the text to be recognized by means of the text classification model, where the text distribution probability includes K first probability values, and each first probability value corresponds to a text type, and K is an integer greater than 1; Determine the target probability value according to the text distribution probability; Determine the text type corresponding to the target probability value as the text classification result; Obtain P emojis, where the P emojis are the emojis adjacent to the voice to be recognized before, or the P emojis are the emojis adjacent to the voice to be recognized after, and P is an integer greater than or equal to 1; Generate the gain text distribution probability according to the types of the P emojis; The determining the target probability value according to the text distribution probability includes: Generate the updated text distribution probability according to the text distribution probability and the gain text distribution probability; Determine the target probability value according to the updated text distribution probability; Obtain the historical voice feature signal corresponding to the historical voice, where the historical voice is a voice adjacent to the voice to be recognized before, the historical voice includes M frames of voice data, the historical voice feature signal includes M signal features, and each signal feature corresponds to one frame of voice data, and M is an integer greater than or equal to 1; Obtain the historical text to be recognized according to the historical voice feature signal; Obtain the historical text distribution probability based on the historical text to be recognized by means of the text classification model, where the historical text distribution probability includes K second probability values, and each second probability value corresponds to a text type; The determining the target probability value according to the text distribution probability includes: Generate the updated text distribution probability according to the text distribution probability and the historical text distribution probability; Determine the target probability value according to the updated text distribution probability.
18. The device according to claim 17, characterized in that, The instant voice communication message includes N frames of voice data, the voice feature signal includes N signal features, and each signal feature corresponds to one frame of voice data, and N is an integer greater than or equal to 1; The display module is specifically configured to obtain the target feature vector based on the voice feature signal by means of the convolutional neural network included in the voice classification model, where the convolutional neural network includes a convolutional layer, a pooling layer and a hidden layer; Obtain the target score based on the target feature vector by means of the time series neural network included in the voice classification model; Determine the voice classification result according to the target score.
19. The device according to claim 18, characterized in that, The obtaining module is further configured to obtain P emoji, where the P emoji are the adjacent emoji that appear before the instant voice communication message, or the P emoji are the adjacent emoji that appear after the instant voice communication message, and P is an integer greater than or equal to 1; The obtaining module is further configured to generate a gain score according to the number of the P emoji; The display module is specifically configured to determine the voice classification result according to the gain score and the target score.
20. The device according to claim 19, wherein, The display module is specifically configured to: If the voice classification result is an excited type and the text classification result is a happy text type, determine that the emotion recognition result corresponding to the voice to be recognized is a happy emotion type; If the voice classification result is a low type and the text classification result is a happy text type, determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type; If the voice classification result is an excited type and the text classification result is an angry text type, determine that the emotion recognition result corresponding to the voice to be recognized is an angry emotion type; If the voice classification result is a low type and the text classification result is an angry text type, determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type; If the voice classification result is an excited type and the text classification result is a sad text type, determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type; If the voice classification result is a low type and the text classification result is a sad text type, determine that the emotion recognition result corresponding to the voice to be recognized is a sad emotion type; If the voice classification result is an excited type and the text classification result is a neutral text type, determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type; If the voice classification result is a low type and the text classification result is a neutral text type, determine that the emotion recognition result corresponding to the voice to be recognized is a no-emotion type.
21. The device according to claim 20, wherein, The display module is specifically configured to: If the emotion recognition result is the happy emotion type, display a first emoji; If the emotion recognition result is the angry emotion type, display a second emoji; If the emotion recognition result is the sad emotion type, display a third emoji.
22. The device according to any one of claims 17 - 21, wherein, The obtaining module is further configured to, after responding to the message content conversion operation on the instant voice communication message and displaying the text message including the emoji corresponding to the instant voice communication message, obtain the setting operation for the emoji; The display module is further configured to respond to the setting operation for the emoji and display at least two selectable emoji; The obtaining module is further configured to obtain the selection operation for the target emoji; The display module is further configured to respond to the selection operation for the target emoji and display the text message including the target emoji corresponding to the instant voice communication message.
23. A computer device, wherein, including: a memory, a processor, and a bus system; Wherein, the memory is used to store programs; The processor is used to execute the programs in the memory, and the processor is used to execute the method according to any one of claims 1 to 11 according to the instructions in the program code; The bus system is used to connect the memory and the processor, so that the memory and the processor can communicate with each other.
24. A computer-readable storage medium comprising instructions which, when run on a computer, cause the computer to execute the method according to any one of claims 1 to 11.
25. A computer program product, wherein, The computer program product includes instructions, and when the instructions are run on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Intelligent interaction method and device, computer equipment and computer readable storage medium
CN108197115A
Classification model training method, classification method and device, equipment and medium
CN109684478A
Service evaluation obtaining method and device based on bimodal emotion recognition network
CN111563422A
Method for auto interpreting using emoticon and apparatus using the same
KR1020160138613A
KR20200082232A