Dialogue text emotion recognition method and system, electronic equipment and storage medium

By tuning multiple basic models parameters and selecting high accuracy models to build a dialogue text emotion recognition model, the problem of low emotional recognition accuracy in dialogue text in the prior art is solved, and higher emotional recognition accuracy is achieved.

CN120216692APending Publication Date: 2025-06-27UNICOM WOYUEDU TECH CULTURE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510214449.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing dialogue text emotion recognition methods often fail to correctly identify the true emotions of dialogue text during the emotional recognition process, resulting in a relatively low accuracy of emotion recognition.

Method used

By obtaining the dialogue text to be identified and the dialogue text data set with emotion labels, dividing the training data set and verification data set, presetting multiple basic models and tuning parameters, training and verifying these models, selecting models with high accuracy to build a dialogue text emotion recognition model.

Benefits of technology

By focusing on different features and combining the advantages of multiple models, we ensure that the true emotions of the dialogue text can be correctly identified and the accuracy of emotional recognition of dialogue text is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216692A_ABST
    Figure CN120216692A_ABST
Patent Text Reader

Abstract

The invention discloses a dialogue text emotion recognition method and system, electronic equipment and a storage medium. The method comprises the following steps: performing parameter tuning on a plurality of basic models to obtain a plurality of tuned basic models; training each adjusted and optimized basic model to obtain a plurality of trained adjusted and optimized basic models; verifying each trained optimized basic model for multiple times to obtain a plurality of verification results corresponding to each optimized basic model; according to a plurality of verification results corresponding to each adjusted and optimized basic model, calculating the accuracy rate of each adjusted and optimized basic model; based on the accuracy rate of each adjusted and optimized basic model, selecting a plurality of adjusted and optimized basic models ranked in the top; constructing a dialogue text emotion recognition model according to the plurality of optimized basic models ranked in the top; and performing emotion recognition on the to-be-recognized dialogue text through the dialogue text emotion recognition model to obtain an emotion recognition result. According to the method, the dialogue text emotion recognition accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of text sentiment recognition, and in particular, to a method, a system, an electronic device, and a storage medium for recognizing the sentiment of dialogue text. Background Art

[0002] The development of speech recognition technology has gone through several decades. From the initial simple command recognition to the current intelligent speech assistants that support multiple languages and adapt to complex environments, speech recognition has become an indispensable part of our daily lives. Converting speech into text can bring convenience to many fields. For example, for dialogue sentiment recognition, after converting a speech dialogue into dialogue text, it is possible to better recognize the sentiment of the speech dialogue.

[0003] Existing methods for recognizing the sentiment of dialogue text rely on the learning degree of the constructed model and also on the richness, diversity, and correctness of the dataset. However, different features are concerned, and different models have different advantages. Using only one model for prediction cannot well guarantee the accuracy of the prediction. Therefore, existing methods for recognizing the sentiment of dialogue text often cannot correctly recognize the true sentiment of the dialogue text during the sentiment recognition process, resulting in a relatively low sentiment recognition accuracy. Summary of the Invention

[0004] The present application aims to propose a method, a system, an electronic device, and a storage medium for recognizing the sentiment of dialogue text, which can improve the accuracy of recognizing the sentiment of dialogue text.

[0005] In a first aspect, an embodiment of the present application provides a method for recognizing the sentiment of dialogue text, the method including:

[0006] Obtain a dialogue text to be recognized and a dialogue text dataset with sentiment labels, and divide the dialogue text dataset into a dialogue text training dataset and a dialogue text validation dataset;

[0007] Preset a plurality of basic models, and perform parameter tuning on the plurality of basic models to obtain a plurality of tuned basic models;

[0008] Train each of the tuned basic models through the dialogue text training dataset to obtain a plurality of trained tuned basic models;

[0009] Perform multiple validations on each of the trained tuned basic models through the dialogue text validation dataset to obtain a plurality of validation results corresponding to each of the tuned basic models;

[0010] Calculate the accuracy rate of each of the tuned basic models according to the plurality of validation results corresponding to each of the tuned basic models;

[0011] Select a number of top-ranked tuned base models based on the accuracy of each of the tuned base models;

[0012] Construct a dialogue text sentiment recognition model based on a number of top-ranked tuned base models;

[0013] Perform sentiment recognition on the dialogue text to be recognized through the dialogue text sentiment recognition model to obtain a sentiment recognition result.

[0014] Compared with the prior art, the first aspect of the present application has the following beneficial effects:

[0015] In this method, a dialogue text dataset with sentiment labels is obtained together with the dialogue text to be recognized, and the dialogue text dataset is divided into a dialogue text training dataset and a dialogue text validation dataset; multiple base models are preset, and parameter tuning is performed on the multiple base models to obtain multiple tuned base models; each tuned base model is trained through the dialogue text training dataset to obtain multiple trained tuned base models; each trained tuned base model is verified multiple times through the dialogue text validation dataset to obtain multiple verification results corresponding to each tuned base model; the accuracy of each tuned base model is calculated according to the multiple verification results corresponding to each tuned base model; a number of top-ranked tuned base models are selected based on the accuracy of each tuned base model; a dialogue text sentiment recognition model is constructed based on the number of top-ranked tuned base models; and sentiment recognition is performed on the dialogue text to be recognized through the dialogue text sentiment recognition model to obtain a sentiment recognition result. In this way, parameter tuning of multiple base models is to enable the base models to obtain more accurate recognition results, laying a good data foundation for the subsequent construction of the dialogue text sentiment recognition model; then, a dialogue text sentiment recognition model is constructed through a number of top-ranked tuned base models. Multiple models can focus on different features and combine the advantages of multiple models, thereby ensuring that the true sentiment of the dialogue text can be correctly recognized and improving the accuracy of dialogue text sentiment recognition.

[0016] In some embodiments, the parameter tuning of the multiple base models to obtain multiple tuned base models includes:

[0017] Perform parameter tuning on the multiple base models using Bayesian optimization to obtain multiple tuned base models.

[0018] In some embodiments, the calculating of the accuracy of each tuned base model according to the multiple verification results corresponding to each tuned base model includes:

[0019] According to the multiple verification results corresponding to each of the optimized base models, count the number of times each of the optimized base models predicts correctly;

[0020] Divide the number of times each of the optimized base models predicts correctly by the total number of verifications to obtain the accuracy rate of each of the optimized base models.

[0021] In some embodiments, constructing a dialogue text sentiment recognition model based on several of the top-ranked optimized base models includes:

[0022] Calculate the total number of times that several of the top-ranked optimized base models predict correctly;

[0023] Calculate the probability value between the number of times each of the top-ranked optimized base models predicts correctly and the total number of times of correct predictions;

[0024] Use the probability value as the weight of each of the top-ranked optimized base models corresponding thereto;

[0025] Construct a dialogue text sentiment recognition model according to the weights and several of the top-ranked optimized base models. In some embodiments, constructing a dialogue text sentiment recognition model according to the weights and several of the top-ranked optimized base models includes:

[0026]

[0027] wherein, TSRM represents the dialogue text sentiment recognition model, N represents the number of several of the top-ranked optimized base models, P i represents the probability value of the i-th top-ranked optimized base model, and Y i represents the i-th top-ranked optimized base model.

[0028] In some embodiments, obtaining the dialogue text to be recognized and the dialogue text dataset with sentiment labels includes:

[0029] Obtain a set of dialogue audios;

[0030] Use automatic speech recognition to filter out the audios in the set of dialogue audios with a speech recognition rate lower than a preset threshold to obtain a filtered set of dialogue audios;

[0031] Use the Whisper model to convert the audios in the filtered set of dialogue audios into dialogue texts.

[0032] In some embodiments, training each of the optimized base models through the dialogue text training dataset to obtain multiple trained optimized base models includes:

[0033] Based on the dialogue text training dataset, each of the optimized base models is trained using the five-fold cross-validation method to obtain multiple trained optimized base models.

[0034] In a second aspect, an embodiment of the present application further provides a dialogue text sentiment recognition system, which includes:

[0035] A data acquisition unit, configured to acquire the dialogue text to be recognized and the dialogue text dataset with sentiment labels, and divide the dialogue text dataset into a dialogue text training dataset and a dialogue text validation dataset;

[0036] A parameter tuning unit, configured to preset multiple base models and perform parameter tuning on the multiple base models to obtain multiple optimized base models;

[0037] A model training unit, configured to train each of the optimized base models through the dialogue text training dataset to obtain multiple trained optimized base models;

[0038] A model validation unit, configured to perform multiple validations on each of the trained optimized base models through the dialogue text validation dataset to obtain multiple validation results corresponding to each of the optimized base models;

[0039] A data calculation unit, configured to calculate the accuracy rate of each of the optimized base models according to the multiple validation results corresponding to each of the optimized base models;

[0040] A model selection unit, configured to select a number of the top-ranked optimized base models based on the accuracy rate of each of the optimized base models;

[0041] A model construction unit, configured to construct a dialogue text sentiment recognition model according to a number of the top-ranked optimized base models;

[0042] A sentiment recognition unit, configured to perform sentiment recognition on the dialogue text to be recognized through the dialogue text sentiment recognition model to obtain a sentiment recognition result.

[0043] In a third aspect, an embodiment of the present application further provides an electronic device, including at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute a dialogue text sentiment recognition method as described above.

[0044] Fourthly, an embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute a method for recognizing the sentiment of a dialogue text as described above.

[0045] It can be understood that the beneficial effects of the above second to fourth aspects compared with the related art are the same as those of the first aspect compared with the related art. For relevant descriptions, please refer to the relevant descriptions in the first aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:

[0047] Figure 1 is a schematic flowchart of an embodiment of the method for recognizing the sentiment of a dialogue text provided by the present application;

[0048] Figure 2 is a schematic diagram of the tabular output result in the best embodiment of the method for recognizing the sentiment of a dialogue text provided by the present application;

[0049] Figure 3 is a schematic structural diagram of an embodiment of the system for recognizing the sentiment of a dialogue text provided by the present application;

[0050] Figure 4 is a schematic structural diagram of an embodiment of the electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the present application and should not be construed as limiting the present application.

[0052] In the description of the present application, if the first, second, etc. are described only for the purpose of distinguishing technical features, they should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or the sequence of the indicated technical features.

[0053] In the description of the present application, it should be understood that the orientation or positional relationship indicated by terms such as up, down, etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the present application.

[0054] In the description of this application, it should be noted that unless otherwise clearly defined, terms such as setting, installation, and connection should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above terms in this application in combination with the specific content of the technical solution.

[0055] The development of speech recognition technology has gone through several decades. From the initial simple command recognition to today's intelligent speech assistants that support multiple languages and adapt to complex environments, speech recognition has become an indispensable part of our daily lives. Converting speech into text can bring convenience to many fields. For example, for dialogue emotion recognition, after converting the speech dialogue into dialogue text, it is possible to better recognize the emotion of the speech dialogue.

[0056] Existing methods for dialogue text emotion recognition rely on the learning degree of the constructed model and also on the richness, diversity, and correctness of the dataset. However, different features are concerned, and different models have different advantages. Using only one model for prediction cannot well guarantee the accuracy of the prediction. Therefore, existing methods for dialogue text emotion recognition often cannot correctly recognize the true emotion of the dialogue text during the emotion recognition process, resulting in a relatively low emotion recognition accuracy.

[0057] To solve the above problem of relatively low emotion recognition accuracy, this application proposes a method, system, electronic device, and storage medium for dialogue text emotion recognition.

[0058] Refer to Figure 1 , this embodiment of the application provides a method for dialogue text emotion recognition, and the method includes the following steps:

[0059] Step S100: Obtain the dialogue text to be recognized and the dialogue text dataset with emotion labels, and divide the dialogue text dataset into a dialogue text training dataset and a dialogue text validation dataset;

[0060] Step S200: Preset multiple basic models, and perform parameter tuning on the multiple basic models to obtain multiple tuned basic models;

[0061] Step S300: Train each tuned basic model through the dialogue text training dataset to obtain multiple trained tuned basic models;

[0062] Step S400: Perform multiple validations on each trained tuned basic model through the dialogue text validation dataset to obtain multiple validation results corresponding to each tuned basic model;

[0063] Step S500: Calculate the accuracy rate of each tuned basic model according to the multiple validation results corresponding to each tuned basic model;

[0064] Step S600: Select a number of top-ranked optimized base models based on the accuracy of each optimized base model;

[0065] Step S700: Construct a dialogue text sentiment recognition model based on the number of top-ranked optimized base models;

[0066] Step S800: Perform sentiment recognition on the dialogue text to be recognized through the dialogue text sentiment recognition model to obtain a sentiment recognition result.

[0067] In this embodiment, by obtaining the dialogue text to be recognized and a dialogue text data set with sentiment labels, and dividing the dialogue text data set into a dialogue text training data set and a dialogue text validation data set; presetting multiple base models, and performing parameter tuning on the multiple base models to obtain multiple optimized base models; training each optimized base model through the dialogue text training data set to obtain multiple trained optimized base models; performing multiple validations on each trained optimized base model through the dialogue text validation data set to obtain multiple validation results corresponding to each optimized base model; calculating the accuracy of each optimized base model according to the multiple validation results corresponding to each optimized base model; selecting a number of top-ranked optimized base models based on the accuracy of each optimized base model; constructing a dialogue text sentiment recognition model based on the number of top-ranked optimized base models; performing sentiment recognition on the dialogue text to be recognized through the dialogue text sentiment recognition model to obtain a sentiment recognition result. In this way, the purpose of performing parameter tuning on multiple base models is to enable the base models to obtain more accurate recognition results and lay a good data foundation for constructing the dialogue text sentiment recognition model in the later stage; then, a dialogue text sentiment recognition model is constructed through a number of top-ranked optimized base models, and multiple models can focus on different features and combine the advantages of multiple models, so as to ensure that the true sentiment of the dialogue text can be correctly recognized and the accuracy of dialogue text sentiment recognition can be improved.

[0068] In some embodiments, performing parameter tuning on multiple base models to obtain multiple optimized base models includes:

[0069] Performing parameter tuning on multiple base models using Bayesian optimization to obtain multiple optimized base models.

[0070] In this embodiment, by performing parameter tuning on multiple base models using Bayesian optimization to obtain multiple optimized base models, the purpose is to enable the base models to obtain more accurate recognition results and lay a good data foundation for constructing the dialogue text sentiment recognition model in the later stage.

[0071] In some embodiments, calculating the accuracy rate of each fine-tuned base model based on multiple verification results corresponding to each fine-tuned base model includes:

[0072] Counting the number of correct predictions of each fine-tuned base model according to the multiple verification results corresponding to each fine-tuned base model;

[0073] Dividing the number of correct predictions of each fine-tuned base model by the total number of verifications to obtain the accuracy rate of each fine-tuned base model.

[0074] In this embodiment, counting the number of correct predictions of each fine-tuned base model according to the multiple verification results corresponding to each fine-tuned base model; dividing the number of correct predictions of each fine-tuned base model by the total number of verifications to obtain the accuracy rate of each fine-tuned base model. Calculating the accuracy rate of each fine-tuned base model lays a good data foundation for selecting a better base model in the later stage, thereby improving the accuracy of later emotion prediction.

[0075] In some embodiments, constructing a dialogue text emotion recognition model based on several top-ranked fine-tuned base models includes:

[0076] Calculating the total number of correct predictions of several top-ranked fine-tuned base models;

[0077] Calculating the probability value between the number of correct predictions of each top-ranked fine-tuned base model and the total number of correct predictions;

[0078] Taking the probability value as the weight of each top-ranked fine-tuned base model corresponding thereto;

[0079] Constructing a dialogue text emotion recognition model according to the weights and several top-ranked fine-tuned base models.

[0080] In this embodiment, constructing a dialogue text emotion recognition model according to the weights and several top-ranked fine-tuned base models. Constructing a dialogue text emotion recognition model through multiple fine-tuned base models can focus on different features and combine the advantages of multiple models. Since the prediction accuracies of different base models are different, in order to improve the recognition accuracy, weights are added to each selected fine-tuned base model, so as to ensure that the true emotion of the dialogue text can be correctly recognized and the precision of dialogue text emotion recognition is improved.

[0081] In some embodiments, constructing a dialogue text emotion recognition model according to the weights and several top-ranked fine-tuned base models includes:

[0082]

[0083] TSRM represents the conversation text emotion recognition model, N represents the number of top-ranked optimized basic models, and P i represents the probability value of the i-th top-ranked optimized basic model, Y i represents the top-ranked tuned base model.

[0084] In some implementations, obtaining a conversation text to be recognized and a conversation text dataset with emotion tags includes:

[0085] Get the conversation audio collection;

[0086] Automatic speech recognition is used to filter out audio with a speech recognition rate lower than a preset threshold in the conversation audio set, thereby obtaining a filtered conversation audio set;

[0087] The Whisper model is used to convert the audio in the filtered conversation audio collection into conversation text.

[0088] In this embodiment, automatic speech recognition is used to filter out the audio with a speech recognition rate lower than a preset threshold in the conversation audio set, and a filtered conversation audio set is obtained. Then, the Whisper model is used to convert the audio in the filtered conversation audio set into conversation text. This can ensure the quality and accuracy of the audio-to-text data, lay a good data foundation for training an accurate basic model in the later stage, and thus build an accurate conversation text emotion recognition model.

[0089] In some implementations, each tuned base model is trained using a conversation text training dataset to obtain a plurality of trained tuned base models, including:

[0090] Based on the conversation text training dataset, the five-fold cross-validation method is used to train each tuned base model to obtain multiple trained tuned base models.

[0091] In this embodiment, a five-fold cross-validation method is used to train each tuned basic model to obtain multiple trained tuned basic models, which can make full use of data and improve the generalization ability of the model.

[0092] To facilitate understanding by those skilled in the art, a set of best embodiments is provided below:

[0093] The development of speech recognition technology has gone through decades of development. From the initial simple command recognition to today's intelligent voice assistants that support multiple languages ​​and adapt to complex environments, speech recognition has become an indispensable part of our daily lives. Converting speech into text can bring convenience to many fields, such as conversation emotion recognition. After converting speech conversations into text, the emotions of speech conversations can be better recognized.

[0094] Existing dialogue text sentiment recognition methods rely on the learning degree of the constructed model, as well as the richness, diversity, and correctness of the dataset. However, different features are concerned, and different models have different advantages. Using only one model for prediction cannot well guarantee the accuracy of prediction. Therefore, existing dialogue text sentiment recognition methods often cannot correctly recognize the true sentiment of dialogue texts during the sentiment recognition process, resulting in relatively low sentiment recognition accuracy.

[0095] To solve the above problems, in the stage of collecting the dialogue text dataset in this embodiment, the embodiment first collects comprehensive and diverse dialogue audios, and then constructs a dialogue text dataset based on the collected audios. The dialogue text dataset needs to cover a wide range of emotion categories (such as happy, sad, angry, surprised, calm, etc.) and various dialogue styles (such as formal, casual, humorous, serious, etc.). The data sources include but are not limited to public dialogue records on social media platforms, dialogue fragments in movies and TV shows, professional emotion databases, and dialogue samples recorded in specific scenarios. Then, based on the constructed dialogue text dataset, multiple basic models are trained, several basic models with high accuracy are selected from the multiple trained basic models, and finally, a dialogue text sentiment recognition model is constructed according to the selected several basic models with high accuracy. The dialogue text sentiment recognition model is used to perform dialogue text sentiment recognition to obtain the dialogue text sentiment recognition result. The specific contents are as follows:

[0096] 1. Convert dialogue audio to dialogue text.

[0097] Comprehensive and diverse dialogue audios are obtained. To ensure the quality and accuracy of the data, in this embodiment, the automatic speech recognition (ASR) technology is used to initially filter out the audios with a speech recognition rate lower than the preset threshold to reduce the noise interference in subsequent processing, and then the Whisper model is used to convert the filtered audios into dialogue texts.

[0098] 2. Construct a dialogue text dataset.

[0099] After obtaining the dialogue text in step 1 in this embodiment, each sentence in the dialogue text is sentiment-annotated, which can be manually annotated to construct a dialogue text dataset with accurate sentiment labels.

[0100] 3. Construct a dialogue text sentiment recognition model.

[0101] (1) This embodiment selects multiple basic models from the prior art as the basic skeleton of the conversation text emotion recognition model. For example, existing neural network models such as multi-layer perceptron (MLP), convolutional neural network (CNN), recurrent neural network (RNN), long short-term memory network (LSTM), Transformer network model, fully connected neural network (FCN) and large language model are selected as basic models.

[0102] For the large language model used in this embodiment, when the large language model performs emotion recognition, PROMPT prompt words are used for training in the large language model, and / n is used to separate sentences in the dialogue text. At the same time, it is recommended to prohibit the large language model from modifying the sentence situation in PROMPT. In the large language model containing the intelligent agent, a tabular output result can be obtained, such as Figure 2 shown.

[0103] PROMPT needs to clearly express the format and tone, and the definition of emotion in order to generate accurately. The example PROMPT is: Please judge the tone style and emotion of the sentence I will provide next. Each line is regarded as a sentence. Emotions are selected from "happy", "surprised", "sad", "disgusted", "angry", "fear", "neutral", and "doubtful". The tone style is selected from "steady", "high-pitched", "low", "soothing", "whispering", "gentle", "playful", "serious", "loud", "cheerful", "nervous", "confident", and "stuttering". Just judge the emotion of the sentence in the dialogue text, and tell me the judgment result in the form of a table. If you are not sure, you can mark 2 or more labels, put the important ones in front, and separate them with a minus sign. List the emotion categories in Excel in the order of the sentences. The next step is to enter the sentences that need to be judged into the big prediction model.

[0104] (2) The conversation text dataset with accurate emotion labels constructed in step 2 is divided into a conversation text training dataset and a conversation text verification dataset. Before training, Bayesian optimization is used to tune the parameters of each basic model to obtain a tuned basic model, and then a five-fold cross-validation method is used to train each basic model to obtain the accuracy of each tuned basic model; an accuracy threshold is preset, and the accuracy of each tuned basic model is compared with the accuracy threshold. If the accuracy of the tuned basic model is greater than the accuracy threshold, the tuned basic model is used as a basic skeleton screening list for the conversation text emotion recognition model; a preset number of basic models with the top rankings are selected from the basic skeleton screening list.

[0105] Specifically, the accuracy rate of each optimized basic model can be calculated by validating multiple times using the validation set to obtain the number of times of accurate prediction, and then dividing the number of times of accurate prediction by the total number of validations to obtain the accuracy rate of each optimized basic model. After obtaining the accuracy rate of each optimized basic model, a threshold value threshold is given. The specific setting of threshold can be changed according to the actual situation, and this embodiment does not make specific limitations. If the accuracy rate of the optimized basic model is greater than threshold, the optimized basic model will be included in the basic skeleton screening list, and at the same time, a selection of the number N of basic models will be given. The first N optimized basic models in the basic skeleton screening list are selected as the basic skeleton models of the dialogue text sentiment recognition model. The specific number of basic models selected in this embodiment can be changed according to the actual situation, and this embodiment does not make specific limitations.

[0106] (3) Based on the top-ranked preset number of optimized basic models, construct a dialogue text sentiment recognition model. Specifically, the following method is used to assign weights to the top-ranked multiple optimized basic models to construct a dialogue text sentiment recognition model:

[0107] Calculate the total number of times the top-ranked multiple optimized basic models predict correctly:

[0108]

[0109] Among them, Z represents the total number of times of correct prediction, and Y i represents the i-th top-ranked optimized basic model, and N represents the number of top-ranked multiple optimized basic models.

[0110] Calculate the probability value between the number of times each top-ranked optimized basic model predicts correctly and the total number of times of correct prediction:

[0111]

[0112] Among them, P i represents the probability value of the i-th top-ranked optimized basic model.

[0113] Take the calculated probability value P i as the weight of each top-ranked optimized basic model, and construct the dialogue text sentiment recognition model TSRM as:

[0114]

[0115] After assigning weights to multiple optimized base models with top rankings, weights are assigned to the recognition results obtained for each optimized base model to obtain the final recognition result. For example, if 5 optimized base models are selected, the constructed dialogue text sentiment recognition model is: 0.1A + 0.2B + 0.3C + 0.3D + 0.1E. The optimized base model A predicts a sentiment of class 1, the optimized base model B predicts a sentiment of class 2, the optimized base model C predicts a sentiment of class 1, the optimized base model D predicts a sentiment of class 2, and the optimized base model E predicts a sentiment of class 3. Then the final recognition result is a sentiment of class 2.

[0116] Referring to Figure 3 , the embodiment of the present application further provides a dialogue text sentiment recognition system, which includes a data acquisition unit 100, a parameter tuning unit 200, a model training unit 300, a model verification unit 400, a data calculation unit 500, a model selection unit 600, a model construction unit 700, and a sentiment recognition unit 800, where:

[0117] The data acquisition unit 100 is used to acquire the dialogue text to be recognized and the dialogue text data set with sentiment labels, and divide the dialogue text data set into a dialogue text training data set and a dialogue text verification data set;

[0118] The parameter tuning unit 200 is used to preset multiple base models and perform parameter tuning on the multiple base models to obtain multiple optimized base models;

[0119] The model training unit 300 is used to train each optimized base model through the dialogue text training data set to obtain multiple trained optimized base models;

[0120] The model verification unit 400 is used to perform multiple verifications on each trained optimized base model through the dialogue text verification data set to obtain multiple verification results corresponding to each optimized base model;

[0121] The data calculation unit 500 is used to calculate the accuracy rate of each optimized base model according to the multiple verification results corresponding to each optimized base model;

[0122] The model selection unit 600 is used to select several optimized base models with top rankings based on the accuracy rate of each optimized base model;

[0123] The model construction unit 700 is used to construct a dialogue text sentiment recognition model according to several optimized base models with top rankings;

[0124] An emotion recognition unit 800 is configured to perform emotion recognition on the dialogue text to be recognized through a dialogue text emotion recognition model, and obtain an emotion recognition result.

[0125] It should be noted that since a dialogue text emotion recognition system in this embodiment and the above-mentioned dialogue text emotion recognition method are based on the same inventive concept, the corresponding content in the method embodiment is equally applicable to the system embodiment of this system, and will not be elaborated here.

[0126] Refer to Figure 4 , this application embodiment also provides an electronic device, which includes:

[0127] At least one memory;

[0128] At least one processor;

[0129] At least one program;

[0130] The program is stored in the memory, and the processor executes at least one program to implement the dialogue text emotion recognition method described above in this disclosure.

[0131] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.

[0132] The electronic device of this application embodiment will be introduced in detail below.

[0133] The processor 1600 can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by this disclosure embodiment;

[0134] The memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1700 can store an operating system and other application programs. When implementing the technical solutions provided by this specification embodiment through software or firmware, the relevant program codes are stored in the memory 1700, and are called by the processor 1600 to execute the dialogue text emotion recognition method of this disclosure embodiment.

[0135] The input / output interface 1800 is configured to implement information input and output;

[0136] A communication interface 1900 is used to implement communication interaction between this device and other devices. It can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0137] A bus 2000 transmits information between various components of the device (such as a processor 1600, a memory 1700, an input / output interface 1800, and a communication interface 1900).

[0138] Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 achieve communication connections with each other inside the device through the bus 2000.

[0139] The embodiments of the present disclosure also provide a storage medium. This storage medium is a computer-readable storage medium, and this computer-readable storage medium stores computer-executable instructions, and these computer-executable instructions are used to cause a computer to execute the above-mentioned dialogue text sentiment recognition method.

[0140] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0141] The embodiments described in the embodiments of the present disclosure are for more clearly explaining the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.

[0142] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than those shown in the figures, or combine some steps, or different steps.

[0143] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0144] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0145] The terms "first", "second", "third", "fourth", etc. (if any) in the description of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0146] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0147] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0148] The unit described as a separation component may or may not be physically separated, and the component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0149] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, can also exist separately as individual physical units, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0150] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store programs. The above has described the embodiments of the present application in detail with reference to the drawings, but the present application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can be made without departing from the purpose of the present application.

[0151] The above has described the embodiments of the present application in detail with reference to the drawings, but the present application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can be made without departing from the purpose of the present application.

Claims

1. A method for emotion recognition of dialogue text, characterized in that: The method comprises: Acquire a dialogue text to be recognized and a dialogue text dataset with emotion labels, and divide the dialogue text dataset into a dialogue text training dataset and a dialogue text verification dataset; Presetting a plurality of basic models, and performing parameter tuning on the plurality of basic models to obtain a plurality of tuned basic models; Training each of the tuned basic models using the conversation text training data set to obtain a plurality of trained tuned basic models; Verifying each of the trained and tuned basic models multiple times using the conversation text verification data set to obtain multiple verification results corresponding to each of the tuned basic models; Calculate the accuracy of each of the tuned basic models according to the multiple verification results corresponding to each of the tuned basic models; Based on the accuracy of each of the tuned basic models, select a number of tuned basic models with top rankings; Build a conversation text emotion recognition model based on several top-ranked optimized basic models; The emotion recognition model is used to perform emotion recognition on the dialogue text to be recognized, and an emotion recognition result is obtained.

2. The method for emotion recognition of dialogue text according to claim 1, characterized in that: The step of performing parameter tuning on the multiple basic models to obtain multiple tuned basic models includes: Bayesian optimization is used to tune the parameters of the multiple basic models to obtain multiple tuned basic models.

3. The method for emotion recognition of dialogue text according to claim 1, characterized in that: The step of calculating the accuracy of each of the tuned basic models according to the multiple verification results corresponding to each of the tuned basic models includes: According to the multiple verification results corresponding to each of the tuned basic models, the number of accurate predictions of each of the tuned basic models is counted; The accuracy rate of each tuned basic model is obtained by dividing the number of accurate predictions of each tuned basic model by the total number of verifications.

4. The method for identifying emotion in dialogue text according to claim 1, characterized in that: The dialogue text emotion recognition model is constructed based on several top-ranked optimized basic models, including: Calculate the total number of correct predictions by the top-ranked optimized basic models; Calculate the probability value between the number of correct predictions of each top-ranked tuned base model and the total number of correct predictions; The probability value is used as the weight of each top-ranked basic model corresponding thereto after optimization; A conversation text emotion recognition model is constructed according to the weights and the top-ranked several optimized basic models.

5. The method for identifying emotion in dialogue text according to claim 4, characterized in that: The step of constructing a dialogue text emotion recognition model according to the weights and the top-ranked optimized basic models includes: Where TSRM represents the conversation text emotion recognition model, N represents the number of top-ranked tuned basic models, P represents the probability value of the i-th top-ranked tuned basic model, and Y i represents the top-ranked tuned base model.

6. The method for identifying emotion in dialogue text according to claim 1, characterized in that: The step of obtaining the dialogue text to be recognized and the dialogue text dataset with emotion labels includes: Get the conversation audio collection; Using automatic speech recognition to filter out audio with a speech recognition rate lower than a preset threshold in the conversation audio set, to obtain a filtered conversation audio set; The Whisper model is used to convert the audio in the filtered conversation audio set into conversation text.

7. The method for identifying emotion in dialogue text according to claim 1, characterized in that: The step of training each of the tuned basic models using the conversation text training data set to obtain a plurality of trained tuned basic models includes: Based on the conversation text training data set, a five-fold cross validation method is used to train each of the tuned basic models to obtain multiple trained tuned basic models.

8. A conversation text emotion recognition system, characterized in that: The system comprises: A data acquisition unit, used to acquire a dialogue text to be recognized and a dialogue text dataset with emotion tags, and divide the dialogue text dataset into a dialogue text training dataset and a dialogue text verification dataset; A parameter tuning unit, used to preset a plurality of basic models and perform parameter tuning on the plurality of basic models to obtain a plurality of tuned basic models; A model training unit, used for training each of the tuned basic models using the dialogue text training data set to obtain a plurality of trained tuned basic models; A model verification unit, used to verify each of the trained and tuned basic models multiple times using the dialogue text verification data set to obtain multiple verification results corresponding to each of the tuned basic models; A data calculation unit, used to calculate the accuracy of each of the tuned basic models according to a plurality of verification results corresponding to each of the tuned basic models; A model selection unit, configured to select a number of top-ranked optimized basic models based on the accuracy of each of the optimized basic models; A model building unit, used to build a dialogue text emotion recognition model based on several top-ranked optimized basic models; The emotion recognition unit is used to perform emotion recognition on the dialogue text to be recognized through the dialogue text emotion recognition model to obtain an emotion recognition result.

9. An electronic device, characterized in that: It includes at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the conversation text emotion recognition method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the conversation text emotion recognition method as described in any one of claims 1 to 7.