An emotional dialogue generation method, device, computer device and storage medium

By introducing semantic analysis and emotion label generation models into the large language model, the problem that chatbots find it difficult to understand user emotions is solved, and emotional supportive dialogue generation is achieved without additional hints.

CN118916466BActive Publication Date: 2025-06-10MIND WITH HEART ROBOTICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411173559.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2025-06-10
Estimated Expiration
2044-08-26

AI Technical Summary

Technical Problem

Existing large language models are difficult to understand and process emotional information in user conversations in the field of chatbots, resulting in a lack of empathy and emotional support for generated conversations.

Method used

By taking natural language text and inputting it into two models for processing, the first model performs semantic analysis and generation of sentiment labels, and the second model generates sentiment-related response text based on these results.

Benefits of technology

It enables the generation of empathetic, emotionally supportive conversations without guiding prompt words or instructions, simplifying operational steps and reducing the difficulty of engineering implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118916466B_ABST
    Figure CN118916466B_ABST
Patent Text Reader

Abstract

The present invention discloses an emotional dialogue generation method, apparatus, computer device and storage medium. The method includes: obtaining a natural language text to be processed; inputting the natural language text to be processed into a first model for processing to output a semantic analysis result and an emotional label; and inputting the natural language text to be processed, the semantic analysis result and the emotional label into a second model for processing to output a final natural language text in response to the natural language text to be processed. When the present invention is applied, the user only needs to input a natural language text, and finally can output text data results with empathy and capable of emotionally responding to the user input. Unlike the prior art, it is not necessary to input guiding prompt words or instructions to extract and stimulate the existing emotional processing content in the model. Compared with the prior art, the operation steps are effectively simplified and the engineering implementation difficulty is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and more specifically, to an emotional dialogue generation method, device, computer device, and storage medium. Background Art

[0002] Large language models are widely used in related fields of artificial intelligence. In the field of chatbots, large language models are used to analyze and understand the language dialogue content of human users, and after a large amount of data processing, output human language dialogue results containing specific useful content to users. So far, various types of large language models have been developed to apply to different scenarios, including medical, legal, computer, etc., but they all lack the ability to show empathy for the emotions in the user's dialogue and are difficult to generate considerate and emotionally supportive results for users.

[0003] In order to enable large language models to understand and process emotional information, generally, guiding prompt words or instructions need to be input into the existing large language models to extract and stimulate the existing emotional processing content in the models. The operation steps are relatively complex, and the workload is also relatively large. Moreover, due to the huge amount of data and numerous contents in the large language models themselves, the processing effect of this way of enabling large language models to partially understand the emotional information in the dialogue content is not good.

[0004] Chinese Patent, Patent Publication No. CN107943974A discloses an automatic conversation method and system considering emotions. First, it obtains the sentence and emotional label input by the user simultaneously; then determines the current semantics and emotions of the user; based on a preset conversation model, determines a reply that conforms to the user's current semantics and emotions according to the user's current semantics and emotions; and finally outputs the reply content. This solution requires the user to input both the sentence and the emotional label at the same time, and the emotional label input by the user will be directly input into the model for processing. Moreover, the reply sentences generated by the model include at least multiple sentences such as the first reply sentence, the second reply sentence, and the third reply sentence, and in its processing steps, these multiple reply sentences all need to be input into the model for processing and analyzing the emotions of the sentences.

[0005] A Chinese patent with the patent publication number CN108874972A discloses a multi-round sentiment dialogue method based on deep learning. It tokenizes the text information input by the user and vectorizes the text through a pre-trained word vector model. It uses a deep learning model to perform sentiment analysis on the text input by the user and analyze the dialogue theme and background. It retrieves the most likely dialogue responses from the sentiment corpus based on a retrieval method. Based on the sentiment category of the user's dialogue, as well as the chat theme and background, it uses a generative adversarial network to generate natural dialogue responses. According to two different dialogue generation methods, it selects a dialogue with the most relevant dialogue sentiment and theme background to the user input and sends it to the user. In this solution, the method for generating the response text is to retrieve the most likely dialogue responses from the sentiment corpus based on the processing results of the user's dialogue. There are pre-stored dialogues in the corpus, and the existing dialogues in the corpus have been pre-indexed with sentiment and subject labels. Summary of the Invention

[0006] The object of the present invention is to overcome the deficiencies of the prior art and provide a sentiment dialogue generation method, device, computer device and storage medium.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] In a first aspect, the present invention provides a sentiment dialogue generation method, including:

[0009] Obtain the natural language text to be processed;

[0010] Input the natural language text to be processed into a first model for processing to output a semantic analysis result and a sentiment label;

[0011] Input the natural language text to be processed, the semantic analysis result and the sentiment label into a second model for processing to output the final natural language text in response to the natural language text to be processed.

[0012] Further, the inputting the natural language text to be processed into a first model for processing to output a semantic analysis result and a sentiment label includes:

[0013] Perform sentiment analysis processing on the natural language text to be processed to obtain a sentiment attribute result;

[0014] Find the words related to the sentiment attribute result from the natural language text to be processed to obtain emotion keywords that can express emotions;

[0015] Perform semantic analysis on the natural language text to be processed to obtain a semantic analysis result;

[0016] Analyze and obtain sentiment labels by combining sentiment attribute results, emotion keywords, and semantic analysis results.

[0017] Further, perform sentiment analysis on the natural language text to be processed to obtain sentiment attribute results, including:

[0018] Extract sentiment-related features from the natural language text to be processed;

[0019] Perform sentiment attribute analysis on the sentiment-related features to obtain sentiment attribute results.

[0020] Further, search for words related to the sentiment attribute results in the natural language text to be processed to obtain emotion keywords that can express emotions, including:

[0021] Perform multi-level analysis on the natural language text to be processed to identify keywords that may express emotions;

[0022] Screen out keywords related to the sentiment attribute results from the keywords that may express emotions as the final emotion keywords.

[0023] Further, perform semantic analysis on the natural language text to be processed to obtain semantic analysis results, including:

[0024] Perform word segmentation and grammar analysis on the natural language text to be processed to identify the basic structure of the text;

[0025] Analyze each word and its position in the text based on the basic structure of the text to obtain the position information of the words;

[0026] Perform analysis and processing based on the basic structure of the text and the position information of the words to obtain semantic analysis results.

[0027] Further, combine the sentiment attribute results, emotion keywords, and semantic analysis results to analyze and obtain sentiment labels, including:

[0028] Integrate the sentiment attribute results, emotion keywords, and semantic analysis results to obtain comprehensive sentiment description information;

[0029] Identify sentiment features based on the comprehensive sentiment description information;

[0030] Match the sentiment features with the words in the built-in sentiment label list;

[0031] Select one or more sentiment labels from the sentiment label list according to the matching results.

[0032] Further, the step of inputting the natural language text to be processed, the semantic analysis result, and the sentiment label into a second model for processing to output a final natural language text in response to the natural language text to be processed includes:

[0033] Setting a starting point for text generation and constructing the text in a word-by-word generation manner;

[0034] Adjusting the sentiment tone of the generation of each word using the sentiment label;

[0035] Checking the generated text according to the semantic analysis result to ensure that the generated text contains the core theme and intention of the natural language text to be processed.

[0036] In a second aspect, the present invention further provides an emotional dialogue generation device, including:

[0037] An acquisition unit for acquiring the natural language text to be processed;

[0038] A first model processing unit for inputting the natural language text to be processed into a first model for processing to output a semantic analysis result and a sentiment label;

[0039] A second model processing unit for inputting the natural language text to be processed, the semantic analysis result, and the sentiment label into a second model for processing to output a final natural language text in response to the natural language text to be processed.

[0040] In a third aspect, the present invention further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the emotional dialogue generation method as described above is implemented.

[0041] In a fourth aspect, the present invention further provides a computer-readable storage medium. The storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the emotional dialogue generation method as described above.

[0042] The beneficial effects of the present invention compared with the prior art are as follows: An emotional dialogue generation method includes: obtaining a natural language text to be processed; inputting the natural language text to be processed into a first model for processing to output a semantic analysis result and an emotional label; inputting the natural language text to be processed, the semantic analysis result, and the emotional label into a second model for processing to output a final natural language text in response to the natural language text to be processed. When the present invention is applied, the user only needs to input a natural language text, and finally can output text data results with empathy and capable of emotionally responding to the user input. Unlike the prior art, it does not require inputting guiding prompt words or instructions to extract and stimulate the existing emotional processing content in the model. Compared with the prior art, the operation steps are effectively simplified, and the engineering implementation difficulty is reduced.

[0043] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given and described in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 It is the main flowchart of an emotional dialogue generation method provided by a specific embodiment of the present invention;

[0046] Figure 2 It is a sub - process of an emotional dialogue generation method provided by a specific embodiment of the present invention Figure 1 ;

[0047] Figure 3 It is a sub - process of an emotional dialogue generation method provided by a specific embodiment of the present invention Figure 2 ;

[0048] Figure 4 It is a sub - process of an emotional dialogue generation method provided by a specific embodiment of the present invention Figure 3 ;

[0049] Figure 5 It is a sub - process of an emotional dialogue generation method provided by a specific embodiment of the present invention Figure 4 ;

[0050] Figure 6A sub - process of an emotional dialogue generation method provided by a specific embodiment of the present invention Figure 5 ;

[0051] Figure 7 A sub - process of an emotional dialogue generation method provided by a specific embodiment of the present invention Figure 6 ;

[0052] Figure 8 A schematic block diagram of an emotional dialogue generation device provided by a specific embodiment of the present invention;

[0053] Figure 9 A schematic block diagram of a computer device provided by a specific embodiment of the present invention. Detailed implementation manners

[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0056] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0057] It should be further understood that the term " / and / or" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.

[0058] Now, the following explanations are made for the technical terms related to the present invention:

[0059] Deep learning: A branch type of machine learning, which is an algorithm for feature learning of data with a multi - layer artificial neural network as the architecture;

[0060] Machine learning: A method to achieve artificial intelligence, mainly involving designing algorithms that enable a computer to automatically analyze and extract patterns from data and use these patterns to predict unknown data.

[0061] Artificial neural network: Abbreviated as neural network. In the field of machine learning, it is a mathematical model or computational model designed by imitating the structure and function of the neural network of organisms, used to implement algorithms for estimating or approximating mathematical functions.

[0062] Natural language: Mainly refers to the language that humans usually use for communication, generally used to distinguish from computer languages.

[0063] The following introduces the present invention through specific embodiments.

[0064] Such as Figure 1 As shown, an embodiment of the present invention provides an emotional dialogue generation method, including the following steps: S10 - S30:

[0065] S10. Obtain the natural language text to be processed;

[0066] The natural language text to be processed is generally input by the user. For example, the user inputs: I tried to bake a sausage and cheese cake last night, but it tasted terrible.

[0067] S20. Input the natural language text to be processed into the first model for processing to output a semantic analysis result and an emotion label.

[0068] A large language model refers to a deep learning model trained based on a large amount of text data. It can process various natural language texts, deeply understand the meaning of the texts, and generate natural language texts as results.

[0069] The first model is a large language model that can reason sequentially by imitating human thinking. It can analyze the emotion that the user wants to express based on the user's text input.

[0070] Such as Figure 2 As shown, in one embodiment, step S20 specifically includes the following steps: S201 - S204:

[0071] S201. Perform emotional analysis on the natural language text to be processed to obtain an emotional attribute result.

[0072] For this step, the main objective is to perform sentiment analysis on the text data input by the user to determine the sentiment attribute expressed by the text, that is, the positivity or negativity of the emotion. The principle of obtaining the sentiment attribute result is as follows: Before use, a large amount of text data is used to perform deep learning training on the first model. These data sets usually contain text samples with various different sentiment labels, such as positive, negative, or neutral texts. The result of the training is saved in a neural network with a large number of parameters. The input text data is processed by the neural network and can output a sentiment attribute result, which can represent the degree of positivity or negativity of the sentiment attribute of the text data.

[0073] As Figure 3 shown, in one embodiment, step S201 specifically includes the following steps: S2011 - S2012:

[0074] S2011. Extract sentiment - related features from the natural - language text to be processed;

[0075] S2012. Perform sentiment - attribute analysis on the sentiment - related features to obtain a sentiment - attribute result;

[0076] For steps S2011 - S2012, when the user actually uses it, a natural - language text is input, such as a paragraph or a sentence. This text is input into the already - trained first model. The neural network of the first model processes this input text and extracts the sentiment - related features from it. After being processed layer by layer by the neural network, the model will finally output a sentiment - attribute result. This result is usually a probability distribution, indicating the tendency of the input text towards positive or negative sentiment. For example, the model may output two values, representing the probabilities of "positive" and "negative" respectively. If the probability of positive is higher, it means that the input text is more inclined to express positive sentiment, and vice versa. The output sentiment - attribute result can be directly used in subsequent steps. For example, in the first model, this sentiment - attribute result will be passed to the next step (such as keyword extraction, semantic analysis, etc.) or used as one of the inputs for subsequent generation of empathetic responses;

[0077] S202. Search for words related to the sentiment - attribute result from the natural - language text to be processed to obtain emotion keywords that can express emotions;

[0078] For this step, the goal is to identify sentiment-related keywords from the text input by the user. These keywords are usually important clues to express the user's emotional state, such as "like", "hate", "angry", etc. The principle of obtaining the sentiment attribute result is as follows: A large amount of text data was previously used for in-depth learning training of the first model, and the training results are stored in a neural network with a large number of parameters. After the input text data is processed by the neural network, the first model can find one or more keywords from the text data, and this keyword is a word that can show the emotion expressed by the user in the sentence.

[0079] As Figure 4 shown, step S202 specifically includes the following steps: S2021 - S2022:

[0080] S2021. Perform multi-level analysis on the natural language text to be processed to identify keywords that may express sentiment;

[0081] S2022. Screen out keywords related to the sentiment attribute result from the keywords that may express sentiment as the final emotion keywords;

[0082] For steps S2021 - S2022, the first model will search each word in the input text and determine whether it is related to a certain sentiment. If a word is considered a sentiment keyword by the model (for example, "like" represents positive sentiment and "hate" represents negative sentiment), then this word will be extracted and recorded as the output result of this step. The extraction process of these keywords depends on the knowledge learned by the model during training. Through its neural network parameters, the model can associate certain words in the input text with specific sentiment categories, thereby identifying these words as sentiment keywords. The identified keywords will be used as intermediate results and passed to subsequent processing steps (for example, the selection of sentiment labels, further analysis of sentiment attributes, etc.);

[0083] S203. Perform semantic analysis on the natural language text to be processed to obtain the semantic analysis result;

[0084] For this step, the main task is to perform semantic analysis on the text input by the user, understand the deep meaning of the text, especially to accurately describe the emotional situation expressed in the text through the context. Analyze the semantic analysis result (the semantic analysis result is a piece of natural language text) from the words and sentence structures, and record the semantic analysis result as the result output by this step. The principle of obtaining the semantic analysis result is that a large amount of text data was previously used to perform deep learning training on the first model, and the training result is saved in a neural network with a large number of parameters. The input text data is processed by the neural network, and the first model can analyze the semantics of the text data and generate another piece of natural language text. The main content of this text is to deeply describe the emotional situation expressed in the input text in combination with the context of the input text.

[0085] As Figure 5 shown, step S203 specifically includes the following steps: S2031 - S2033:

[0086] S2031. Perform word segmentation and syntactic analysis on the natural language text to be processed to identify the basic structure of the text;

[0087] S2032. Analyze each word and its position in the text according to the basic structure of the text to obtain the position information of the words;

[0088] S2033. Perform analysis and processing according to the basic structure of the text and the position information of the words to obtain the semantic analysis result.

[0089] For steps S2031 - S2033,

[0090] When the user inputs a piece of natural language text, the first model will first perform word segmentation and syntactic analysis on the text to identify the basic structure of the sentence, such as the subject, predicate, object, etc. The model will further analyze each word in the text and its position in the sentence, especially paying attention to the influence of the context on the meaning of these words. For example, a certain word may express different emotions or semantics in different contexts. By analyzing the context, the first model can understand the deep meaning of the entire sentence, rather than just looking at each word in isolation. For example, the model can distinguish the semantic difference between "I like you" and "I don't like you", although they both contain the word "like". The first model will comprehensively consider the semantics of the words and the modification of their emotions by the context, and generate a more accurate semantic analysis result. This result not only describes the surface meaning of the text, but also reflects the deep emotions expressed by the context. After completing the context analysis, the model will generate a piece of natural language text, which is a further explanation and description of the semantics and emotions of the input text;

[0091] S204. Analyze and obtain emotion tags by combining the emotion attribute results, emotion keywords, and semantic analysis results.

[0092] For this step, based on the emotion attributes, keywords, and semantic analysis results generated in the previous steps, one or more emotion tags that best match these analysis results are selected from the built-in emotion tag list. The principle of obtaining emotion tags is as follows: A large amount of text data was previously used for in-depth learning training of the first model, and the training results are stored in a neural network with a large number of parameters. When the emotion attribute results, keywords, and semantic analysis results generated in the previous three steps are input into the first model, the first model integrates these data items and processes them through the neural network. The model can select several words from the built-in natural language vocabulary list as emotion tags, and these emotion tags are the words in the vocabulary list that can best reflect the emotions in the text data input by the user.

[0093] As Figure 6 shown, step S204 specifically includes the following steps: S2041 - S2044:

[0094] S2041. Integrate the emotion attribute results, emotion keywords, and semantic analysis results to obtain comprehensive emotion description information;

[0095] S2042. Identify emotion features based on the comprehensive emotion description information;

[0096] S2043. Match the emotion features with the words in the built-in emotion tag list;

[0097] S2044. Select one or more emotion tags from the emotion tag list according to the matching results;

[0098] For steps S2041 - S2044, the first model receives the results from the previous three steps: the sentiment attribute (such as the probability distribution of positive or negative), keywords (such as "like", "angry", etc.), and the semantic analysis result (description of the deep meaning of the text). The first model integrates these input data together to form a comprehensive sentiment description. This integration helps the model better understand the sentiment state of the user input because these results reflect the sentiment information in the text from different perspectives. The integrated data is input into the neural network of the first model for further processing. The hierarchical structure of the neural network allows the model to analyze the complex relationships between these data in a high - dimensional space. Through these processes, the model can identify the features that best represent the sentiment of the user input and match these features with the words in the built - in sentiment label list. According to the matching results, the model selects one or more sentiment labels from the built - in sentiment label list that best match the analysis results of the previous steps. These labels may include "happy", "angry", "sad", etc., depending on the vocabulary of the model and the sentiment features of the user input. The selected sentiment label is output as the final result, marking the completion of the first model's understanding and classification of the sentiment of the user input text;

[0099] S30. Input the natural language text to be processed, the semantic analysis result, and the sentiment label into the second model for processing to output the final natural language text in response to the natural language text to be processed;

[0100] The second model is a large - language model that can perform sequential reasoning by imitating human thinking. It can analyze the emotions that the user wants to express based on the user's text input;

[0101] The input data, including the natural language text to be processed, the semantic analysis result, and the selected sentiment label, contains the original information of the user input text and the model's understanding of its sentiment. The second model will use this information to generate an appropriate response. The principle of being able to output the final natural language text is that the second model has been deeply learned and trained with a large amount of text data and sentiment analysis data, etc. previously. The training results are stored in a neural network with a large number of parameters. The input text data, semantic analysis result, sentiment label, and other data are processed by the neural network and can finally output a natural language text that can respond to the user input text data in terms of text content and emotion in an empathetic way.

[0102] As Figure 7 shown, in one embodiment, step S30 specifically includes the following steps: S301 - S303:

[0103] S301. Set the starting point of text generation and construct the text in a word - by - word generation manner;

[0104] S302, using the emotion tag to adjust the emotional tone of each word;

[0105] S303: Check the generated text according to the semantic analysis result to ensure that the generated text contains the core theme and intention of the natural language text to be processed.

[0106] The second model generates text in a word-by-word manner. Each step generates a word, and the next word is determined based on the current generated word sequence and the input sentiment and semantic information. The second model continuously uses the previous context (i.e., the words that have been generated and the input information) to adjust the generation of the next word. This mechanism ensures that the generated text is coherent and consistent. In each generated word, the second model adjusts its output based on the sentiment label. For example, when dealing with positive sentiment labels, the second model may tend to choose more positive vocabulary and sentence patterns to ensure that the final generated text can convey the expected sentiment. The second model uses the results of semantic analysis to ensure that the generated text is logically and semantically consistent with the user input. It achieves this goal by checking whether the generated text contains the core theme and intent of the input text and whether it responds to these contents in an appropriate way. Semantic consistency helps the model generate text that is not only emotionally appropriate, but also content-related to the topic of the user's communication. After a series of iterations and adjustments, the model generates a complete natural language text. This text is output as a response to the user's input.

[0107] It is worth noting that both the first model and the second model in the present invention have the following characteristics: both are based on the Transformer architecture, but compared with the traditional Transformer architecture, the model in the present invention only retains the Decoder module, omits the Encoder module, and uses RoPE position encoding on the Query matrix and the Key matrix, which is conducive to better capturing the sequential position information of the text.

[0108] In the working process of the thought chain model and the reply model in the present invention, the user inputs text data, which is processed by the model and finally outputs the result. The principle of this process is as follows:

[0109] First, the tokenizer splits the sentence (text) into a corresponding integer sequence according to the tokenization table. Subsequently, the Embedding (embedding layer) converts the integer sequence into individual digital vectors, which store the position information of each token in the semantic space. Subsequently, the RMSNorm (root mean square layer normalization) component is responsible for normalizing the input vector data. After being processed by RMSNorm, the vectors have the same scale and distribution, which helps the model converge faster during training. Before the self-attention layer, rotary position encoding (RoPE) is applied to the Query matrix and the Key matrix to better capture the position information in the sequence and improve the model's expressive ability. Subsequently, the data passes through the self-attention layer; self-attention means that the data itself serves as both the query vector and the key vector; then the attention weights are calculated through matrix multiplication, where the query vector, key vector, and value vector are all learnable weights. The resulting attention weight matrix is multiplied by the value matrix through matrix multiplication and then undergoes a Softmax operation to ensure that the matrix weights sum to 1 and each weight value is between 0 and 1. Finally, the output of the self-attention layer is obtained. After calculating the output of the self-attention mechanism, the model performs a residual connection with the input and then passes it to the feed-forward neural network (FFN) to calculate the output; the feed-forward neural network consists of two linear transformations and an activation function. In this case, the input vector first passes through a linear transformation, then the activation function (such as SwiGLU) is applied, and finally another linear transformation is performed to generate the final output vector. Finally, the vector undergoes a Softmax operation to obtain the probability of the next possible word, and the sum of the probabilities is 1. And so on, the results of the next and the next words are obtained until the results of all words are calculated.

[0110] The following uses an example to specifically illustrate the processing process of the present invention:

[0111] User input: Last night I tried to bake a sausage and cheese cake, but it tasted terrible. Then the sentiment attribute obtained by the first model is "negative", the keyword is "terrible", the semantic analysis result is "tried a baked good with a bad taste", and the sentiment label is "disgust, sadness". Based on the content input by the user, the semantic analysis result and sentiment label obtained by the first model, the second model finally outputs: "That sounds really bad, but remember: Failure is the mother of success."

[0112] From the above example, it can be seen that by using the method of the present invention, when a user inputs a natural language text, text data results with empathy and capable of emotionally responding to the user input can be finally output, without the need to input guiding prompt words or instructions as in the prior art to extract and stimulate the existing emotion processing content in the model. Compared with the prior art, the operation steps are effectively simplified and the engineering implementation difficulty is reduced.

[0113] An embodiment of the present invention further provides an emotional dialogue generation device, which is used to execute the steps in any one of the foregoing embodiments of the emotional dialogue generation method. Specifically, please refer to Figure 8 , Figure 8 FIG. shows a schematic block diagram of an emotional dialogue generation device 100 provided by an embodiment of the present application. The emotional dialogue generation device 100 specifically includes:

[0114] An acquisition unit 110, configured to acquire a natural language text to be processed;

[0115] A first model processing unit 120, configured to input the natural language text to be processed into a first model for processing to output a semantic analysis result and an emotion label.

[0116] In one embodiment, the first model processing unit 120 includes:

[0117] An emotion analysis module, configured to perform emotion analysis processing on the natural language text to be processed to obtain an emotion attribute result. A keyword search module, configured to search for words related to the emotion attribute result from the natural language text to be processed to obtain emotion keywords that can express emotions. A semantic analysis module, configured to perform semantic analysis on the natural language text to be processed to obtain a semantic analysis result. An emotion label processing module, configured to analyze and obtain an emotion label by combining the emotion attribute result, the emotion keywords, and the semantic analysis result.

[0118] In one embodiment, the emotion analysis module includes:

[0119] An extraction sub-module, configured to extract emotion-related features from the natural language text to be processed. An emotion attribute analysis sub-module, configured to perform emotion attribute analysis on the emotion-related features to obtain an emotion attribute result.

[0120] In one embodiment, the keyword search module includes:

[0121] A multi-level analysis sub-module, configured to perform multi-level analysis on the natural language text to be processed to identify keywords that may express emotions. A screening sub-module, configured to screen out keywords related to the emotion attribute result from the keywords that may express emotions as the final emotion keywords.

[0122] In one embodiment, the semantic analysis module includes:

[0123] The first recognition sub-module is used to perform word segmentation and syntactic analysis on the natural language text to be processed, so as to identify the basic structure of the text. The first analysis sub-module is used to analyze each word and the position of each word in the text according to the basic structure of the text, so as to obtain the position information of the words. The second analysis sub-module is used to perform analysis and processing according to the basic structure of the text and the position information of the words, so as to obtain the semantic analysis result.

[0124] In one embodiment, the emotion label processing module includes:

[0125] The integration sub-module is used to integrate the emotion attribute result, the emotion keyword and the semantic analysis result to obtain comprehensive emotion description information. The second recognition sub-module is used to identify emotion features according to the comprehensive emotion description information. The matching sub-module is used to match the emotion features with the vocabulary in the built-in emotion label list. The selection sub-module is used to select one or more emotion labels from the emotion label list according to the matching result.

[0126] The second model processing unit 130 is used to input the natural language text to be processed, the semantic analysis result and the emotion label into the second model for processing, so as to output the final natural language text in response to the natural language text to be processed.

[0127] In one embodiment, the second model processing unit 130 includes:

[0128] The setting module is used to set the starting point of text generation and construct the text in a word-by-word generation manner. The adjustment module is used to adjust the emotional tone of the generation of each word by using the emotion label. The detection module is used to check the generated text according to the semantic analysis result to ensure that the generated text contains the core theme and intention of the natural language text to be processed.

[0129] It should be noted that those skilled in the art can clearly understand that the specific implementation processes of the above emotion dialogue generation device 100 and each unit can refer to the corresponding descriptions in the foregoing method embodiments. For the convenience and conciseness of description, they will not be elaborated here.

[0130] The above emotion dialogue generation device can be implemented in the form of a computer program, and the computer program can run on a computer device as Figure 9 shown.

[0131] Please refer to Figure 9 , Figure 9 which is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 700 can be a server. Among them, the server can be an independent server or a server cluster composed of multiple servers.

[0132] Such asFigure 9 As shown, the computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the emotional dialogue generation method as described above are implemented.

[0133] The computer device 700 can be a terminal or a server. The computer device 700 includes a processor 720, a memory, and a network interface 750 connected through a system bus 710. Among them, the memory can include a non-volatile storage medium 730 and an internal memory 740.

[0134] The non-volatile storage medium 730 can store an operating system 731 and a computer program 732. When the computer program 732 is executed, the processor 720 can be made to execute the emotional dialogue generation method.

[0135] The processor 720 is used to provide computing and control capabilities to support the operation of the entire computer device 700.

[0136] The internal memory 740 provides an environment for the operation of the computer program 732 in the non-volatile storage medium 730. When the computer program 732 is executed by the processor 720, the processor 720 can be made to execute the emotional dialogue generation method.

[0137] The network interface 750 is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 9 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device 700 to which the solution of this application is applied. The specific computer device 700 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements. Among them, the processor 720 is used to run the program code stored in the memory to implement the steps of the emotional dialogue generation method.

[0138] Those skilled in the art can understand that Figure 9 the embodiments of the computer device shown in do not constitute a limitation on the specific composition of the computer device. In other embodiments, the computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements. For example, in some embodiments, the computer device may only include a memory and a processor. In such an embodiment, the structures and functions of the memory and the processor are the same as those in Figure 9 the shown embodiment and will not be elaborated here.

[0139] It should be understood that in the embodiments of the present application, the processor 720 may be a central processing unit (CPU), and the processor 720 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0140] In another embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium may be a non-volatile computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the emotional dialogue generation method disclosed in the embodiments of the present invention is implemented.

[0141] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices, apparatuses, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0142] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. Units with the same function can also be aggregated into a single unit. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other may be indirect couplings or communication connections through some interfaces, devices, or units, and may also be electrical, mechanical, or other forms of connection.

[0143] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.

[0144] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0145] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes.

[0146] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for generating emotional dialogue, characterized in that: include: Obtain natural language text to be processed; Inputting the natural language text to be processed into the first model for processing, so as to output a semantic analysis result and a sentiment label; Inputting the natural language text to be processed, the semantic analysis result and the sentiment tag into the second model for processing, so as to output a final natural language text that responds to the natural language text to be processed; The step of inputting the natural language text to be processed into the first model for processing to output a semantic analysis result and a sentiment label includes: Perform sentiment analysis on the natural language text to be processed to obtain sentiment attribute results; Find words related to the emotional attribute results from the natural language text to be processed to obtain emotional keywords that can express emotions; Performing semantic analysis on the natural language text to be processed to obtain semantic analysis results; Combine the sentiment attribute results, sentiment keywords and semantic analysis results to get sentiment labels; The emotional attribute results, emotional keywords and semantic analysis results are combined to obtain emotional tags, including: Integrate the sentiment attribute results, sentiment keywords and semantic analysis results to obtain comprehensive sentiment description information; Identify emotional features based on comprehensive emotional description information; Match sentiment features with words from a built-in list of sentiment tags; According to the matching results, one or more emotion tags are selected from the emotion tag list; The step of inputting the natural language text to be processed, the semantic analysis result and the sentiment tag into the second model for processing to output a final natural language text that responds to the natural language text to be processed comprises: Set the starting point for text generation and construct the text word by word; Use emotional tags to adjust the emotional tone of each word generated; The generated text is checked based on the semantic analysis results to ensure that the generated text contains the core topics and intentions of the natural language text to be processed.

2. The method for generating emotional dialogue according to claim 1, characterized in that: The sentiment analysis process is performed on the natural language text to be processed to obtain the sentiment attribute result, including: Extract sentiment-related features from the natural language text to be processed; Perform sentiment attribute analysis on sentiment-related features to obtain sentiment attribute results.

3. The method for generating emotional dialogue according to claim 1, characterized in that: The step of searching for words related to the emotional attribute result from the natural language text to be processed to obtain emotional keywords that can express emotions includes: Perform multi-level analysis on the natural language text to be processed to identify keywords that may express emotions; Keywords related to emotional attribute results are selected from keywords that may express emotions as the final emotional keywords.

4. The method for generating emotional dialogue according to claim 1, characterized in that: The semantic analysis of the natural language text to be processed to obtain a semantic analysis result includes: Perform word segmentation and grammatical analysis on the natural language text to be processed to identify the basic structure of the text; Analyze each word and its position in the text according to the basic structure of the text to obtain the position information of the word; The basic structure of the text and the position information of the words are analyzed and processed to obtain the semantic analysis results.

5. An emotional dialogue generation device, which, when in operation, executes the emotional dialogue generation method according to any one of claims 1 to 4, characterized in that: include: An acquisition unit, used for acquiring a natural language text to be processed; A first model processing unit, used for inputting the natural language text to be processed into the first model for processing, so as to output a semantic analysis result and a sentiment label; The second model processing unit is used to input the natural language text to be processed, the semantic analysis result and the sentiment tag into the second model for processing, so as to output the final natural language text that responds to the natural language text to be processed.

6. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for generating emotional dialogue as described in any one of claims 1 to 4 is implemented.

7. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, the processor executes the emotional dialogue generation method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Automatic conversation method and system considering emotion

    CN107943974A

  • Multi-round emotional dialogue method based on deep learning

    CN108874972A

  • Intelligent question answering method, device and equipment based on emotion recognition and storage medium

    CN114999533A