Conversation reply generation method and device, equipment and storage medium

Through the dialogue reply generation method, emotion recognition and cognitive distortion analysis technology are used to generate replies for users' negative distortion thinking, solving the problem of high cost of traditional psychological counseling and achieving efficient mental health assistance.

CN120067286APending Publication Date: 2025-05-30PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411795160.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, psychological counselors use cognitive behavioral therapy to help clients recognize and challenge negatively distorted thinking methods, and the cost of time and money is high and difficult to popularize.

Method used

By obtaining user discourse, conducting emotional recognition and cognitive distortion analysis, and generating dialogue replies based on preset response strategies, helping users soften negative distortion thinking and establish a positive perspective.

Benefits of technology

Reducing the time and money cost of conversations provides an efficient way to help users deal with negative distorted thinking and achieve the improvement of mental health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067286A_ABST
    Figure CN120067286A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides a dialogue reply generation method and device, equipment and a storage medium, and the method comprises the steps: obtaining a user utterance; the user utterance is input into an emotion recognizer for emotion recognition, an emotion recognition result of the user utterance is obtained, the emotion recognition result comprises a first recognition result, and the first recognition result shows that the user utterance is an utterance with emotion; if the recognition result is the first recognition result, the user utterance is input into a cognitive distortion classifier for cognitive distortion analysis, a cognitive distortion analysis result of the user utterance is obtained, the cognitive distortion analysis result comprises a first analysis result, and the first analysis result shows that the user utterance is a cognitive distortion utterance; and if the result is the first analysis result, generating a reply utterance corresponding to the user utterance based on a preset first reply strategy, thereby reducing the dialogue cost. The invention further relates to a block chain technology. User utterances and reply utterances can be stored in block chain nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, device, and storage medium for generating dialogue responses. Background Art

[0002] Nowadays, various factors such as the fast pace of life and the constantly changing external environment have increased people's stress, made them in a bad mood, and generated negative distorted thinking, which affects people's mental health. At present, for the problem of mental sub-health, people mainly seek help from psychological counselors. The psychological counselors help the clients to more clearly understand their emotional feelings through dialogue and help the clients establish new ways of thinking. Among them, for negative distorted thinking, mainly the psychological counselors use cognitive behavioral therapy to help the clients perceive and challenge these negative distorted thinking. Such a method has relatively high time and money costs. Summary of the Invention

[0003] This application provides a method, apparatus, device, and storage medium for generating dialogue responses, aiming to reduce the cost of dialogue.

[0004] To achieve the above object, this application provides a method for generating dialogue responses. The method for generating dialogue responses includes:

[0005] Obtain a user's utterance;

[0006] Input the user's utterance into an emotion recognizer for emotion recognition to obtain an emotion recognition result of the user's utterance. The emotion recognition result includes a first recognition result and a second recognition result. The first recognition result indicates that the user's utterance is an emotional utterance, and the second recognition result indicates that the user's utterance is a non-emotional utterance;

[0007] If the emotion recognition result is the first recognition result, input the user's utterance into a cognitive distortion classifier for cognitive distortion analysis to obtain a cognitive distortion analysis result of the user's utterance. The cognitive distortion analysis result includes a first analysis result and a second analysis result. The first analysis result indicates that the user's utterance is a cognitively distorted utterance, and the second analysis result indicates that the user's utterance is not a cognitively distorted utterance;

[0008] If the cognitive distortion analysis result is the first analysis result, generate a response utterance corresponding to the user's utterance based on a preset first response strategy.

[0009] In addition, to achieve the above object, this application also provides a device for generating dialogue responses. The device for generating dialogue responses includes:

[0010] An obtaining module, configured to obtain a user's utterance;

[0011] The first recognition module is used to input the user's utterance into an emotion recognizer for emotion recognition, and obtain the emotion recognition result of the user's utterance. The emotion recognition result includes a first recognition result and a second recognition result. The first recognition result indicates that the user's utterance is an emotional utterance, and the second recognition result indicates that the user's utterance is a non-emotional utterance;

[0012] The second recognition module is used to, if the emotion recognition result is the first recognition result, input the user's utterance into a cognitive distortion classifier for cognitive distortion analysis, and obtain the cognitive distortion analysis result of the user's utterance. The cognitive distortion analysis result includes a first analysis result and a second analysis result. The first analysis result indicates that the user's utterance is a cognitively distorted utterance, and the second analysis result indicates that the user's utterance is not a cognitively distorted utterance;

[0013] The reply generation module is used to, if the cognitive distortion analysis result is the first analysis result, generate a reply utterance corresponding to the user's utterance based on a preset first reply strategy.

[0014] In addition, to achieve the above object, the present application further provides a computer device, which includes a memory and a processor;

[0015] The memory is used to store a computer program;

[0016] The processor is used to execute the computer program and, when executing the computer program, implement the dialogue reply generation method as described above.

[0017] In addition, to achieve the above object, the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the above dialogue reply generation method are implemented.

[0018] The present application discloses a dialogue reply generation method, device, equipment and storage medium. By obtaining a user's utterance, inputting the user's utterance into an emotion recognizer for emotion recognition, obtaining the emotion recognition result of the user's utterance, the emotion recognition result includes a first recognition result, the first recognition result indicates that the user's utterance is an emotional utterance, if it is the first recognition result, input the user's utterance into a cognitive distortion classifier for cognitive distortion analysis, obtaining the cognitive distortion analysis result of the user's utterance, the cognitive distortion analysis result includes a first analysis result, the first analysis result indicates that the user's utterance is a cognitively distorted utterance, if it is the first analysis result, generate a reply utterance corresponding to the user's utterance based on a preset first reply strategy, thereby helping the user soften negative distorted thinking and establish a positive perspective. Compared with the traditional human consultation dialogue method, the time and money required by the user are greatly reduced, that is, the cost of the dialogue is reduced. Brief Description of the Drawings

[0019] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 It is a schematic flowchart of the steps of a method for generating a dialogue response provided by an embodiment of the present application;

[0021] Figure 2 It is a schematic block diagram of the structure of a dialogue software system provided by an embodiment of the present application;

[0022] Figure 3 It is a schematic diagram of implementing an input interpretation and recognition task based on the Roberta model provided by an embodiment of the present application;

[0023] Figure 4 It is a schematic diagram of implementing a response generation task based on the CDial-GPT model provided by an embodiment of the present application;

[0024] Figure 5 It is a schematic flowchart of the steps of another method for generating a dialogue response provided by an embodiment of the present application;

[0025] Figure 6 It is a schematic flowchart of the steps of generating a response corresponding to the user's utterance based on a preset second response strategy if the cognitive distortion analysis result is the second analysis result provided by an embodiment of the present application;

[0026] Figure 7 It is a schematic flowchart of the steps of another method for generating a dialogue response provided by an embodiment of the present application;

[0027] Figure 8 It is a schematic diagram of the processing flow of a dialogue based on a dialogue software system provided by an embodiment of the present application;

[0028] Figure 9 It is a schematic block diagram of a dialogue response generation device provided by an embodiment of the present application;

[0029] Figure 10 It is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. Detailed Description of the Embodiments

[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0031] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may change according to the actual situation.

[0032] It should be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0033] It should also be understood that the term "and / or" used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.

[0034] Currently, various factors such as the fast pace of life and the constantly changing external environment have increased people's stress, made them in a bad mood, generated negative distorted thinking, and affected people's mental health. At present, for the problem of mental sub-health, mainly seeking help from a psychological counselor. The psychological counselor helps the client know their emotional feelings more clearly through dialogue and helps the client establish a new way of thinking. Among them, for negative distorted thinking, mainly the psychological counselor uses cognitive behavioral therapy to help the client become aware of and challenge these negative distorted thinking. According to the data of One Psychology, the price of a single consultation is usually starting from 300 yuan to 400 yuan, usually once a week for 50 minutes each time. Such a method has relatively high time and money costs and is not easy to popularize.

[0035] To solve the above problems, embodiments of the present application provide a method, device, equipment, and storage medium for generating dialogue responses to achieve reducing the cost of dialogue.

[0036] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for generating dialogue responses provided by an embodiment of the present application. This method can be applied to a computer device, and the application scenario of this method is not limited in the present application.

[0037] As Figure 1As shown in the figure, the method for generating dialogue responses specifically includes steps S101 to S104.

[0038] S101. Obtain the user's utterance.

[0039] First, construct a dialogue software system. After installing the dialogue software system on products such as computer devices, users can have conversations through the dialogue software system. When the dialogue software system receives the user's utterance input by the user, it automatically replies to the user's utterance. Among them, the user's utterance includes but is not limited to utterances of the voice type, text type, etc., and is not specifically limited in this application.

[0040] Exemplarily, if the user's utterance is of the voice type, after receiving the user's voice-type utterance, perform speech-to-text processing on the voice-type utterance to obtain the text-type utterance corresponding to the voice-type utterance.

[0041] S102. Input the user's utterance into an emotion recognizer for emotion recognition to obtain the emotion recognition result of the user's utterance. The emotion recognition result includes a first recognition result and a second recognition result. The first recognition result indicates that the user's utterance is an emotional utterance, and the second recognition result indicates that the user's utterance is a non-emotional utterance.

[0042] Taking the user's text-type utterance as an example, after the user inputs the user's utterance based on the dialogue software system, input the user's utterance into the emotion recognizer to perform keyword extraction processing on the user's utterance, and perform emotion recognition on the user's utterance through the extracted keywords. The emotion recognition result output by the emotion recognizer includes a first recognition result and a second recognition result. Exemplarily, the first recognition result is "1" and the second recognition result is "0". Of course, the first recognition result and the second recognition result can also be other expressions, which are not limited in the embodiments of this application. Among them, the first recognition result indicates that the user's utterance is an emotional utterance, and the second recognition result indicates that the user's utterance is a non-emotional utterance.

[0043] Exemplarily, the emotion recognition result output by the emotion recognizer further includes an emotion type. Among them, the emotion type includes but is not limited to happy, calm, angry, frustrated, tired, sad, anxious, etc.

[0044] Exemplarily, such as Figure 2As shown in the figure, the dialogue software system includes an encoder, an emotion cause judgment device, an emotion recognition and classification device, etc. The obtained user's speech is encoded by the encoder and then input into the emotion cause judgment device. The emotion cause judgment device determines whether there are emotion cause words in the user's speech. If there are emotion cause words in the user's speech, it outputs a judgment result that the user's speech is an emotional speech; otherwise, if there are no emotion cause words in the user's speech, it outputs a judgment result that the user's speech is a non-emotional speech. If the user's speech is an emotional speech, the emotion recognition and classification device determines the emotion type corresponding to the user's speech.

[0045] For example, if the user's speech input by the user is "I feel a bit lonely recently", input this user's speech into the emotion recognition and classification device, and the emotion recognition and classification device outputs that the emotion type corresponding to this user's speech is depression.

[0046] For another example, if the user's speech input by the user is "We broke up", input this user's speech into the emotion recognition and classification device, and the emotion recognition and classification device outputs that the emotion type corresponding to this user's speech is sadness.

[0047] S103. If the emotion recognition result is the first recognition result, input the user's speech into the cognitive distortion classification device for cognitive distortion analysis to obtain the cognitive distortion analysis result of the user's speech. The cognitive distortion analysis result includes a first analysis result and a second analysis result. The first analysis result indicates that the user's speech is a cognitively distorted speech, and the second analysis result indicates that the user's speech is not a cognitively distorted speech.

[0048] Generally, negative distorted thinking is often taken for granted by the parties concerned, so it is very difficult to detect it by oneself. Generally, it requires the judgment and assistance of others to have the opportunity to detect it. Therefore, in order to take care of this part of users at the same time, if the user's speech is an emotional speech, further cognitive distortion analysis is performed on the user's speech. If the user's speech is a cognitively distorted speech, that is to say, the user has negative distorted thinking, then help the user soften the negative distorted thinking and establish a positive perspective.

[0049] Exemplarily, as Figure 2 shown in the figure, the dialogue software system further includes a cognitive distortion classification device. Input the user's speech into the cognitive distortion classification device, perform cognitive distortion analysis on the user's speech, and obtain the cognitive distortion analysis result corresponding to the user's speech.

[0050] Exemplarily, the cognitive distortion analysis results include but are not limited to the first analysis result, the second analysis result, the cognitive distortion type, etc. Exemplarily, the first analysis result is "yes", and the second analysis result is "no". Of course, the first analysis result and the second analysis result can also be other expressions, and the embodiments of the present application are not limited thereto. Among them, the first analysis result indicates that the user's speech is a cognitively distorted speech, and the second analysis result indicates that the user's speech is not a cognitively distorted speech.

[0051] Types of cognitive distortions include but are not limited to emotionality, overgeneralization, psychological filters, should statements, black-and-white, jumping to conclusions, exaggeration, personalization, labeling, etc. Among them, emotionality refers to using emotional reasoning and believing "I feel this way, so it must be this way"; overgeneralization refers to drawing negative conclusions based on limited negative experiences; psychological filters refer to focusing only on the negative part and ignoring the positive part; should statements refer to expecting things or personal behaviors to be a certain way; black-and-white refers to extreme binary thinking, believing that any imperfection is a failure; jumping to conclusions refers to concluding that others deny oneself without factual basis, and always expecting things to go to the worst result; exaggeration refers to exaggerating certain behaviors or results; personalization refers to attributing things that cannot be controlled by oneself to oneself; labeling refers to labeling oneself or others.

[0052] For example, the types of cognitive distortions are shown in Table 1:

[0053] Table 1

[0054]

[0055] In some embodiments, inputting the user speech into a cognitive distortion classifier for cognitive distortion analysis to obtain a cognitive distortion analysis result of the user speech includes:

[0056] Performing semantic recognition on the user's speech by the cognitive distortion classifier, and judging whether the user's cognition belongs to one of multiple cognitive distortion types according to the semantic recognition result;

[0057] If yes, then the first analysis result is obtained;

[0058] If not, the second analysis result is obtained.

[0059] Semantically recognize the user's utterance through a cognitive distortion classifier, and based on the semantic recognition result, determine whether the user's cognition belongs to one of emotional, overgeneralization, mental filter, should statement, black-and-white thinking, jumping to conclusions, exaggeration, personalization, labeling, etc. For example, if it is determined according to the semantic recognition result that the user's cognition belongs to the cognitive distortion type of exaggeration, a first analysis result is obtained, that is, the user's utterance is a cognitive distortion utterance. On the contrary, if it is determined according to the semantic recognition result that the user's cognition does not belong to any of emotional, overgeneralization, mental filter, should statement, black-and-white thinking, jumping to conclusions, exaggeration, personalization, labeling, etc., a second analysis result is obtained, that is, the user's utterance is not a cognitive distortion utterance.

[0060] S104. If the cognitive distortion analysis result is the first analysis result, generate a reply utterance corresponding to the user's utterance based on a preset first reply strategy.

[0061] Among them, the first reply strategy is proposed based on positive psychology theory, which is difficult for ordinary people to use naturally without training. It is a positive reply strategy. If it is determined that the user's utterance is a cognitive distortion utterance, a reply utterance corresponding to the user's utterance is generated based on the positive first reply strategy, so as to help the user break negative distorted thinking, establish a positive perspective, and achieve the effect of physical and mental regulation.

[0062] Exemplarily, as Figure 2 shown, the dialogue software system further includes a decoder. After generating a reply utterance corresponding to the user's utterance based on the positive first reply strategy, it is decoded by the decoder and then the reply utterance is output.

[0063] For example, if the user's input utterance is "This is the worst week of my life, and everything is not going according to plan", by performing emotion recognition on this user's utterance, the emotion type is determined to be depression. Then, through the cognitive distortion classifier, cognitive distortion analysis is performed on this user's utterance, and it is determined that this user's utterance is a cognitive distortion utterance. The reply utterance corresponding to the user's utterance generated based on the positive first reply strategy is "It's not easy. You've had a very challenging week".

[0064] Exemplarily, the first reply strategy includes multiple positive reply sub-strategies, including but not limited to growth belief, neutrality, acceptance, optimism, self-affirmation, gratitude, etc. This application embodiment is not limited to this.

[0065] Exemplarily, each type of cognitive distortion matches a corresponding positive response sub-strategy. If the user's utterance is a cognitively distorted one, the cognitive distortion type of the user's utterance is identified by the cognitive distortion classifier. Then, based on the cognitive distortion type of the user's utterance, a positive response sub-strategy that matches the cognitive distortion type of the user's utterance is selected from various positive response sub-strategies such as growth belief, neutrality, acceptance, optimism, self-affirmation, gratitude, etc. as the target positive response sub-strategy. Then, a response utterance corresponding to the user's utterance is generated based on the target positive response sub-strategy.

[0066] For example, if the cognitive distortion type of the user's utterance is identified as exaggeration by the cognitive distortion classifier, the target positive response sub-strategy can be determined as optimism, and a response utterance corresponding to the user's utterance is generated based on the target positive response sub-strategy of optimism.

[0067] Another example, if the cognitive distortion type of the user's utterance is identified as overgeneralization by the cognitive distortion classifier, the target positive response sub-strategy can be determined as neutrality, and a response utterance corresponding to the user's utterance is generated based on the target positive response sub-strategy of neutrality.

[0068] Exemplarily, the dialogue software system can use Roberta (Robustly Optimized BERT Pretraining Approach, a natural language processing model) to implement the input interpretation and recognition task. The Roberta model has a 24-layer Transformer Encoder structure. As Figure 3 shown, the Roberta model includes an encoding layer, a multi-head attention layer, layer normalization, a feed-forward network, a max pooling layer, an emotion recognition linear layer, an emotion cause extraction linear layer, an emotion event detection linear layer, a cognitive distortion recognition linear layer, a response strategy prediction linear layer, a softmax (normalized exponential) function, etc. The emotion cause corresponding to the user's utterance can be obtained through the emotion cause extraction linear layer. The emotion type of the user's utterance can be identified through the emotion recognition linear layer. The specific event corresponding to the user's utterance can be identified through the emotion event detection linear layer. The user's utterance can be recognized for cognitive distortion through the cognitive distortion recognition linear layer. The response strategy corresponding to the user's utterance can be predicted through the response strategy prediction linear layer. For example, if the user inputs the user's utterance "There is so much homework and it's so annoying", the emotion type recognition result corresponding to this user's utterance obtained through the Roberta model is angry, the emotion cause is indignation, the event recognition result is yes, the cognitive distortion recognition result is yes, and the predicted response strategy result corresponding to the user's utterance is neutral.

[0069] Exemplarily, the dialogue software system can use CDial-GPT (Chinese dialogue pre-training model) to implement the response generation task. The CDial-GPT model has a 12-layer Transformer Decoder structure. As Figure 4 shown, the CDial-GPT model includes an encoding layer, a multi-head attention layer, layer normalization, a feed-forward network, a softmax function, etc. The result of input interpretation and recognition task of the Roberta model is input into the CDial-GPT model, and the response is generated through the CDial-GPT model. For example, still taking the above-listed example, input "There are so many assignments and it's so annoying [Label][angry][furious][cognitive distortion-yes][event-yes][neutral]" into the CDial-GPT model, and based on the neutral positive response sub-strategy, generate the corresponding response and output the response.

[0070] In some embodiments, as Figure 5 shown, after step S103, step S105 may be included.

[0071] S105. If the cognitive distortion analysis result is the second analysis result, generate a response corresponding to the user's utterance based on a preset second response strategy.

[0072] If it is identified by the cognitive distortion classifier that the user's utterance is not a cognitively distorted utterance, a second response strategy different from the first response strategy is used to generate a response corresponding to the user's utterance. Among them, the second response strategy is an empathic response strategy that can be naturally used by most people.

[0073] Exemplarily, the second response strategy includes a variety of empathic response sub-strategies, including but not limited to sympathy, concern, advice, blessing, comfort, agreement, encouragement, sharing views, sharing experiences, etc. This is not limited in the embodiments of the present application.

[0074] In some embodiments, as Figure 6 shown, step S105 may include sub-step S1051 and sub-step S1052.

[0075] S1051. If the cognitive distortion analysis result is the second analysis result, determine the target empathic response sub-strategy corresponding to the user's utterance from a variety of the empathic response sub-strategies;

[0076] S1052. Generate the response based on the target empathic response sub-strategy.

[0077] If the user's utterance is not cognitively distorted, a corresponding empathy response sub-strategy is selected from a variety of empathy response sub-strategies such as sympathy, suggestion, blessing, approval, encouragement, etc. as the target empathy response sub-strategy, and then a response utterance corresponding to the user's utterance is generated based on the target empathy response sub-strategy.

[0078] Exemplarily, the conversation reply generating method further includes:

[0079] If the emotion recognition result is the first recognition result, determining the emotion type corresponding to the user speech;

[0080] If the cognitive distortion analysis result is the second analysis result, determining a target empathy response sub-strategy corresponding to the user's speech from the multiple empathy response sub-strategies includes:

[0081] If the cognitive distortion analysis result is the second analysis result, the target empathy response sub-strategy is determined from the plurality of empathy response sub-strategies according to the emotion type corresponding to the user utterance.

[0082] If the user's utterance is not cognitively distorted, the emotion type corresponding to the user's utterance is identified through the emotion recognition classifier. Then, based on the emotion type corresponding to the user's utterance, an empathy response sub-strategy that matches the emotion type of the user's utterance is selected from a variety of empathy response sub-strategies such as sympathy, suggestion, blessing, approval, encouragement, etc. as the target empathy response sub-strategy. Then, a reply utterance corresponding to the user's utterance is generated based on the target empathy response sub-strategy.

[0083] For example, if the emotion type corresponding to the user's speech is identified as frustration through the emotion recognition classifier, the target empathy response sub-strategy can be determined to be encouragement, and a response speech corresponding to the user's speech is generated based on the target empathy response sub-strategy of encouragement.

[0084] For another example, if the emotion type corresponding to the user's speech is identified as sadness through the emotion recognition classifier, the target empathy response sub-strategy can be determined to be comfort, and a response speech corresponding to the user's speech is generated based on the target empathy response sub-strategy of comfort.

[0085] In some embodiments, Figure 7 As shown, step S102 may include step S106.

[0086] S106: If the emotion recognition result is the second recognition result, generating a reply utterance corresponding to the user utterance based on a third reply strategy, wherein the third reply strategy includes a reply strategy of asking questions or listening.

[0087] For example, if the user's input is "I'm very idle now", and the emotion reason detector determines that the user's utterance does not contain emotion reason words, and the user's utterance is a non-emotional one. At this time, a response utterance such as "Do you want to chat with me?" is generated based on the response strategy of asking questions.

[0088] For example, taking a chat statement with the user's utterance as the text type as an example, as Figure 8 shown, the processing flow of the dialogue based on the dialogue software system is as follows:

[0089] 1. The user inputs a chat statement.

[0090] 2. Determine whether there is an emotion reason in the statement; if not, execute step 3; if so, execute step 4.

[0091] 3. Generate a corresponding response statement based on the response strategy of in-depth questioning or active listening.

[0092] 4. Determine whether the statement is a cognitive distortion statement; if not, execute step 5; if so, execute step 6.

[0093] 5. Generate a corresponding response statement based on the empathy response strategy.

[0094] 6. Generate a corresponding response statement based on the positive response strategy.

[0095] For example, still taking the user's input "I've been feeling a bit lonely lately" as an example, the identified emotion type corresponding to this user's utterance is depression, and the corresponding response statement generated based on the empathy response strategy is "I'm here with you. Do you want to chat with me?"

[0096] Another example, still taking the user's input "We broke up" as an example, the identified emotion type corresponding to this user's utterance is sadness, and the corresponding response statement generated based on the empathy response strategy is "Are you okay? Did you have a fight?"

[0097] Another example, taking the user's input "I feel like I'm not good at socializing and can't make friends" as an example, the identified cognitive distortion type corresponding to this user's utterance is exaggeration, and response statements such as "You're not alone. I'm here with you" and "Get out more and do activities with people you know. Maybe you can make friends" are generated based on the positive response strategy.

[0098] Based on the empathy dialogue of the dialogue software system, for inputs with negative distorted thinking, positive responses that do not deny the user's original meaning are generated based on the positive response strategy, helping the user soften the negative distorted thinking, establish a positive perspective, achieve the role of physical and mental regulation, and at the same time reduce the time and money costs required for traditional human consultation.

[0099] In the above embodiments, by obtaining the user's utterance, inputting the user's utterance into an emotion recognizer for emotion recognition, obtaining the emotion recognition result of the user's utterance, the emotion recognition result includes a first recognition result, and the first recognition result indicates that the user's utterance is an emotional utterance. If it is the first recognition result, then input the user's utterance into a cognitive distortion classifier for cognitive distortion analysis, obtain the cognitive distortion analysis result of the user's utterance, the cognitive distortion analysis result includes a first analysis result, and the first analysis result indicates that the user's utterance is a cognitively distorted utterance. If it is the first analysis result, then generate a reply utterance corresponding to the user's utterance based on a preset first reply strategy, thereby helping the user soften negative distorted thinking and establish a positive perspective. Compared with the traditional human consultation dialogue method, the time and money required by the user are greatly reduced, that is, the cost of the dialogue is reduced.

[0100] Please refer to Figure 9 , Figure 9 which is a schematic block diagram of a dialogue reply generation device provided by an embodiment of the present application. The dialogue reply generation device can be configured in a computer device and is used to execute the foregoing dialogue reply generation method.

[0101] As Figure 9 shown, the dialogue reply generation device 1000 includes: an acquisition module 1001, a first recognition module 1002, a second recognition module 1003, and a reply generation module 1004.

[0102] The acquisition module 1001 is used to acquire the user's utterance.

[0103] The first recognition module 1002 is used to input the user's utterance into an emotion recognizer for emotion recognition, obtain the emotion recognition result of the user's utterance. The emotion recognition result includes a first recognition result and a second recognition result. The first recognition result indicates that the user's utterance is an emotional utterance, and the second recognition result indicates that the user's utterance is a non-emotional utterance.

[0104] The second recognition module 1003 is used to, if the emotion recognition result is the first recognition result, input the user's utterance into a cognitive distortion classifier for cognitive distortion analysis, obtain the cognitive distortion analysis result of the user's utterance. The cognitive distortion analysis result includes a first analysis result and a second analysis result. The first analysis result indicates that the user's utterance is a cognitively distorted utterance, and the second analysis result indicates that the user's utterance is not a cognitively distorted utterance.

[0105] The reply generation module 1004 is used to, if the cognitive distortion analysis result is the first analysis result, generate a reply utterance corresponding to the user's utterance based on a preset first reply strategy.

[0106] In one embodiment, the second recognition module 1003 is further configured to:

[0107] Perform semantic recognition on the user's utterance through the cognitive distortion classifier, and determine whether the user's cognition belongs to one of multiple cognitive distortion types according to the semantic recognition result;

[0108] If so, obtain the first analysis result;

[0109] If not, obtain the second analysis result.

[0110] In one embodiment, the cognitive distortion analysis result further includes cognitive distortion types, there are multiple cognitive distortion types, the first reply strategy includes multiple positive reply sub-strategies, and each cognitive distortion type matches one positive reply sub-strategy.

[0111] In one embodiment, the reply generation module 1004 is further configured to:

[0112] If the cognitive distortion analysis result is the second analysis result, generate a reply utterance corresponding to the user's utterance based on a preset second reply strategy.

[0113] In one embodiment, the second reply strategy includes multiple empathy reply sub-strategies, and the reply generation module 1004 is further configured to:

[0114] If the cognitive distortion analysis result is the second analysis result, determine the target empathy reply sub-strategy corresponding to the user's utterance from multiple empathy reply sub-strategies;

[0115] Generate the reply utterance based on the target empathy reply sub-strategy.

[0116] In one embodiment, the reply generation module 1004 is further configured to:

[0117] If the emotion recognition result is the first recognition result, determine the emotion type corresponding to the user's utterance;

[0118] If the cognitive distortion analysis result is the second analysis result, determine the target empathy reply sub-strategy from multiple empathy reply sub-strategies according to the emotion type corresponding to the user's utterance.

[0119] In one embodiment, the reply generation module 1004 is further configured to:

[0120] If the emotion recognition result is the second recognition result, generate a reply utterance corresponding to the user's utterance based on a third reply strategy, and the third reply strategy includes a reply strategy of asking questions or listening.

[0121] Among them, each module in the above-mentioned dialogue response generation device 1000 corresponds to each step in the above-mentioned dialogue response generation method embodiment, and its functions and implementation processes will not be elaborated here one by one.

[0122] The dialogue response generation device 1000 can execute the dialogue response generation method provided by the embodiments of the present application. Therefore, the beneficial effects that can be achieved by the dialogue response generation method provided by the embodiments of the present application can be realized. For details, please refer to the previous embodiments and will not be elaborated here.

[0123] The methods and devices of the present application can be used in many general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0124] Exemplarily, the above-mentioned method and device can be implemented in the form of a computer program, and this computer program can run on a computer device as Figure 10 shown.

[0125] Please refer to Figure 10 , Figure 10 which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application.

[0126] Please refer to Figure 10 , the computer device includes a processor and a memory connected through a system bus. Among them, the memory can include a non-volatile storage medium and an internal memory.

[0127] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0128] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When this computer program is executed by the processor, the processor can execute any dialogue response generation method.

[0129] It should be understood that the processor can be a Central Processing Unit (CPU), and the processor can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0130] Among them, in one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:

[0131] Obtain the user's utterance;

[0132] Input the user's utterance into an emotion recognizer for emotion recognition to obtain the emotion recognition result of the user's utterance. The emotion recognition result includes a first recognition result and a second recognition result. The first recognition result indicates that the user's utterance is an emotional utterance, and the second recognition result indicates that the user's utterance is a non-emotional utterance;

[0133] If the emotion recognition result is the first recognition result, input the user's utterance into a cognitive distortion classifier for cognitive distortion analysis to obtain the cognitive distortion analysis result of the user's utterance. The cognitive distortion analysis result includes a first analysis result and a second analysis result. The first analysis result indicates that the user's utterance is a cognitively distorted utterance, and the second analysis result indicates that the user's utterance is not a cognitively distorted utterance;

[0134] If the cognitive distortion analysis result is the first analysis result, generate a response utterance corresponding to the user's utterance based on a preset first response strategy.

[0135] In one embodiment, when the processor implements inputting the user's utterance into a cognitive distortion classifier for cognitive distortion analysis to obtain the cognitive distortion analysis result of the user's utterance, it is used to implement:

[0136] Perform semantic recognition on the user's utterance through the cognitive distortion classifier, and judge whether the user's cognition belongs to one of multiple cognitive distortion types according to the semantic recognition result;

[0137] If so, obtain the first analysis result;

[0138] Otherwise, obtain the second analysis result.

[0139] In one embodiment, the cognitive distortion analysis result further includes cognitive distortion types, there are multiple cognitive distortion types, the first reply strategy includes multiple positive reply sub-strategies, and each cognitive distortion type matches one positive reply sub-strategy.

[0140] In one embodiment, after the processor implements that if the emotion recognition result is the first recognition result, input the user's speech into a cognitive distortion classifier for cognitive distortion analysis to obtain the cognitive distortion analysis result of the user's speech, it is used to implement:

[0141] If the cognitive distortion analysis result is the second analysis result, generate a reply speech corresponding to the user's speech based on a preset second reply strategy.

[0142] In one embodiment, the second reply strategy includes multiple empathy reply sub-strategies. When the processor implements that if the cognitive distortion analysis result is the second analysis result, generate a reply speech corresponding to the user's speech based on a preset second reply strategy, it is used to implement:

[0143] If the cognitive distortion analysis result is the second analysis result, determine the target empathy reply sub-strategy corresponding to the user's speech from multiple empathy reply sub-strategies;

[0144] Generate the reply speech based on the target empathy reply sub-strategy.

[0145] In one embodiment, the processor is further used to implement:

[0146] If the emotion recognition result is the first recognition result, determine the emotion type corresponding to the user's speech;

[0147] When the processor implements that if the cognitive distortion analysis result is the second analysis result, determine the target empathy reply sub-strategy corresponding to the user's speech from multiple empathy reply sub-strategies, it is used to implement:

[0148] If the cognitive distortion analysis result is the second analysis result, determine the target empathy reply sub-strategy from multiple empathy reply sub-strategies according to the emotion type corresponding to the user's speech.

[0149] In one embodiment, after the processor implements emotion recognition on the user's speech, it is used to implement:

[0150] If the emotion recognition result is the second recognition result, a reply utterance corresponding to the user's utterance is generated based on a third reply strategy, where the third reply strategy includes a reply strategy of asking questions or listening.

[0151] This computer device can execute the dialogue reply generation method provided in the embodiments of the present application. Therefore, the beneficial effects that can be achieved by the dialogue reply generation method provided in the embodiments of the present application can be realized. For details, refer to the previous embodiments and will not be elaborated here.

[0152] The embodiments of the present application further provide a computer-readable storage medium.

[0153] A computer program is stored on the computer-readable storage medium of the present application. When the computer program is executed by a processor, the steps of the dialogue reply generation method as described above are implemented.

[0154] Among them, the computer-readable storage medium may be an internal storage unit of the dialogue reply generation device or computer device described in the foregoing embodiments, such as the hard disk or memory of the dialogue reply generation device or computer device. The computer-readable storage medium may also be an external storage device of the dialogue reply generation device or computer device, such as a plug-in hard disk equipped on the dialogue reply generation device or computer device, a Smart Media Card (SMC), a Secure Digital Card (SD Card), a Flash Card, etc.

[0155] Since the computer program stored in the computer-readable storage medium can execute any of the dialogue reply generation methods provided in the embodiments of the present application, the beneficial effects that can be achieved by any of the dialogue reply generation methods provided in the embodiments of the present application can be realized. For details, refer to the previous embodiments and will not be elaborated here.

[0156] Further, the computer-readable storage medium may mainly include a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function, etc.; the storage data area may store data created according to the use of the blockchain node, etc.

[0157] The blockchain referred to in the present application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain may include a blockchain underlying platform, a platform product service layer, an application service layer, etc.

[0158] It should be noted that in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or system comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of additional identical elements in the process, method, article or system comprising such element.

[0159] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.

Claims

1. A method for generating a dialogue reply, characterized in that: The dialogue reply generating method comprises: Get user words; Inputting the user speech into an emotion recognizer for emotion recognition, and obtaining an emotion recognition result of the user speech, wherein the emotion recognition result includes a first recognition result and a second recognition result, wherein the first recognition result indicates that the user speech is an emotion-containing speech, and the second recognition result indicates that the user speech is an emotion-free speech; If the emotion recognition result is the first recognition result, the user speech is input into a cognitive distortion classifier for cognitive distortion analysis to obtain a cognitive distortion analysis result of the user speech, the cognitive distortion analysis result including a first analysis result and a second analysis result, the first analysis result indicating that the user speech is a cognitively distorted speech, and the second analysis result indicating that the user speech is not a cognitively distorted speech; If the cognitive distortion analysis result is the first analysis result, a reply utterance corresponding to the user utterance is generated based on a preset first reply strategy.

2. The method for generating a dialogue response according to claim 1, wherein: The step of inputting the user speech into a cognitive distortion classifier for cognitive distortion analysis to obtain a cognitive distortion analysis result of the user speech comprises: Performing semantic recognition on the user's speech by the cognitive distortion classifier, and judging whether the user's cognition belongs to one of multiple cognitive distortion types according to the semantic recognition result; If yes, then the first analysis result is obtained; If not, the second analysis result is obtained.

3. The method for generating a dialogue response according to claim 1, wherein: The cognitive distortion analysis result also includes cognitive distortion types, and there are multiple cognitive distortion types. The first response strategy includes multiple positive response sub-strategies, and each cognitive distortion type matches one positive response sub-strategy.

4. The method for generating a dialogue response according to claim 1, wherein: If the emotion recognition result is the first recognition result, the user speech is input into a cognitive distortion classifier for cognitive distortion analysis. After obtaining the cognitive distortion analysis result of the user speech, the method further includes: If the cognitive distortion analysis result is the second analysis result, a reply utterance corresponding to the user utterance is generated based on a preset second reply strategy.

5. The method for generating a dialogue response according to claim 4, wherein: The second reply strategy includes a plurality of empathy reply sub-strategies. If the cognitive distortion analysis result is the second analysis result, a reply utterance corresponding to the user utterance is generated based on the preset second reply strategy, including: If the cognitive distortion analysis result is the second analysis result, determining a target empathy response sub-strategy corresponding to the user's utterance from the plurality of empathy response sub-strategies; The response utterance is generated based on the target empathy response sub-strategy.

6. The method for generating a dialogue response according to claim 5, wherein: The method further comprises: If the emotion recognition result is the first recognition result, determining the emotion type corresponding to the user speech; If the cognitive distortion analysis result is the second analysis result, determining a target empathy response sub-strategy corresponding to the user's speech from the multiple empathy response sub-strategies includes: If the cognitive distortion analysis result is the second analysis result, the target empathy response sub-strategy is determined from the plurality of empathy response sub-strategies according to the emotion type corresponding to the user utterance.

7. The method for generating a dialogue response according to any one of claims 1 to 6, characterized in that: After the emotion recognition is performed on the user's speech, the method further includes: If the emotion recognition result is the second recognition result, a reply utterance corresponding to the user utterance is generated based on a third reply strategy, where the third reply strategy includes a reply strategy of asking questions or listening.

8. A dialogue reply generating device, characterized in that: The dialogue reply generating device comprises: The acquisition module is used to obtain user speech; A first recognition module, configured to input the user speech into an emotion recognizer for emotion recognition, and obtain an emotion recognition result of the user speech, wherein the emotion recognition result includes a first recognition result and a second recognition result, wherein the first recognition result indicates that the user speech is an emotion-containing speech, and the second recognition result indicates that the user speech is an emotion-free speech; a second recognition module, configured to input the user speech into a cognitive distortion classifier for cognitive distortion analysis if the emotion recognition result is the first recognition result, and obtain a cognitive distortion analysis result of the user speech, wherein the cognitive distortion analysis result includes a first analysis result and a second analysis result, wherein the first analysis result indicates that the user speech is a cognitively distorted speech, and the second analysis result indicates that the user speech is not a cognitively distorted speech; A reply generation module is used to generate a reply utterance corresponding to the user utterance based on a preset first reply strategy if the cognitive distortion analysis result is the first analysis result.

9. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the method for generating a dialogue response according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the dialog reply generating method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Emotional dialogue generation method and device, electronic equipment and storage medium

    CN122264086A