Response generation method and device for generating co-event response and electronic equipment

By constructing an iterative update module within the dialogue, simulating the update of emotional and cognitive states of users and intelligent agents, the problem of difficulty in generating high-quality empathy responses in the prior art is solved, and more accurate state representation and better empathy effects are achieved.

CN120012786AActive Publication Date: 2025-05-16NAT UNIV OF DEFENSE TECH

Patent Information

Application Number
CN202510178096.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-16
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The prior art is difficult to generate a response with better empathy, especially when ignoring the role differences between users and smart agents and the nuances in conversations.

Method used

By constructing an iterative update module within the dialogue, simulate the update of emotional and cognitive states of users and intelligent agents, treat each sentence as a continuous time node, and the status of users and intelligent agents is updated accordingly to simulate their interaction.

Benefits of technology

More accurate cognitive and emotional state representations are achieved, and the significant two-way impacts inherent in human-computer interaction are modeled, thereby obtaining a more empathetic response with better empathy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012786A_ABST
    Figure CN120012786A_ABST
Patent Text Reader

Abstract

The invention discloses a response generation method, and particularly relates to a response generation method and device for generating co-estrus response and electronic equipment, and the method comprises the steps: inputting a historical dialogue context of a user and an intelligent agent into a context coding module for coding processing, and obtaining a context representation; performing intra-dialogue state iterative updating on the emotional state and the cognitive state corresponding to each sentence in the historical dialogue context by using an intra-dialogue state iterative updating learning module to obtain the final emotional state of the user and the final cognitive state of the intelligent agent; and inputting the context representation, the final emotional state of the user and the final cognitive state of the intelligent agent into a response generation module to obtain a target response. According to the method, a response with a better common situation effect can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a response generation method, and in particular to a response generation method, device and electronic equipment for generating empathetic responses. Background Art

[0002] Empathy response generation refers to the ability of a dialogue system to recognize and interpret the user's cognitive and emotional states, thereby designing responses that focus on the user's needs and feelings. As a key task to improve the quality of dialogue, empathy response generation is invaluable in open-domain applications, such as psychological counseling and emotional support.

[0003] Most methods in the prior art treat all utterances as a long continuous sequence to deduce user states at the simulation level. These methods ignore the role distinction between users and intelligent agents, as well as the subtle differences in each utterance, resulting in less accurate emotion perception and state recognition. However, dialogue is a process of mutual adaptation and collaboration, and both parties in the dialogue will affect each other's understanding and response, especially in emotional support dialogues, where there is a significant mutual influence between the two parties in terms of emotions and cognition. Therefore, it is not enough to simply distinguish between self-awareness and other-awareness, that is, modeling the states of users and intelligent agents as independent individuals without considering the cross-influence between them makes it difficult to obtain high-quality responses. In addition, at the dialogue level, methods in the prior art also ignore the changes between dialogues with the same emotional labels. Therefore, it is difficult for the commonly used methods in the prior art to generate responses with good empathy effects. Summary of the invention

[0004] The technical problem to be solved by the present invention is that it is difficult for the commonly used methods in the prior art to generate responses with good empathy effects. In order to solve the above problem, the present invention provides a response generation method, device and electronic device for generating empathy responses.

[0005] The content of the present invention includes:

[0006] In a first aspect, an embodiment of the present invention provides a response generation method for generating an empathy response, comprising:

[0007] The historical conversation context between the user and the intelligent agent is input into the context encoding module for encoding processing to obtain the context representation;

[0008] Using the in-dialogue state iterative update learning module to iteratively update the emotional state and cognitive state corresponding to each utterance in the historical dialogue context, to obtain the final emotional state of the user and the final cognitive state of the intelligent agent;

[0009] The context representation, the user's final emotional state and the intelligent agent's final cognitive state are input into a response generation module to obtain a target response.

[0010] Optionally, before inputting the historical conversation context between the user and the intelligent agent into the context encoding module for encoding processing to obtain the context representation, the method further includes:

[0011] Iteratively training the ININ model based on multiple training dialogue contexts to obtain a training response corresponding to a target training dialogue context, wherein the multiple training dialogue contexts include the target training dialogue context, and the ININ model includes the context encoding module, the in-dialogue state iterative update learning module, the response generation module, and the inter-dialogue emotion adjacency comparison learning module;

[0012] Based on the multiple training dialogue contexts, using the inter-dialogue emotion adjacency contrast learning module to determine the contrast loss in the current iterative training process, and calculating the emotion prediction loss, generation loss and diversity loss in the current iterative training process based on the training responses;

[0013] Determining a final loss value based on the contrast loss, the sentiment prediction loss, the generation loss, and the diversity loss;

[0014] The parameters of the ININ model are adjusted based on the final loss value, and the training is terminated when the training end condition is met. When the training end condition is not met, the next iterative training is performed.

[0015] Optionally, the determining the contrast loss in the current iterative training process based on the multiple training dialogue contexts and using the inter-dialogue emotion adjacency contrast learning module includes:

[0016] Inputting the multiple training dialogue contexts into K parallel enhanced encoders to obtain K enhanced views, where K is an integer greater than 1;

[0017] Generate an emotion adjacency matrix corresponding to the K enhanced views, where the emotion adjacency matrix is ​​used to indicate whether the emotion labels of any two training dialogue contexts in the multiple training dialogue contexts are the same;

[0018] Any one of the K enhanced views is used as a target view, and the contrast loss is calculated based on a neighbor contrast loss between each other enhanced view and the target view.

[0019] Optionally, the contrast loss is:

[0020]

[0021] in, is the x-th enhanced view, the x-th enhanced view is the target view, is the kth enhanced view, The embedding of the i-th training dialogue context learned for the k-th augmented view, The embedding of the jth training conversation context learned for the kth augmentation view, The embedding of the i-th training dialogue context learned for the x-th augmented view, is the embedding of the jth training dialogue context learned for the xth augmented view, N is the number of training sample data, θ(·) is the inner product function measuring similarity, τ is the temperature coefficient, are the positive samples in the two enhanced views, is the dialogue C in the training batch j Collection of Indicates dialogue C j and dialogue C i have the same emotion category.

[0022] Optionally, the step of inputting the historical conversation context between the user and the intelligent agent into a context encoding module for encoding to obtain a context representation includes:

[0023] Connecting all utterances in the historical conversation context to obtain an input sequence, wherein the input sequence includes a special mark for identifying the start of context input in the historical conversation context;

[0024] Processing the input sequence using a word embedding layer and a position embedding layer to obtain a word embedding and a position embedding of the input sequence;

[0025] Generate a final vector representation based on the word embedding, position embedding and role embedding, wherein the role embedding is used to distinguish the dialogue subject in the dialogue context, and the dialogue subject is a user or an intelligent agent;

[0026] The final vector representation is input into a Transformer encoder for processing to obtain the context representation.

[0027] Optionally, the using the in-dialogue state iterative updating learning module to perform in-dialogue state iterative updating on the emotional state and cognitive state corresponding to each utterance in the historical dialogue context to obtain the final emotional state of the user and the final cognitive state of the intelligent agent includes:

[0028] Extracting common sense knowledge from each utterance of the historical conversation context and obtaining a hidden vector of the common sense knowledge, wherein the common sense knowledge includes an emotional state and a cognitive state;

[0029] Using the hidden vector of the common sense knowledge, generating an initial dialogue graph through an image constructor;

[0030] Iteratively updating the cognitive state and the emotional state based on the initial dialogue graph to obtain the cognitive state and the emotional state corresponding to each of the utterances;

[0031] The final emotional state of the user and the final cognitive state of the intelligent agent are determined based on the dialogue subject in the historical dialogue context.

[0032] Optionally, the iterative updating of the cognitive state and the emotional state based on the initial dialogue graph to obtain the cognitive state and the emotional state corresponding to each utterance includes:

[0033] Determine the user’s initial cognitive state The user's initial emotional state Initial cognitive state of the intelligent agent and the initial emotional state of the agent

[0034]

[0035] in, and is randomly initialized. represents the cognitive state of the first utterance in the historical conversation context, Indicates the emotional state of the first sentence in the historical dialogue context, Enc ini Initial encoders for representing the user’s cognitive and emotional states;

[0036] Based on the user's initial cognitive state The user's initial emotional state The initial cognitive state of the intelligent agent and the initial emotional state of the intelligent agent Perform iterative updates to obtain the emotional state and cognitive state corresponding to each sentence in the historical conversation context. The iterative update process of the cognitive state is as follows:

[0037]

[0038] in, For the current discourse U m The cognitive state, m∈(1,M], M is the number of utterances in the historical dialogue context, ⊙ is used to represent element-wise multiplication, σ is used to represent the sigmoid activation function, is the trainable weight matrix of the linear layer, is a trainable bias term;

[0039] The iterative update process of the emotional state is:

[0040]

[0041] in, For the current discourse U m emotional state, b emo is a trainable bias term, W emo is the trainable weight matrix of the linear layer.

[0042] In a second aspect, an embodiment of the present invention provides a response generation device for generating an empathy response, comprising:

[0043] The context encoding module is used to encode the historical conversation context between the user and the intelligent agent to obtain the context representation;

[0044] An intra-dialogue state iterative update learning module is used to iteratively update the intra-dialogue state of the emotional state and cognitive state corresponding to each utterance in the historical dialogue context to obtain the final emotional state of the user and the final cognitive state of the intelligent agent;

[0045] A response generation module is used to obtain a target response based on the context representation, the user's final emotional state and the intelligent agent's final cognitive state.

[0046] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: a memory, a processor, and a program stored in the memory and executable on the processor; the processor is used to read the program in the memory to implement the steps in the response generation method for generating an empathetic response as described in the first aspect.

[0047] In a fourth aspect, an embodiment of the present invention provides a readable storage medium for storing a program, which, when executed by a processor, implements the steps in the response generation method for generating an empathetic response as described in the first aspect.

[0048] The beneficial effect of the present invention is that, in the embodiment of the present invention, an intra-dialogue state iteration update module is constructed to simulate the emotional state and cognitive state update of the user and the intelligent agent, each sentence is regarded as a continuous time node, and the state of the user and the intelligent agent is updated accordingly to simulate their interaction. The intra-dialogue state iteration update module captures the mutual influence between the user and the intelligent agent in a cross-iteration manner, thereby achieving a more accurate representation of cognitive and emotional states, modeling the significant two-way influence inherent in the human-computer interaction process, and obtaining a more empathetic and better empathetic response. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Attached Figure 1 A flow chart of a response generation method for generating an empathic response provided by an embodiment of the present invention;

[0050] Attached Figure 2a A schematic diagram of the framework of the ININ model provided by an embodiment of the present invention;

[0051] Attached Figure 2b for Figure 2a Schematic diagram of the context encoding module in the ININ model provided;

[0052] Attached Figure 2c for Figure 2a A schematic diagram of the in-dialogue state iterative update learning module in the provided ININ model;

[0053] Attached Figure 2d for Figure 2a Schematic diagram of the response generation module in the ININ model provided;

[0054] Attached Figure 2e for Figure 2a The schematic diagram of the inter-dialogue sentiment adjacency comparison learning module in the provided ININ model;

[0055] Attached Figure 3 A schematic diagram of a response generation device for generating an empathic response provided by an embodiment of the present invention;

[0056] Attached Figure 4 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0057] In the embodiments of the present application, the term "and / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. In the embodiments of the present application, the term "multiple" refers to two or more, and other quantifiers are similar. The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first" and "second" are generally of one type, and the number of objects is not limited. For example, the first object can be one or more.

[0058] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0060] The embodiments of the present application provide a response generation method, device and electronic device for generating empathetic responses, which are intended to generate empathetic responses that are full of empathy and have speech-level and dialogue-level understanding capabilities.

[0061] See also Figure 1 , Figure 1 A flow chart of a response generation method for generating an empathy response provided by an embodiment of the present invention, the method specifically comprises the following steps:

[0062] Step 101, inputting the historical conversation context between the user and the intelligent agent into a context encoding module for encoding processing to obtain a context representation;

[0063] Step 102, using an in-dialogue state iterative update learning module to perform in-dialogue state iterative update on the emotional state and cognitive state corresponding to each utterance in the historical dialogue context, to obtain the user's final emotional state and the intelligent agent's final cognitive state;

[0064] Step 103, input the context representation, the user's final emotional state and the intelligent agent's final cognitive state into a response generation module to obtain a target response.

[0065] The response generation method for generating empathic responses provided in the embodiment of the present application can be used to implement the task of generating empathic responses in a multi-round dialogue system. In this embodiment, the task of generating empathic responses in a multi-round dialogue system is converted into a probability solving task of generating responses based on a given historical dialogue context for calculation. Specifically, a given dialogue D can be represented as an M+1 speech sequence D={U 1 ,…,U M+1}. Consider D as two parts (C, Y), where C = {U 1 ,…,U M} represents the conversation context, that is, the historical conversation context, and Y represents the target response U M+1 Each utterance U in the historical conversation context m (m=1,2,…,M) is a string containing any length N m The emotion category (also called sentiment category) e is obtained through empathy knowledge learning. Therefore, the task of generating empathetic responses can be characterized as calculating the probability P(Y|C,e) of generating a response Y based on a given historical conversation context C.

[0066] In the embodiment of the present application, an intra-dialogue state iteration update module is constructed to simulate the emotional state (also called affective state) and cognitive state update of the user and the intelligent agent, treating each utterance as a continuous time node, and updating the state of the user and the intelligent agent accordingly to simulate their interaction. The intra-dialogue state iteration update module captures the mutual influence between the user and the intelligent agent in a cross-iteration manner, thereby achieving a more accurate representation of cognitive and emotional states, modeling the significant two-way influence inherent in the human-computer interaction process, and obtaining a more empathetic and better empathetic response.

[0067] In some embodiments, the context encoding module, the in-dialogue state iterative update learning module and the response generation module are pre-trained so that the response generated by the method has the ability to understand the speech level and the dialogue level. Optionally, before step 101, the method further includes:

[0068] Iteratively training the ININ model based on multiple training dialogue contexts to obtain a training response corresponding to a target training dialogue context, wherein the multiple training dialogue contexts include the target training dialogue context, and the ININ model includes the context encoding module, the in-dialogue state iterative update learning module, the response generation module, and the inter-dialogue emotion adjacency comparison learning module;

[0069] Based on the multiple training dialogue contexts, using the inter-dialogue emotion adjacency contrast learning module to determine the contrast loss in the current iterative training process, and calculating the emotion prediction loss, generation loss and diversity loss in the current iterative training process based on the training responses;

[0070] Determining a final loss value based on the contrast loss, the sentiment prediction loss, the generation loss, and the diversity loss;

[0071] The parameters of the ININ model are adjusted based on the final loss value, and the training is terminated when the training end condition is met. When the training end condition is not met, the next iterative training is performed.

[0072] See also Figure 2a , Figure 2a The schematic diagram of the ININ model framework is shown in Figure 1. The ININ model consists of four modules: context encoding module, in-dialogue state iterative update learning module, response generation module, and inter-dialogue emotion adjacency comparison learning module. During the training process, the target training dialogue context including its corresponding reasoning knowledge is regarded as internal information, and other training dialogue contexts are regarded as external information.

[0073] See also Figure 2b, the context encoding module learns the context and common sense knowledge in the conversation through multiple encoders. Figure 2b As shown in the upper part of , optionally, in some embodiments, step 101 specifically includes:

[0074] Connecting all utterances in the historical conversation context to obtain an input sequence, wherein the input sequence includes a special mark for identifying the start of context input in the historical conversation context;

[0075] Processing the input sequence using a word embedding layer and a position embedding layer to obtain a word embedding and a position embedding of the input sequence;

[0076] Generate a final vector representation based on the word embedding, position embedding and role embedding, wherein the role embedding is used to distinguish the dialogue subject in the dialogue context, and the dialogue subject is a user or an intelligent agent;

[0077] The final vector representation is input into a Transformer encoder for processing to obtain the context representation.

[0078] First, each sentence in the historical dialogue context is connected into a long word sequence to obtain the input sequence C. Then the preset token [CLS] is used as the starting token of the context input, that is, C = [CLS] ⊕ U 1 ⊕U m ⊕…U M-1 ⊕U M , where the symbol “⊕” represents the connection operation, U m is the mth utterance in the historical dialogue context. Similarly, [CLS] also appears in the final hidden representation of the entire sequence.

[0079] The word embedding layer and position embedding layer are used to obtain the word embedding E of the input sequence C respectively. w (C) and position embedding E p (C). Since it is necessary to distinguish between users and intelligent agents in the iterative state update learning module within the dialogue, the role embedding E is inserted into the input sequence C r (C), and get the final vector representation. The final vector representation E(C) of the historical dialogue context is the sum of the above-mentioned types of embeddings:

[0080] E(C)=E w (C)+E p (C)+E r (C);

[0081] in, l≤MN M +1 is the number of words in the input sequence C, M is the maximum number of words in the sentence, N M is the number of sentences in the dialogue history, demb is the embedding dimension.

[0082] By feeding E(C) into the Transformer encoder for processing, we can obtain the context representation:

[0083] H ctx =Enc(E(C));

[0084] in, d h is the hidden layer size of the encoder.

[0085] It is difficult for the dialogue sequence representation to reflect the dialogue turns in chronological order. However, the node-line structure of the knowledge graph can solve this problem well. In order to enable the emotional and cognitive states of the user and the intelligent agent to be updated and iterated with the dialogue turns, a dialogue graph is constructed in this embodiment. It treats each dialogue turn as a node, which contains the common sense knowledge node and emotion node of each utterance, the cognitive state node and emotion state node of the user, and the cognitive state node and emotion state node of the agent.

[0086] Optionally, in some embodiments, the using the in-dialogue state iterative updating learning module to perform in-dialogue state iterative updating on the emotional state and cognitive state corresponding to each utterance in the historical dialogue context to obtain the final emotional state of the user and the final cognitive state of the intelligent agent includes:

[0087] Extracting common sense knowledge from each utterance of the historical conversation context and obtaining a hidden vector of the common sense knowledge, wherein the common sense knowledge includes an emotional state and a cognitive state;

[0088] Using the hidden vector of the common sense knowledge, generating an initial dialogue graph through an image constructor;

[0089] Iteratively updating the cognitive state and the emotional state based on the initial dialogue graph to obtain the cognitive state and the emotional state corresponding to each of the utterances;

[0090] The final emotional state of the user and the final cognitive state of the intelligent agent are determined based on the dialogue subject in the historical dialogue context.

[0091] See also Figure 2c Specifically, we first divide the common sense knowledge of each sentence into two parts: cognitive state and emotional state, as a dialogue graph. The initial state of the sentence is extracted from the dialogue context C. Five common sense relations ([xReact], [xWant], [xNeed], [xIntent], [xEffect]) are extracted from each sentence Um in the dialogue context C, where "[xReact]" represents the emotional state and the other four relations represent the cognitive state. Among them, the common sense knowledge is inferred from the dialogue context by the model COMET and the external knowledge base ATOMIC.

[0092] Then, two independent encoders are used to obtain the hidden vector of common sense knowledge:

[0093]

[0094] in, is the cognitive state of the historical conversation context, is the emotional state of the historical conversation context, d ck is the number of words in common sense knowledge, C rel is the representation of the common sense relation text corresponding to the dialogue, C react It is a representation of the emotional state of the conversation, Enc rel is the common sense relation encoder, Enc emo is the emotional state encoder, refers to the hidden representation of the preset tag [CLS], rel∈{xWant,xNeed,xIntent,xEffect}, ∥ is the connection operation, using The initial dialogue graph can be generated through the graph constructor

[0095] In human-computer dialogue, the words of the dialogue subjects will affect each other. This influence may be a change in emotional state or an extension of cognitive state. This is the same as human dialogue. Both parties are constantly enriching common sense knowledge in the dialogue and influencing each other's emotions. Therefore, in the process of dialogue, it is necessary to continuously learn the status of users and intelligent agents.

[0096] First, determine the initial state of the user and the intelligent agent:

[0097]

[0098] in, and is randomly initialized. represents the cognitive state of the first utterance in the historical conversation context, Indicates the emotional state of the first sentence in the historical dialogue context, Enc ini It is the initial encoder of the user's cognitive and emotional state.

[0099] The iterative update of cognitive state satisfies:

[0100]

[0101] in, For the current discourse U m The cognitive state, m∈(1,M], M is the number of utterances in the historical dialogue context, ⊙ is used to represent element-wise multiplication, σ is used to represent the sigmoid activation function, is the trainable weight matrix of the linear layer, is a trainable bias term.

[0102] Similarly, the iterative update of the emotional state satisfies:

[0103]

[0104] in, For the current discourse U m emotional state, b emo is a trainable bias term, W emo is the trainable weight matrix of the linear layer.

[0105] Then, based on the role embedding E r (C) It can be determined whether the iteratively updated state belongs to the user or the intelligent agent, and the updated state is further classified to obtain the user's cognitive state, the user's emotional state, the intelligent agent's cognitive state, and the intelligent agent's emotional state as follows:

[0106]

[0107] After iterative updates, the user’s final emotional state can be obtained and the final cognitive state of the intelligent agent and will and Input to the response generation module.

[0108] Optionally, in some embodiments, inputting the context representation, the user's final emotional state, and the intelligent agent's final cognitive state into a response generation module to obtain a target response includes:

[0109] fusing the final emotional state of the user and the final cognitive state of the intelligent agent into the context representation to obtain a final context representation;

[0110] The final context representation is input into the decoder for processing to obtain a target response.

[0111] See also Figure 2d In the response generation module, the user's final emotional state is first Fusion to the context representation H ctx , and obtain the emotional state enhanced context representation

[0112]

[0113] in, Enhance contextual representation for emotional states, for Feature representation after dimension transformation, dim = l × d h ,Enc ctx-emo and Enc ctx-cog They are emotion enhancement encoder and cognitive enhancement encoder respectively, and expand(dim) is the dimension conversion function.

[0114] Similarly, the agent's final cognitive state Fusion to the context representation H ctx , and obtain the cognitive state enhanced context representation

[0115]

[0116] Finally, based on and Together, they are used to generate the final context representation H′ ctx :

[0117]

[0118] in MLP is a multilayer perceptron with ReLU activation, the symbol σ is the sigmoid activation function, and the symbol ⊙ represents element-wise multiplication.

[0119] The final context representation is input into the decoder to obtain the target response. Specifically, the target response U M+1 =Y = (y 1 ,…,y T ) is generated by the decoder token by token:

[0120] P(y t ∣y 1:t-1 ,C)=Dec(Y 1:t-1 ,H′ ctx );

[0121] Among them, Y 1:t-1 represents the embedding of the token generated before time t and Dec() is the decoder.

[0122] In this embodiment, the final loss value is determined based on contrast loss, sentiment prediction loss, generation loss and diversity loss. The calculation methods of each loss are introduced below.

[0123] Users have similar emotional needs in different conversation situations. By learning and comparing conversation situation features of the same emotion category at different scales, the intelligent agent can better adapt to changes in conversation situations. Therefore, an inter-conversation emotion neighbor comparison learning module is set up to extract similar features between conversations of the same emotion category.

[0124] See also Figure 2b Optionally, in some embodiments, the determining the contrast loss in the current iterative training process based on the multiple training dialogue contexts and using the inter-dialogue emotion adjacency contrast learning module includes:

[0125] Inputting the multiple training dialogue contexts into K parallel enhanced encoders to obtain K enhanced views, where K is an integer greater than 1;

[0126] Generate an emotion adjacency matrix corresponding to the K enhanced views, where the emotion adjacency matrix is ​​used to indicate whether the emotion labels of any two training dialogue contexts in the multiple training dialogue contexts are the same;

[0127] Any one of the K enhanced views is used as a target view, and the contrast loss is calculated based on a neighbor contrast loss between each other enhanced view and the target view.

[0128] In order to ensure sufficient learning samples, multiple training dialogue contexts are input into K parallel augmented encoders to obtain different dialogue scenarios, which are specifically represented as follows:

[0129]

[0130] in, is the conversation context representation generated by the k-th augmented view, k∈[1,K]. Originated from Transformer, the main difference between the two lies in the number of heads in the multi-head attention mechanism. This design aims to obtain dialogue representations in different spatial dimensions, thereby appropriately broadening the context learning space.

[0131] See also Figure 2e , then generate the emotion adjacency matrix corresponding to the K enhanced views. The specific method of generating the emotion adjacency matrix is ​​as follows:

[0132] Given a batch size B = {C 1 , C 2 , ..., C N}, represents the emotion label vector, t i Represents the i-th training dialogue context C iThe label, sentiment adjacency matrix M e The construction formula is as follows:

[0133]

[0134] Among them, M e ∈R N×N , If and only if dialogue i and dialogue j have the same emotion label, otherwise

[0135] In the specific implementation, the number of enhanced views K is arbitrarily selected. When K is greater than 2, one of the enhanced views is randomly selected. Determine the target view, use the target view as the anchor view, and compare each enhanced view with the target view. The contrast loss is calculated by the neighbor contrast loss between them. The details are as follows:

[0136]

[0137] To enhance the view and enhanced views For example, let and Enhanced Vision Figure 1 And the training dialogue context C learned from enhanced view 2 i Embed, select As the current training dialogue context, positive samples can be divided into three categories:

[0138] (1) The same training dialogue context in different augmented views, i.e.

[0139] (2) Enhanced Vision Figure 1 The training dialogue context with the same sentiment label in Right now

[0140] (3) Enhance the training dialogue context with the same emotion label in view 2 Right now

[0141] The other training dialogue contexts are negative samples. The number of relevant positive pairs should be in It is c j With enhanced vision Figure 1 middle Related Enhanced Vision Figure 1 The neighbor contrast loss formula between the enhanced view 2 is calculated as follows:

[0142]

[0143] Among them, θ(·) is the inner product function to measure the similarity, τ is the temperature coefficient, are the positive samples of the two enhanced views. Similarly, we can get Related Enhanced Vision Figure 1 and the neighbor contrast loss between enhanced view 2

[0144] Will enhance visual Figure 1 The total neighbor contrast loss between and enhanced view 2 is averaged over all training dialogue contexts and is defined as:

[0145]

[0146] This embodiment also includes emotion prediction loss, generation loss and diversity loss. Specifically, in order to predict emotions more accurately, use The [CLS] hidden representation is used for emotion classification, which contains the user’s last emotional state in the context of the target training dialogue:

[0147]

[0148] in,

[0149] Then, the Softmax operation is used to feed e into the linear layer to obtain the emotion category distribution P emo :

[0150] P emo =Softmax(W e e)

[0151] Among them, P emo ∈R s , s is the total number of emotion categories available in the dataset. During training, the emotion category distribution P is calculated emo The weight parameters are updated by minimizing the cross entropy loss between the real label e′ and the emotion detection loss:

[0152]

[0153] After getting the training response, use the negative log-likelihood as the generative loss function:

[0154]

[0155] The diversity loss is as follows:

[0156]

[0157] Where T is the total number of decoding time steps, V is the number of vocabulary in the dataset, and a i is a candidate token for the decoded word, δ t (a i ) is the indicator function.

[0158] The final loss value is as follows:

[0159]

[0160] In this embodiment, an ININ model is proposed for generating empathic responses, which has both discourse-level and conversation-level understanding capabilities. Specifically, at the discourse level, an intra-conversation state iterative update strategy is created, which regards each utterance as a continuous time node and updates the states of the user and agent accordingly to simulate their interactions. At the conversation level, an emotional adjacency matrix is ​​constructed based on emotional labels, and neighbor contrast learning is applied to capture the overall differences between conversations in the same emotional category. In addition, this embodiment also combines representations from the discourse and conversation levels and integrates them into the empathic response generation process, thereby generating responses with better empathic effects.

[0161] See also Figure 3 The embodiment of the present invention further provides a response generating device 300 for generating an empathy response, comprising:

[0162] A context encoding module 301 is used to encode the historical conversation context between the user and the intelligent agent to obtain a context representation;

[0163] The in-dialogue state iterative update learning module 302 is used to iteratively update the in-dialogue state of the emotional state and cognitive state corresponding to each utterance in the historical dialogue context to obtain the final emotional state of the user and the final cognitive state of the intelligent agent;

[0164] The response generation module 303 is used to obtain a target response based on the context representation, the user's final emotional state and the intelligent agent's final cognitive state.

[0165] Optionally, the response generating device 300 for generating an empathy response further includes:

[0166] A training module, configured to iteratively train the ININ model based on a plurality of training dialogue contexts to obtain a training response corresponding to a target training dialogue context, wherein the plurality of training dialogue contexts include the target training dialogue context, and the ININ model includes the context encoding module, the in-dialogue state iterative update learning module, the response generation module, and the inter-dialogue emotion adjacency comparison learning module;

[0167] A determination module, configured to determine the contrast loss in the current iterative training process based on the multiple training dialogue contexts and using the inter-dialogue emotion adjacency contrast learning module, and calculate the emotion prediction loss, generation loss and diversity loss in the current iterative training process based on the training responses;

[0168] Determining a final loss value based on the contrast loss, the sentiment prediction loss, the generation loss, and the diversity loss;

[0169] The parameters of the ININ model are adjusted based on the final loss value, and the training is terminated when the training end condition is met. When the training end condition is not met, the next iterative training is performed.

[0170] Optionally, the determining the contrast loss in the current iterative training process based on the multiple training dialogue contexts and using the inter-dialogue emotion adjacency contrast learning module includes:

[0171] Inputting the multiple training dialogue contexts into K parallel enhanced encoders to obtain K enhanced views, where K is an integer greater than 1;

[0172] Generate an emotion adjacency matrix corresponding to the K enhanced views, where the emotion adjacency matrix is ​​used to indicate whether the emotion labels of any two training dialogue contexts in the multiple training dialogue contexts are the same;

[0173] Any one of the K enhanced views is used as a target view, and the contrast loss is calculated based on a neighbor contrast loss between each other enhanced view and the target view.

[0174] Optionally, the contrast loss is:

[0175]

[0176] in, is the x-th enhanced view, the x-th enhanced view is the target view, is the kth enhanced view, The embedding of the i-th training dialogue context learned for the k-th augmented view, The embedding of the jth training conversation context learned for the kth augmentation view, The embedding of the i-th training dialogue context learned for the x-th augmented view, is the embedding of the jth training dialogue context learned for the xth augmented view, N is the number of training sample data, θ(·) is the inner product function measuring similarity, τ is the temperature coefficient, are the positive samples in the two enhanced views, is the dialogue C in the training batch jCollection of Indicates dialogue C j and dialogue C i have the same emotion category.

[0177] Optionally, the context encoding module 301 includes:

[0178] a connection unit, used to connect all the utterances in the historical conversation context to obtain an input sequence, wherein the input sequence includes a special mark for identifying the start of the context input in the historical conversation context;

[0179] A processing unit, configured to process the input sequence using a word embedding layer and a position embedding layer to obtain a word embedding and a position embedding of the input sequence;

[0180] A generating unit, configured to generate a final vector representation based on the word embedding, the position embedding and the role embedding, wherein the role embedding is used to distinguish a dialogue subject in the dialogue context, wherein the dialogue subject is a user or an intelligent agent;

[0181] The encoding unit is used to input the final vector representation into a Transformer encoder for processing to obtain the context representation.

[0182] Optionally, the in-dialog state iterative update learning module 302 includes:

[0183] An extraction unit, configured to extract common sense knowledge from each utterance of the historical conversation context and obtain a hidden vector of the common sense knowledge, wherein the common sense knowledge includes an emotional state and a cognitive state;

[0184] A construction unit, configured to generate an initial dialogue graph through an image constructor using the hidden vector of the common sense knowledge;

[0185] An updating unit, configured to iteratively update the cognitive state and the emotional state based on the initial dialogue graph to obtain the cognitive state and the emotional state corresponding to each of the utterances;

[0186] A determination unit is used to determine the final emotional state of the user and the final cognitive state of the intelligent agent based on the dialogue subject in the historical dialogue context.

[0187] Optionally, the updating unit is specifically used for:

[0188] Determine the user’s initial cognitive state The user's initial emotional state Initial cognitive state of the intelligent agent and the initial emotional state of the agent

[0189]

[0190] in, and is randomly initialized. represents the cognitive state of the first utterance in the historical conversation context, Indicates the emotional state of the first sentence in the historical dialogue context, Enc ini Initial encoders for representing the user’s cognitive and emotional states;

[0191] Based on the user's initial cognitive state The user's initial emotional state The initial cognitive state of the intelligent agent and the initial emotional state of the intelligent agent Perform iterative updates to obtain the emotional state and cognitive state corresponding to each sentence in the historical conversation context. The iterative update process of the cognitive state is as follows:

[0192]

[0193] in, For the current discourse U m The cognitive state, m∈(1,M], M is the number of utterances in the historical dialogue context, ⊙ is used to represent element-wise multiplication, σ is used to represent the sigmoid activation function, is the trainable weight matrix of the linear layer, is a trainable bias term;

[0194] The iterative update process of the emotional state is:

[0195]

[0196] in, For the current discourse U m emotional state, b emo is a trainable bias term, W emo is the trainable weight matrix of the linear layer.

[0197] The response generation device 300 for generating an empathic response provided in the embodiment of the present application can execute the above-mentioned method embodiment, and its implementation principle and technical effect are similar, which will not be repeated in this embodiment.

[0198] It should be noted that the division of units in the embodiments of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional unit in each embodiment of the present application may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0199] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a processor-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program codes.

[0200] like Figure 4 As shown, an embodiment of the present application provides an electronic device 400, including: a memory 402, a processor 401, and a program stored in the memory 402 and executable on the processor 401; the processor 401 is used to read the program in the memory 402 to implement the steps in the response generation method for generating an empathetic response as described above.

[0201] An embodiment of the present application also provides a readable storage medium, on which a program is stored. When the program is executed by a processor, the various processes of the above-mentioned response generation method embodiment for generating empathetic responses are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here. Among them, the readable storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic storage (such as floppy disk, hard disk, magnetic tape, magneto-optical disk (MO), etc.), optical storage (such as compact disk (CD), digital video disc (DVD), Blu-ray Disc (BD), high-definition versatile disc (HVD), etc.), and semiconductor memory (such as read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read only memory (EEPROM), non-volatile memory (NAND FLASH), solid-state drive (Solid State Disk or Solid State Drive, SSD)), etc.

[0202] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0203] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, disk, CD), and includes a number of instructions for a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0204] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A response generation method for generating an empathic response, characterized in that: include: The historical conversation context between the user and the intelligent agent is input into the context encoding module for encoding processing to obtain the context representation; Using the in-dialogue state iterative update learning module to iteratively update the emotional state and cognitive state corresponding to each utterance in the historical dialogue context, to obtain the final emotional state of the user and the final cognitive state of the intelligent agent; The context representation, the user's final emotional state and the intelligent agent's final cognitive state are input into a response generation module to obtain a target response.

2. The method according to claim 1, characterized in that: Before inputting the historical conversation context between the user and the intelligent agent into the context encoding module for encoding processing to obtain the context representation, the method further includes: Iteratively training the ININ model based on multiple training dialogue contexts to obtain a training response corresponding to a target training dialogue context, wherein the multiple training dialogue contexts include the target training dialogue context, and the ININ model includes the context encoding module, the in-dialogue state iterative update learning module, the response generation module, and the inter-dialogue emotion adjacency comparison learning module; Based on the multiple training dialogue contexts, using the inter-dialogue emotion adjacency contrast learning module to determine the contrast loss in the current iterative training process, and calculating the emotion prediction loss, generation loss and diversity loss in the current iterative training process based on the training responses; Determining a final loss value based on the contrast loss, the sentiment prediction loss, the generation loss, and the diversity loss; The parameters of the ININ model are adjusted based on the final loss value, and the training is terminated when the training end condition is met. When the training end condition is not met, the next iterative training is performed.

3. The method according to claim 2, characterized in that: The determining the contrast loss in the current iterative training process based on the multiple training dialogue contexts and using the inter-dialogue emotion adjacency contrast learning module includes: Inputting the multiple training dialogue contexts into K parallel enhanced encoders to obtain K enhanced views, where K is an integer greater than 1; Generate an emotion adjacency matrix corresponding to the K enhanced views, where the emotion adjacency matrix is ​​used to indicate whether the emotion labels of any two training dialogue contexts in the multiple training dialogue contexts are the same; Any one of the K enhanced views is used as a target view, and the contrast loss is calculated based on a neighbor contrast loss between each other enhanced view and the target view.

4. The method according to claim 3, characterized in that: The contrast loss is: in, is the xth enhanced view, the xth enhanced view is the target view, is the kth enhanced view, The embedding of the i-th training dialogue context learned for the k-th augmented view, The embedding of the jth training conversation context learned for the kth augmented view, The embedding of the i-th training dialogue context learned for the x-th augmented view, is the embedding of the jth training dialogue context learned for the xth augmented view, N is the number of training sample data, θ(·) is the inner product function measuring similarity, τ is the temperature coefficient, are the positive samples in the two enhanced views, is the dialogue C in the training batch jj Collection of Indicates dialogue C jj and dialogue C ii have the same emotion category.

5. The method according to claim 1, characterized in that: The inputting the historical conversation context between the user and the intelligent agent into the context encoding module for encoding processing to obtain the context representation includes: Connecting all utterances in the historical conversation context to obtain an input sequence, wherein the input sequence includes a special mark for identifying the start of context input in the historical conversation context; Processing the input sequence using a word embedding layer and a position embedding layer to obtain a word embedding and a position embedding of the input sequence; Generate a final vector representation based on the word embedding, position embedding and role embedding, wherein the role embedding is used to distinguish the dialogue subject in the dialogue context, and the dialogue subject is a user or an intelligent agent; The final vector representation is input into a Transformer encoder for processing to obtain the context representation.

6. The method according to claim 1, characterized in that: The in-dialogue state iterative updating learning module performs in-dialogue state iterative updating on the emotional state and cognitive state corresponding to each utterance in the historical dialogue context to obtain the final emotional state of the user and the final cognitive state of the intelligent agent, including: Extracting common sense knowledge from each utterance of the historical conversation context and obtaining a hidden vector of the common sense knowledge, wherein the common sense knowledge includes an emotional state and a cognitive state; Using the hidden vector of the common sense knowledge, generating an initial dialogue graph through an image constructor; Iteratively updating the cognitive state and the emotional state based on the initial dialogue graph to obtain the cognitive state and the emotional state corresponding to each of the utterances; The final emotional state of the user and the final cognitive state of the intelligent agent are determined based on the dialogue subject in the historical dialogue context.

7. The method according to claim 6, characterized in that: The iterative updating of the cognitive state and the emotional state based on the initial dialogue graph to obtain the cognitive state and the emotional state corresponding to each utterance includes: Determine the user’s initial cognitive state The user's initial emotional state Initial cognitive state of the intelligent agent and the initial emotional state of the agent in, and is randomly initialized. represents the cognitive state of the first utterance in the historical conversation context, Indicates the emotional state of the first sentence in the historical dialogue context, Enc iieeii Initial encoders for representing the user’s cognitive and emotional states; Based on the user's initial cognitive state The user's initial emotional state The initial cognitive state of the intelligent agent and the initial emotional state of the intelligent agent Perform iterative updates to obtain the emotional state and cognitive state corresponding to each sentence in the historical conversation context. The iterative update process of the cognitive state is as follows: in, For the current discourse U ee The cognitive state, m∈(1,M], M is the number of utterances in the historical dialogue context, ⊙ is used to represent element-wise multiplication, σ is used to represent the sigmoid activation function, is the trainable weight matrix of the linear layer, is a trainable bias term; The iterative update process of the emotional state is: in, For the current discourse U ee emotional state, b eeeecc is a trainable bias term, W eeeecc is the trainable weight matrix of the linear layer.

8. A response generating device for generating an empathic response, characterized in that: include: The context encoding module is used to encode the historical conversation context between the user and the intelligent agent to obtain the context representation; An intra-dialogue state iterative update learning module is used to iteratively update the intra-dialogue state of the emotional state and cognitive state corresponding to each utterance in the historical dialogue context to obtain the final emotional state of the user and the final cognitive state of the intelligent agent; A response generation module is used to obtain a target response based on the context representation, the user's final emotional state and the intelligent agent's final cognitive state.

9. An electronic device, comprising: A memory, a processor, and a program stored in the memory and executable on the processor; wherein the processor is used to read the program in the memory to implement the steps in the response generation method for generating an empathetic response as described in any one of claims 1 to 7.

10. A readable storage medium for storing a program, characterized in that: When the program is executed by a processor, the steps of the response generation method for generating an empathic response as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Common-condition reply generation method and device, terminal and storage medium

    CN115934909A

  • Method and apparatus for determining a dialog state, dialog system, computer device, and storage medium

    US20200335104A1

Cited By

  • Emotional dialogue generation method and system based on large model

    CN120851041A