Estrus-sharing dialogue generation method and system fusing psychological information

By constructing a role-aware emotion embedding representation and a multimodal attention mechanism, dynamically fusing the emotional states of the speaker and the listener, and combining an external knowledge base to generate multi-hop knowledge reasoning paths, the problem of insufficient emotion understanding in existing dialogue generation technologies is solved, and more natural and accurate empathetic dialogue generation is achieved.

CN121786754APending Publication Date: 2026-04-03ZHEJIANG NORMAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-15
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing dialogue generation technologies lack a deep understanding and empathy for human emotions, making it difficult to achieve natural and warm human-computer dialogue interaction. Furthermore, the introduction of common sense knowledge cannot be effectively filtered and utilized, resulting in a lack of emotional adaptability in the generated dialogues.

Method used

By constructing an emotional embedding representation of role perception and splicing it with the role context representation, the emotional states of the speaker and the listener are dynamically fused. Combined with a multimodal attention mechanism and an external knowledge base, multi-hop knowledge reasoning paths are generated, enhancing the emotional expression and semantic understanding of dialogue generation.

Benefits of technology

It improves the naturalness and accuracy of dialogue generation, enhances the system's ability to understand users' emotional states, and generates responses that are more empathetic and emotionally resonant.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786754A_ABST
    Figure CN121786754A_ABST
Patent Text Reader

Abstract

The invention discloses an emotional dialogue generation method and system fused with psychological information, and the emotional expression capability of generating response can be improved through the obtained global emotional latent variable fused with a speaker and a listener. By determining a multi-hop knowledge reasoning path set, obtaining context semantic representation of each knowledge reasoning path, and fusing the context semantic representation with dialogue history global context representation to obtain knowledge-enhanced context representation, the knowledge reasoning ability and semantic relevance in a dialogue context are improved; by constructing a knowledge representation matrix, taking a global emotion latent variable as a bias item, and dynamically adjusting the knowledge representation matrix and a value vector generated by knowledge-enhanced context representation through semantic features, multi-scale attention fusion representation is obtained, and the context adaptive capacity and emotion representation naturalness of generated response are effectively improved. According to the method, the psychological state of the user can be identified more accurately, and the dialogue response with humanization, estrus-sharing expression and interaction adaptability is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing technology, and more specifically to a method and system for generating empathic dialogues that integrates psychological information. Background Technology

[0002] With the continuous development of artificial intelligence technology, dialogue generation technology has become an important research direction in the field of natural language processing, demonstrating wide application value in various scenarios such as intelligent customer service, medical consultation, psychological counseling, and social companionship. Although current dialogue generation technology exhibits strong language understanding and information retrieval capabilities in human-computer interaction, significant gaps still exist in human-computer communication. This is because existing technologies primarily focus on surface-level semantic matching of language, lacking a deep understanding and empathy for human emotions, making it difficult to achieve natural and warm human-computer dialogue interaction.

[0003] Therefore, enhancing the emotional perception capabilities of dialogue systems has become a key research objective in order to build more intelligent and human-like systems. Empathy, as an important component of interpersonal communication, is becoming a core direction for the further development of dialogue generation technology. Empathy, as a key social cognitive ability, refers to an individual's ability to perceive and understand the emotions of others and respond appropriately. Empathy not only enhances emotional connections between individuals but also plays a vital role in mental health and social interaction. Dialogue generation systems incorporating empathic elements need to be able to identify user emotions, understand the reasons for emotional changes, and provide appropriate responses in the dialogue to enhance the naturalness and humanization of the interaction. The goal of empathic dialogue generation that integrates psychological information is to capture the user's emotional state during the dialogue process and generate responses that match the user's emotions and the context of the communication, making human-computer dialogue more interactive and realistic.

[0004] Research on empathic dialogue generation originated from the exploration of emotion perception in dialogue systems. The Emotional Chatbot ECM first proposed a dialogue system incorporating empathy; this model can express specific emotional information when generating responses based on predefined emotion tags. Subsequently, Rashkin et al. constructed the large-scale empathic dialogue dataset *EmpatheticDialogues* and, by adding emotion feature vectors to the model, enabled the system to understand and simulate the user's emotional state to a certain extent. These studies laid the foundation for empathic dialogue generation that integrates psychological information, leading to some progress in the emotional expression of dialogue systems. In terms of emotion modeling, existing research mainly focuses on modeling and utilizing emotional information. For example, the MoEL model sets different decoders for different contextual emotions and generates more empathetic responses by combining a shared decoder with predicted emotion distribution. While these methods improve the quality of empathic dialogue generation to some extent, they mostly focus only on the speaker's single emotional state, neglecting the emotional interaction between the listener and speaker during the dialogue, resulting in a lack of coherence in empathic expression in multi-turn dialogues.

[0005] However, considering only emotional factors is insufficient to generate high-quality empathetic dialogues; common-sense knowledge must also be incorporated. Common-sense knowledge plays a crucial role in human communication, enabling speakers to reason and express themselves based on real-world knowledge. Empathetic dialogue systems lacking common-sense knowledge may struggle to accurately understand the emotional intentions expressed by users and fail to provide responses appropriate to the dialogue context. Traditional approaches to incorporating common-sense knowledge include the PostKS model, which proposes a knowledge selection framework based on prior-posterior distribution co-optimization to improve knowledge adaptability and make dialogue generation more coherent; the KEMP model, which constructs a multimodal sentiment context graph integrating the ConceptNet common-sense knowledge graph and the NRC-VAD sentiment lexicon, and extracts cross-modal sentiment features through a graph attention network to establish a knowledge-enhanced sentiment reasoning mechanism; and the DCKS model, which proposes a context-aware common-sense reasoning mechanism that dynamically selects relevant common-sense knowledge and sentiment state classifications to ensure the coordination between common-sense logic and sentiment evolution, thereby generating more empathetic dialogue responses. These studies demonstrate that the appropriate incorporation of common-sense knowledge can effectively improve the quality of empathetic dialogue generation, but how to efficiently select and utilize external knowledge remains a significant challenge.

[0006] In summary, although existing research has made significant progress in generating empathic dialogues, limitations in emotion modeling remain. Most existing models rely on a single emotion label as a supervisory signal, neglecting the dynamic evolution of emotions and the complexity of two-way emotional interaction in dialogues. Although the introduction of common-sense knowledge helps improve the logic and coherence of dialogues, the inability of existing methods to effectively filter and utilize external knowledge results in a lack of emotional adaptability in the generated dialogues. Summary of the Invention

[0007] To address the problems existing in the aforementioned fields, this invention proposes a method and system for generating empathic dialogues that integrates psychological information. The method acquires global emotional latent variables, which can dynamically integrate the emotional changes of the speaker and listener; the acquired knowledge-enhanced contextual representation can effectively filter external knowledge and align with the context in the dialogue information. Through the fusion of psychological dialogue semantic features using a multimodal attention mechanism, the system's comprehensive capabilities in emotion recognition and knowledge reasoning are enhanced, significantly improving the depth and naturalness of dialogue generation in terms of emotional expression and semantic understanding, thereby generating more context-appropriate and empathetic responses.

[0008] To address the aforementioned technical problems, this invention discloses a method for generating empathic dialogues that integrates psychological information, comprising the following steps: Obtain the role context representation of the dialogue sequence, the global context representation of the dialogue history, and the target response representation; By constructing and concatenating the emotional embedding representation of role perception with the role context representation, an emotionally enhanced context representation is generated; based on the emotionally enhanced context representation and the target response representation, the global emotional latent variable that integrates the emotional states of the speaker and the listener is obtained by identifying the emotional states of the speaker and the listener and dynamically fusing them. Construct a keyword set for the current speaker's dialogue sequence, and determine a knowledge subset by retrieving and constructing a set of candidate knowledge tuples corresponding to the keywords; when the target entity in the knowledge subset matches the candidate keyword, an initial path connection is formed, and the path connection is gradually expanded to the maximum allowed reasoning depth to determine the set of multi-hop knowledge reasoning paths; obtain the contextual semantic representation of each knowledge reasoning path, and fuse it with the global contextual representation of the dialogue history to obtain a knowledge-enhanced contextual representation; Based on the knowledge-enhanced contextual representation, a keyword-guided knowledge representation matrix is ​​constructed; multi-scale semantic features of the query vector of the knowledge-enhanced contextual representation are extracted, and the global sentiment latent variable is used as a bias term. The semantic features are used to dynamically adjust the value vector generated by the knowledge representation matrix and the knowledge-enhanced contextual representation to obtain a multi-scale attention fusion representation. The multi-scale attention fusion representation is concatenated with global sentiment latent variables and then input into the pointer network. By predicting the probability distribution of the current word, an empathetic response is generated.

[0009] Preferably, the acquisition of the role context representation of the dialogue sequence, the global context representation of the dialogue history, and the target response representation is obtained by encoding the input dialogue sequence through a context encoding module, including the following steps: The context encoding module includes a GloVe model, a Transformer-based encoder, and a Transformer-based interactive encoder. Dialogue history Dialogue sequence divided according to roles Conversation sequence with listeners And introduce dialogue history The corresponding target response sequence of the reference response ; On the history of dialogue Target response sequence Speaker dialogue sequence Conversation sequence with listeners Each statement in the text is marked with a special tag [CLS] and then concatenated to obtain the expanded dialogue history. Target response sequence Speaker dialogue sequence Dialogue Sequence with Listeners ; The expanded sequences are mapped to a continuous vector space through the word embedding layer of the GloVe model, thereby obtaining the global contextual semantic representation of the dialogue history. Contextual semantic representation of the target response sequence Contextual semantic representation of speaker dialogue sequence Contextual semantic representation of dialogue sequences with listeners ; Will and Input a Transformer-based encoder to obtain a global context representation of the dialogue history. and target response representation ; Based on pre-trained language models, and Input to a Transformer-based interactive encoder, and As a query, As keys and values, respectively, we obtain the role context representations of the speaker's dialogue sequence and the listener's dialogue sequence. and .

[0010] Preferably, the contextual representation of the emotion enhancement is obtained through the emotion enhancement module, specifically including: The sentiment enhancement module includes a sentiment enhancement context representation generation module, a recognition network, and a normalized flow-based conditional variational autoencoder; Based on the context of the speaker's and listener's roles and The hidden state vectors corresponding to the [CLS] labels are extracted respectively, and obtained by linear mapping and the Softmax function. and The probability distribution of the sentiment category; according to and The probability distribution of sentiment categories is used, and the sentiment category label corresponding to the maximum value of the probability distribution is mapped to a low-dimensional vector representation through the sentiment embedding layer. Sentiment embedding representations for the speaker and the listener are constructed respectively. and ; Will and Representation of the corresponding role context and By splicing the data and introducing a self-attention mechanism, an emotionally enhanced contextual representation is obtained. and .

[0011] Preferably, the emotion enhancement module further includes a recognition network and a conditional variational autoencoder based on normalized flow. The recognition network identifies the emotional states of the speaker and listener to obtain initial latent emotional variables for both. The conditional variational autoencoder based on normalized flow dynamically fuses the initial latent emotional variables of the speaker and listener to obtain global latent emotional variables that fuse the emotional states of both the speaker and listener. Specifically, this includes: Initialize the latent emotional variables of the speaker and listener separately from a standard normal distribution. Contextual representation based on sentiment enhancement and Representation of target response The mean and standard deviation of the speaker's and listener's emotional latent variables are estimated using the recognition network. Based on the mean and standard deviation, a reparameterization technique is introduced to sample the emotional latent variables, obtaining initial emotional latent variables with emotional characteristics. and ; The conditional variational autoencoder based on normalized flow is a conditional variational autoencoder that sequentially introduces a normalized flow mechanism and a dynamic fusion mechanism. A normalized flow mechanism is introduced to analyze the initial emotional latent variables. and By performing multi-level reversible flow transformation, the emotional latent variables of the speaker and listener after transformation are obtained. and ; A dynamic fusion mechanism is introduced to dynamically adjust the emotional latent variables of the speaker and listener after transformation. and The weights in the global affective latent variable are used to obtain the global affective latent variable that integrates the speaker's and listener's emotions. .

[0012] Preferably, the keyword set of the current speaker's dialogue sequence is constructed through a semantic deduplication module, specifically including: The semantic deduplication module includes a pre-trained language model, the TF-IDF method, and a scaled dot product attention mechanism. Extract each candidate word The word vector representation is used to calculate the semantic importance score of candidate words through the multilayer perceptron in the pre-trained language model, thus obtaining the candidate word score calculated based on the pre-trained model; and the statistical weight score of candidate words is calculated based on the TF-IDF method, thus obtaining the candidate word score calculated based on the TF-IDF method. The top K words with scores calculated using both the pre-trained model and the TF-IDF method are selected to form a preliminary keyword set and a TF-IDF candidate set, respectively. Lexical normalization and semantic vector fusion are then performed on the preliminary keyword set and the TF-IDF candidate set, respectively, to obtain a fused keyword set. Finally, a deduplication operation is performed on the fused keyword set to remove redundant terms after semantic merging, resulting in a deduplicated keyword set. ; Based on the word vector representations of the acquired candidate words and the global context representation of the dialogue history, the matching degree between each candidate word and the context is calculated through a scaling dot product attention mechanism to obtain the attention weight between the candidate word and the current dialogue context. The attention weights between candidate words and the current dialogue context are fused with candidate word scores calculated based on the pre-trained model and the TF-IDF method to construct a context-aware keyword importance scoring function. The context-aware keyword importance scoring function, the candidate word scores calculated based on the pre-trained model and the TF-IDF method are then weighted and fused to obtain a comprehensive score. Candidate words are ranked according to the comprehensive score, and the top K words with the highest scores are selected to form the final keyword set. That is, the set of keywords in the current speaker's dialogue sequence. .

[0013] Preferably, determining the multi-hop knowledge reasoning path set through the knowledge reasoning module specifically includes: The knowledge reasoning module includes a semantic deduplication module and a knowledge retrieval and filtering module; The knowledge retrieval and filtering module includes a knowledge retrieval module, a context fusion encoder, and a knowledge reasoning path generation module. Keyword set based on the current speaker's dialogue sequence The system retrieves knowledge triples related to the current dialogue context from the external knowledge base of the knowledge retrieval module, obtains their corresponding confidence scores, and constructs a set of candidate knowledge quadruples with confidence scores. By determining a comprehensive scoring function between the target entity and its corresponding context feature vector, the candidate knowledge quadruples are filtered and reordered, and the top K quadruples with the highest scores are selected to form a knowledge subset. ; Extract keywords using a context fusion encoder Semantic feature representation in the current context The knowledge reasoning path generation module is used to calculate knowledge subsets based on semantic feature representations. target entity in With keywords Semantic similarity; determining the target entity in a quadruple. Whether there is a semantic connection with other keywords; when the semantic similarity exceeds a set threshold, i.e., the target concept... With a certain keyword If a match is found, a semantic jump path exists, forming a path connection; recursive reasoning is used to expand the path to the set maximum allowed reasoning depth. Generate a set of multi-hop paths The paths are then sorted, and the top K paths are selected as the optimal set of multi-hop knowledge reasoning paths.

[0014] Preferably, obtaining the knowledge-enhanced contextual representation through the knowledge enhancement module specifically includes: The knowledge enhancement module includes a knowledge reasoning path semantic generation module and a knowledge fusion module; The context semantic representation generation module for the knowledge reasoning path set is used to convert the knowledge reasoning paths in the knowledge reasoning path set into natural language sequences, and input them into a bidirectional gated loop unit to obtain the context semantic representation of the knowledge reasoning path set. ; Obtain the global context representation of the dialogue history The first in The representation of a token ; The scaling dot product attention mechanism is used to compute the first... Knowledge fusion weights corresponding to each position ; The knowledge fusion module is used to integrate... and Contextual representation of knowledge reasoning paths through element-wise weighting Representation of the context of the dialogue Dynamic fusion is performed at various locations to obtain a knowledge-enhanced contextual representation. .

[0015] Preferably, the keyword-guided knowledge representation matrix is ​​constructed through the speech intent prediction module, specifically including: From knowledge-enhanced contextual representation Extract the hidden state corresponding to the [CLS] marker. Based on the hidden state The probability distribution of intent categories in the current dialogue is generated through linear transformation and Softmax prediction. The predicted intent category is mapped to a high-dimensional vector representation through the embedding layer in the dialogue intent prediction module, generating a dialogue intent embedding vector. ; Embedding dialogue intent into vectors The starting tag vector of the decoder generated by the empathic dialogue The vectors are concatenated to form a fused vector, which is then mapped to the dimensions required by the decoder to obtain the initial decoding state. , as the starting input of the decoder; Based on the global context representation of the keyword set and dialogue history, obtain the keyword attention weight vector g; Enhancing the context of knowledge As the key and value inputs in the cross-attention mechanism of the decoder generated by the empathic dialogue, and combined with the keyword attention weight vector g, based on the initial decoding state A knowledge representation matrix is ​​constructed through element-wise weighted operations. .

[0016] Preferably, the multi-scale attention fusion representation is obtained by decoding through a context decoding module, including the following steps: The context decoding module includes a mask self-attention mechanism, a cross attention mechanism, a multi-scale multi-head attention mechanism, and a feedforward neural network; Embed the generated dialogue intent vector As input to the decoder, the dependencies of the dialogue history rounds are modeled through a masked self-attention mechanism to obtain the intent representation of the current dialogue; Based on the intent representation of the current dialogue and the global context representation of the dialogue history, the external context is fused through the key and value of the cross-attention mechanism to obtain a global context representation that incorporates external information; By employing a multi-scale, multi-head attention mechanism, the global context representation query vector, generated internally by the decoder and incorporating external information, is input into multiple parallel one-dimensional convolutional modules to extract knowledge and enhance the context representation. Multi-scale semantic features of query vectors; By integrating multi-scale semantic features through a feedforward neural network, a unified representation of multi-scale semantic features is generated. Introducing latent emotional variables As a bias term, the value vector generated by the knowledge representation matrix and the knowledge-enhanced context representation is dynamically adjusted through multi-scale semantic features of a unified representation. Attention scores are calculated using Softmax, and the query vector at each scale is adjusted accordingly. (i=1, 2, 3) Perform multi-head attention fusion respectively to obtain multi-scale attention fusion output. .

[0017] Preferably, it also includes an empathic dialogue generation system that integrates psychological information, comprising: The context encoding module is used to obtain the role context representation of the dialogue sequence, the global context representation of the dialogue history, and the target response representation; The emotion enhancement module includes an emotion enhancement context representation generation module, a recognition network, and a normalized flow-based conditional variational autoencoder. The emotion enhancement context representation generation module generates an emotion enhancement context representation by concatenating a role-aware emotion embedding representation with a role context representation. The recognition network identifies the emotional states of the speaker and listener based on the emotion enhancement context representation and the target response representation, obtaining initial emotional latent variables for the speaker and listener. The normalized flow-based conditional variational autoencoder dynamically fuses the initial emotional latent variables of the speaker and listener to obtain a global emotional latent variable that integrates the emotional states of both the speaker and listener. The knowledge enhancement module includes a knowledge reasoning module and a knowledge enhancement module. The knowledge reasoning module is used to construct a set of keywords for the current speaker's dialogue sequence. By retrieving and constructing a set of candidate knowledge tuples corresponding to the keywords, a knowledge subset is determined. When a target entity in the knowledge subset matches a candidate keyword, an initial path connection is formed, and the path connection is gradually expanded to the maximum allowed reasoning depth to determine a set of multi-hop knowledge reasoning paths. The knowledge enhancement module is used to obtain the contextual semantic representation of each knowledge reasoning path and fuse it with the global contextual representation of the dialogue history to obtain a knowledge-enhanced contextual representation. The response generation module includes a speech intent prediction module and a context decoding module. The speech intent prediction module is used to construct a keyword-guided knowledge representation matrix based on the knowledge-enhanced context representation. The context decoding module is used to extract multi-scale semantic features from the query vector of the knowledge-enhanced context representation, using the global sentiment latent variable as a bias term, and dynamically adjusting the value vector generated by the knowledge representation matrix and the knowledge-enhanced context representation through semantic features to obtain a multi-scale attention fusion representation. The multi-scale attention fusion representation is concatenated with the global sentiment latent variable and input into the pointer network. By predicting the probability distribution of the current word, an empathetic response is generated.

[0018] Compared with the prior art, the present invention has the following beneficial effects: This invention proposes an empathic dialogue generation method that integrates psychological information. By constructing an emotion-embedded representation of role perception, it enhances the emotion of the role's contextual representation, enabling dynamic learning of richer emotional representations and improving the accuracy of empathic expression. By identifying the emotional states of the speaker and listener and dynamically fusing the identification results, it encompasses both the identification of explicit emotions and the extraction of deep emotional latent variables, providing a foundation for emotion regulation in the response generation stage. By selecting a set of knowledge relevant to the current context from an external knowledge base and forming a multi-hop knowledge reasoning path, knowledge is introduced into the global contextual representation of the dialogue history in a structured manner. The keyword set generated through multi-source feature fusion can explore semantic association paths based on the structured knowledge base, thereby enhancing the hierarchy and contextual fit of the reasoning chain. The resulting knowledge-enhanced contextual representation provides a unified semantic foundation for the response generation stage. The obtained multi-scale attention fusion representation integrates a context-dependent structure with multi-semantic granularity and introduces emotion bias regulation, providing a richer, more coherent, and emotionally expressive semantic foundation for response generation. By combining multi-scale attention fusion representation with global sentiment latent variables, a pointer generation network is used to predict the generation probability distribution of the current word. This pointer generation network integrates vocabulary generation and input replication mechanisms, balancing language fluency with knowledge accuracy. This method effectively improves the system's ability to understand the user's emotional state and the naturalness and accuracy of empathic responses. Attached Figure Description

[0019] Figure 1 This is a flowchart of the empathic dialogue generation method that integrates psychological information proposed in this invention; Figure 2 The network architecture for an empathic dialogue process that integrates psychological information is provided in the embodiments of the present invention; Figure 3 This is the network architecture for the context encoding process of this invention; Figure 4This is a flowchart of the multi-scale multi-head attention mechanism and emotion injection method in the response generation stage of the present invention. Detailed Implementation

[0020] The following will refer to the appendices in the embodiments of the present invention. Figures 1-4 The technical solutions in the embodiments of the present invention will be clearly and completely described. It should be understood that the terminology used in the present invention is only for describing particular implementation methods and is not intended to limit the present invention.

[0021] Example like Figure 1 As shown, this invention proposes a method for generating empathic dialogues that integrates psychological information, comprising the following steps: S1: Obtain the role context representation of the dialogue sequence, the global context representation of the dialogue history, and the target response representation; S2: By constructing the emotional embedding representation of role perception and splicing it with the role context representation, an emotionally enhanced context representation is generated; based on the emotionally enhanced context representation and the target response representation, the global emotional latent variable that integrates the emotional states of the speaker and the listener is obtained by identifying the emotional states of the speaker and the listener and dynamically fusing them. S3: Construct a keyword set for the current speaker's dialogue sequence, and determine a knowledge subset by retrieving and constructing a set of candidate knowledge tuples corresponding to the keywords; when the target entity in the knowledge subset matches the candidate keyword, an initial path connection is formed, and the path connection is gradually expanded to the maximum allowed reasoning depth to determine the set of multi-hop knowledge reasoning paths; obtain the contextual semantic representation of each knowledge reasoning path, and fuse it with the global contextual representation of the dialogue history to obtain a knowledge-enhanced contextual representation. S4: Based on the knowledge-enhanced contextual representation, construct a keyword-guided knowledge representation matrix; extract multi-scale semantic features from the query vector of the knowledge-enhanced contextual representation, use the global sentiment latent variable as a bias term, and dynamically adjust the value vector generated by the knowledge representation matrix and the knowledge-enhanced contextual representation through semantic features to obtain a multi-scale attention fusion representation; S5: The multi-scale attention fusion representation is concatenated with the global sentiment latent variable and then input into the pointer network. By predicting the probability distribution of the current word, an empathetic response is generated.

[0022] Specifically, the implementation process of steps S1-S5 above includes four stages in sequence: context encoding stage, sentiment enhancement stage, knowledge reasoning stage, and response generation stage, wherein: like Figure 2As shown, the empathic dialogue generation network architecture integrating psychological information provided in this embodiment includes a context encoding module, a context decoding module, and a pointer network. The context encoding module mainly involves the context encoding stage, the emotion enhancement stage, and the knowledge reasoning stage. The context decoding module and the pointer network are involved in the response generation stage and are used to guide response generation.

[0023] Context encoding stage To extract multi-layered semantic features of the dialogue and provide a unified and fine-grained representation foundation for the subsequent emotion enhancement, knowledge reasoning, and response generation stages, this stage uses the context encoding module of the empathic dialogue generation network to perform structured, hierarchical encoding of the dialogue history and its role information, such as... Figure 3 As shown.

[0024] The context encoding module includes the GloVe model, a Transformer-based encoder, and a Transformer-based interactive encoder.

[0025] First, the history of dialogue Dialogue sequence divided according to roles Conversation sequence with listeners .in, Indicates the speaker's first One comment, Indicates the listener's corresponding number One piece of feedback.

[0026] At the same time, it introduces the history of current dialogue. Corresponding reference response target response sequence The target response sequence Used in the training phase to guide supervised learning in the emotion enhancement and response generation phases.

[0027] Subsequently, , , , Each statement in the text is marked with a special tag [CLS] and then concatenated to obtain the expanded dialogue history. Target response sequence Speaker dialogue sequence Dialogue Sequence with Listeners The word embedding layer of the GloVe model maps the expanded sequences to a continuous vector space, constructing the contextual semantic representation of the initial dialogue history. Contextual semantic representation of the target response sequence Contextual semantic representation of speaker dialogue sequence Contextual semantic representation of dialogue sequences with listeners :

[0028] ; ; ; ; in, , For vocabulary size, is the dimension of the word vector.

[0029] In order to capture the history of dialogue With target response sequence Global semantic information, and The input is fed into a Transformer-based encoder, denoted as... Each obtains a global context representation of the dialogue history. and target response representation : ; ; Given the significant differences between speakers and listeners in language style and emotional expression, and The inputs are respectively fed into a Transformer-based interactive encoder, denoted as... ,in, and For query, Using keys and values, retrieve the role context representations of the speaker and listener, respectively: ; ; in, and All are interactive encoders based on Transformer, with output vector dimensions of [missing information]. ; , , , ,in and These represent the global context representation of the dialogue history. Target response representation Contextual representation of the roles of speaker and listener and The length.

[0030] Emotion Enhancement Phase In the emotion enhancement stage, in order to enhance the ability of both the speaker and the listener to perceive and express their emotional states in the dialogue, this stage introduces an emotion enhancement mechanism through an emotion enhancement module, which includes three parts: emotion recognition, emotion feature fusion, and emotion latent variable representation.

[0031] The sentiment enhancement module includes a sentiment enhancement context representation generation module, a recognition network, and a normalized flow-based conditional variational autoencoder.

[0032] (1) Emotion recognition To accurately identify changes in emotional state during dialogue, we independently model the emotions of the speaker and the listener to capture the differences and interactivity of their emotions during the communication process.

[0033] Speaker and listener role context representations extracted during the context encoding stage Obtain the hidden state vector corresponding to its [CLS] tag, denoted as and And through linear projection and softmax operation, the probability distributions of the speaker's and listener's emotional categories were obtained respectively: ; ; in, These are the emotion classification parameters for the speaker and the listener, respectively. Number of sentiment categories; The dimension representing the context.

[0034] Specifically, during the training phase, the cross-entropy loss function is used to measure the difference between the predicted probability and the true sentiment label: ; ; in, and These are the speaker's and listener's genuine emotional labels, respectively.

[0035] (2) Integration of emotional characteristics To further enhance the understanding and expression of fine-grained emotional information, after completing the emotion classification, the category label corresponding to the maximum value in the probability distribution of emotion categories is mapped to a low-dimensional vector representation, constructing the emotion embedding representation of the speaker and listener for role perception: ; ; ; ; in, For learnable embedding layers, the output dimension is It is used to transform discrete sentiment tags into semantic vectors in a continuous space.

[0036] To enhance the perception of emotional signals through contextual representation, an emotion-enhanced contextual representation generation module is used to represent the role context of the speaker and listener. and Corresponding sentiment embedding representation and Perform the splicing operation: ; ; Building upon this, a self-attention mechanism is introduced to enhance the interaction between key semantics and emotion, resulting in an emotion-enhanced contextual representation. and : ; ; in, This represents the self-attention mechanism; This represents a linear transformation layer used for projection onto the target spatial dimension; This represents the learnable linear transformation parameters.

[0037] (3) Representation of latent affective variables To capture deep emotional factors and their changing trends during dialogue, a conditional variational autoencoder based on normalized flow is further introduced to enhance the expressive power and flexible control of emotional latent variables. This normalized flow-based conditional variational autoencoder combines the controllable generation characteristics of conditional variational autoencoders with the modeling ability of normalized flow for complex distributions. It can effectively capture the dynamic changes in the distribution of emotional latent variables between the speaker and listener, improving the diversity and precision of emotional expression.

[0038] Step 1 Initial Emotional Latent Variable Calculation First, initialize the initial latent emotional variables of the sampled speakers and listeners using a standard normal distribution: ; Subsequently, contextual representation based on sentiment enhancement. , Representation of target response By identifying the network (denoted as Estimate the mean and standard deviation of the latent affective variables: ; ; in, It represents the target response, assists in modeling latent emotional variables, and improves the targeting of emotion modeling; and It is a contextual representation for emotion enhancement, capturing personalized semantic and emotional cues from the speaker and listener respectively, for fine-grained emotion enhancement modeling; , and , This corresponds to the mean and standard deviation of the speaker's and listener's emotional latent variables.

[0039] To achieve differentiable training of the model, a reparameterization technique is introduced to sample the initial sentiment latent variables: ; in, and These are the initial latent emotional variables for the speaker and the listener, respectively, carrying emotional characteristics. and These are the latent variable distribution parameters of the speaker. and These are the latent variable distribution parameters of the listener.

[0040] Step 2: Normalized Flow Transformation To enhance the expressive power of latent variable distributions and alleviate the pattern collapse problem, this invention introduces a normalization flow mechanism into the Conditional Variational Autoencoder (CVAE) framework, namely, the Normalizing Flow Conditional Variational Autoencoder (NF-CVAE). Through a series of invertible transformation functions, the initial sentiment latent variables are mapped layer by layer to a more complex distribution space, thereby improving the ability of sentiment latent variables to model diverse and nonlinear features.

[0041] Introducing a normalized flow mechanism to analyze initial emotional latent variables and By performing multi-level reversible flow transformation, the emotional latent variables of the speaker and listener after transformation are obtained. and This process can be represented as: ; ; in, It is a set of invertible transformation functions, each of which is invertible and has a computable Jacobian determinant.

[0042] go through Laminar transformation ultimately yields more complex and realistic latent emotional variables of the speaker and listener. and ; Step 3: Integration of Dual Role Emotional Latent Variables Since emotions in a dialogue are often reflected in the interaction between the speaker and the listener, a dynamic fusion mechanism is introduced to dynamically adjust the latent emotional variables of the speaker and the listener after the changes. and The weights in the global affective latent variable are used to integrate the affective latent variables of both the speaker and the listener, resulting in a global affective latent variable that incorporates both the speaker's and listener's affective latent variables. ; ; in, As a global emotional latent variable that integrates the speaker and the listener, it is used to guide the attention mechanism and generation in the response generation stage, thereby enhancing the emotional resonance and psychological adaptability of the generated content. This is the emotional fusion gating coefficient, used to control the dynamic weighting of emotional information from both characters; This represents vector concatenation. For learnable weights, It is the Sigmoid activation function.

[0043] This dual-role emotional latent variable fusion strategy can dynamically control the contribution weight of the emotional information of the two roles according to the current context, thereby improving the empathic adaptability of the generated response.

[0044] (4) NF-CVAE joint training objectives To train the NF-CVAE, the following three joint loss terms are introduced: the flow transform density term, the KL divergence loss, and the reconstruction error term. The total NF-CVAE loss is constructed by combining these three terms, where: The flow transform density term used to model the logarithmic density variation of the flow transform: ; KL divergence loss used to maintain the regularity of the initial latent variable distribution: ; The reconstruction error term used to optimize the generated response: ; The total loss of NF-CVAE is: ; The design of the NF-CVAE total loss ensures that the model can obtain highly expressive sentiment latent variables during the optimization process, while maintaining the rationality of the latent variable distribution and the ability to guide response generation.

[0045] By constructing hierarchical emotion recognition, emotion feature fusion, and emotion latent variable representation, we not only improve our ability to model complex emotional states, but also provide a multi-dimensional, dynamically adaptive emotional semantic foundation for the response generation stage.

[0046] This stage encompasses both the identification of explicit emotions in the speaker and listener and the extraction of latent emotional variables, providing a foundation for emotion regulation in the response generation stage. In particular, the introduction of NF-CVAE proposed in this invention effectively alleviates the problem of latent variable degradation in traditional conditional variational autoencoders, enhancing the emotional controllability and expressive richness of the model-generated content.

[0047] Knowledge Reasoning Stage To enhance the model's knowledge acquisition capabilities and semantic reasoning accuracy within a dialogue context, this invention proposes a reasoning mechanism that integrates multi-strategy keyword extraction and multi-hop knowledge reasoning path expansion. This mechanism generates a keyword set for the current speaker's dialogue sequence through multi-source feature fusion and explores semantic association paths based on a structured knowledge base, thereby strengthening the hierarchy and contextual fit of the reasoning chain.

[0048] (1) Multi-strategy keyword extraction To extract semantically representative keywords that are highly relevant to the context from dialogue text, this invention designs a semantic deduplication module that integrates semantic representation and statistical features for multi-strategy keyword extraction. This semantic deduplication module includes a pre-trained language model, the TF-IDF method, and a scaled dot product attention mechanism. This semantic deduplication module comprehensively utilizes the context modeling capability of the pre-trained language model (PLM), the term weight evaluation capability of TF-IDF, and the context adaptability of the scaled dot product attention mechanism to construct a robust and highly adaptable keyword set to support the subsequent knowledge reasoning process.

[0049] The multi-strategy keyword extraction process includes the following steps: Step 1: Calculate keyword importance scores based on the pre-trained language model. First, the dialogue sequence of the current speaker. Perform word segmentation and obtain a context-aware representation of each token using a pre-trained language model (such as BERT): ; in, , The number of tokens in the dialogue sequence. For the hidden layer dimension.

[0050] For each candidate word The word vector representations are extracted, and the importance scores of the candidate words are calculated using a multilayer perceptron (MLP) of a pre-trained language model. ; ; in, , These are trainable parameters.

[0051] Select the top scorers The candidate words form a preliminary keyword set: ; Step 2: Calculate keyword weights based on TF-IDF. To enhance the interpretability of keyword selection, this invention uses the TF-IDF index to measure the importance of terms, defining the importance score as follows: ; in, Candidate words Frequency of appearance in the current conversation.

[0052] Its inverse document frequency is defined as follows: ; in, This represents the total number of documents. Indicates inclusion The number of documents.

[0053] Similarly, select the top scorers. Terms constitute the TF-IDF candidate set: ; Step 3: Keyword semantic deduplication Since the above two methods may generate semantically redundant keywords, in order to improve the efficiency and coverage of knowledge retrieval, this invention introduces a word form normalization and semantic vector fusion mechanism to remove redundant items.

[0054] First, merge the two candidate sets. Perform word form normalization: ; Next, word vectors obtained based on the pre-trained language model and Calculate the semantic similarity between any pair of words: ; If the similarity of a pair Exceeding the preset threshold If it is, then it is determined to be a redundant term.

[0055] For each pair of redundant terms, a weighted fusion is performed based on their importance scores in the pre-trained language model: ; in, The semantic similarity threshold is determined through optimization on the validation set and is used to control the granularity of semantic merging of keywords.

[0056] Perform a deduplication operation on the merged keyword set to remove redundant terms after semantic merging, resulting in a deduplicated keyword set: ; Step 4 Contextual Relevance Modeling Since different keywords have different importance in specific contexts, this invention uses a scaling dot product attention mechanism to calculate the semantic coupling between candidate words and the global dialogue context in order to improve the contextual adaptability of keyword extraction.

[0057] Get candidate words The word vector representations and the encoded representations of the global dialogue context extracted by the encoder: ; ; A scaled dot product attention mechanism is used to calculate the degree of matching between candidate words and context: ; in, , , For a trainable parameter matrix, This represents the dimension of the context representation, used to calculate the degree of matching between candidate words and the current context.

[0058] Attention weights between candidate words and the current dialogue context Indicate candidate words The degree of semantic matching with the current dialogue context.

[0059] Based on this, the attention weights between candidate words and the current dialogue context are fused with the word scores calculated using the pre-trained language model and the TF-IDF method to construct a context-aware keyword importance scoring function: ; in, Keyword importance scores are calculated based on pre-trained language models and are used to measure semantic relevance. The keyword information content score is calculated based on the TF-IDF method and is used to capture high-information words; The two score contributions are dynamically adjusted to adapt to the specific dialogue context.

[0060] Step 5 Keyword Fusion Score Calculation To integrate the three scoring strategies and further improve the accuracy and representativeness of keyword selection, the final keyword score is defined as a weighted fusion of the three scores: ; in, , These are used to control the fusion weights for different scores; Used to measure semantic representativeness and ensure that keywords have global semantic relevance; Used to evaluate the information contribution, emphasizing the information richness of keywords; It is used to reflect contextual adaptability and improve the actual applicability of keywords in the current dialogue.

[0061] Candidate words are ranked according to their overall scores, and the top ones are selected. The highest-scoring terms constitute the final keyword set, which is the keyword set of the current speaker's dialogue sequence. : ; The final keyword set obtained through the above process It possesses high semantic representativeness, context adaptability, and information redundancy removal capabilities, providing robust information support for subsequent knowledge retrieval and the construction of multi-hop knowledge reasoning paths.

[0062] (2) Knowledge retrieval and filtering After completing keyword extraction and obtaining the final keyword set Then, the model needs to introduce semantic support information from an external common sense knowledge base to enhance the knowledge reasoning ability of the dialogue system.

[0063] This invention designs a keyword-driven knowledge retrieval mechanism and a multi-level filtering strategy for the knowledge retrieval and filtering module. This knowledge retrieval and filtering module, together with the aforementioned semantic deduplication module, constitutes the content of the knowledge reasoning module. The knowledge retrieval and filtering module is used to filter out a set of knowledge that is highly relevant to the current context and has a reliable structure from a large-scale knowledge graph.

[0064] The knowledge retrieval and filtering module includes a knowledge retrieval module, a context fusion encoder, and a knowledge reasoning path generation module.

[0065] Step 1: Keyword-driven knowledge retrieval With the final set of keywords As the entry point for retrieval, the corresponding triples are retrieved from the external commonsense graph (such as ConceptNet) of the knowledge retrieval module. And obtain its corresponding confidence score. This forms a set of candidate knowledge quadruples with confidence, denoted as: ; in, As keywords, This refers to the type of relationship between keywords and target concepts. The target concept in a triplet. The knowledge confidence score represents the degree of credibility of the triple in the knowledge base. This is a set of relationships related to the keywords.

[0066] Step 2: Knowledge Quadruple Filtering and Reordering To select knowledge content that is highly relevant to the current dialogue context and semantically reliable, a comprehensive scoring mechanism integrating semantic similarity and knowledge confidence scores is introduced to evaluate the candidate knowledge quadruple set. Perform filtering and reordering.

[0067] Define the concept of computational target Its corresponding context feature vector The comprehensive scoring function between them: ; in, By calculating the target concept With context feature vector Cosine similarity is used to measure the semantic closeness between the target concept and the semantic vector of the current context; The knowledge confidence score measures the credibility of the quadruple in the knowledge base.

[0068] Based on the comprehensive scoring function, the candidate knowledge quadruple set is evaluated. Sort and select the top These high-quality knowledge items constitute the final knowledge subset: ; This knowledge set It serves as the input foundation for constructing multi-hop knowledge reasoning paths and knowledge fusion.

[0069] (3) Construction of multi-hop reasoning paths The quadruplets obtained during the knowledge retrieval stage often only provide local explicit semantic associations, which are insufficient to meet the reasoning needs in complex contexts. Therefore, to improve the model's knowledge reasoning ability, this invention introduces a knowledge reasoning path generation module based on a keyword-driven multi-hop knowledge path construction mechanism to establish deeper semantic association chains, enabling the model to make reasonable transitions and reasoning completions between potential concepts.

[0070] Step 1: Keyword Context Feature Modeling To achieve multi-hop path expansion starting from an entity, the keyword set is first calculated. Chinese keywords The contextual semantic features are represented. Specifically, this is combined with keyword embedding representation. Dialogue Context Encoding Extracting keywords through a context fusion encoder Semantic feature representation in the current context:

[0071] ; in, Keywords Embedded representation, A semantic encoding representation of the complete dialogue context. Keywords Semantic feature representation in the current context.

[0072] Step 2: Construct the reasoning path Based on keyword context feature modeling, a multi-hop knowledge reasoning path is further constructed based on entity semantic similarity.

[0073] In the knowledge reasoning path generation module, the knowledge set obtained in the previous step is combined... Determine the target concept in the quadruple. If a semantic connection exists between the target concept and other keywords, and if the condition is met, a one-hop path is established. The specific determination method involves calculating the target concept. With keywords Semantic similarity:

[0074] ; If the similarity exceeds a set threshold (i.e., the target concept) With a certain keyword If a match is found, it is considered that a valid semantic jump path exists, forming the following connection: ; Building upon this, recursive reasoning is employed to construct longer knowledge paths, gradually expanding to the maximum permissible depth of reasoning. Form a set of multi-hop paths: ; in, For a set of knowledge reasoning paths, For the updated keyword set, The current depth of the path. The maximum allowed inference depth is used to control the path expansion boundary.

[0075] Finally, the top-K paths with the best overall quality are selected from all construction paths to form the set of multi-hop knowledge reasoning paths.

[0076] Step 3: Path Attention Weighting To enhance the guiding ability of reasoning paths during the response generation stage, this invention introduces an attention mechanism to weight different paths, strengthening paths with greater reasoning value in the context. The specific calculation is as follows: By assigning higher weights to high-quality reasoning paths through an attention mechanism, the final fused knowledge is made more valuable. To select the most relevant reasoning paths, the attention weights of each knowledge reasoning path are calculated. :

[0077] in, The attention weights corresponding to each knowledge reasoning path, This refers to the final set of keywords involved in the path. This represents the global context of the dialogue history.

[0078] Step 4: Exception Handling Mechanism Considering that paths in knowledge graphs may be incomplete or interrupted, this invention designs the following two anomaly handling mechanisms to improve the robustness and usability of the model: (1) Parameter expansion strategy. Appropriately increase path depth. Or relax the similarity matching threshold to increase the number of available paths. (2) Fallback mechanism. If the path is unreachable or the path construction is interrupted, fall back to use the original one-hop triple knowledge as a substitute.

[0079] The above processing mechanism ensures that each round of dialogue can obtain the best available knowledge support, avoiding semantic interruption or generation crash due to missing paths.

[0080] (4) Knowledge integration After completing keyword extraction, knowledge retrieval, and multi-hop reasoning path construction, it is still necessary to integrate external knowledge into the dialogue context representation in a structured manner to improve the knowledge support capability and contextual fit of the generated response. To this end, this invention designs a knowledge enhancement module that integrates the semantic representation of knowledge reasoning paths with an attention mechanism of contextual information, constructing a deep fusion representation of knowledge and context, and providing a unified semantic foundation for the response generation stage.

[0081] The knowledge enhancement module includes a knowledge reasoning path semantic generation module and a knowledge fusion module.

[0082] Step 1: Semantic Representation of Knowledge Paths The knowledge reasoning path semantic generation module is used to convert the multi-hop knowledge reasoning path obtained by reasoning into a natural language sequence according to the structural order, and input it into a bidirectional gated recurrent unit (Bi-GRU) to obtain its contextual semantic representation: ; in, This represents a text sequence composed of knowledge reasoning paths. This serves as a contextual representation of the knowledge reasoning path. The path sequence length. This refers to the hidden layer dimension of the knowledge reasoning path, which is the dimension of the feature vector used to represent the semantic information in the knowledge reasoning path.

[0083] Step 2: Integrating Knowledge with Dialogue Context To achieve knowledge-guided response generation, a knowledge fusion module is used to represent knowledge... Representation of global context of dialogue history Dynamic fusion is performed to obtain knowledge-enhanced contextual representations. .

[0084] For the context of the first One token: ; ; in, For the first in the dialogue context The representation of a token, Indicates the first The knowledge fusion weights corresponding to each position are calculated through a scaling dot product attention mechanism, enabling the model to dynamically adjust its dependence on different knowledge.

[0085] Dialogue intention refers to the interactive intent or functional goal expressed in utterances within a specific context, such as asking questions, making suggestions, offering encouragement, or refuting statements. Dialogue intention not only reflects the speaker's understanding and situational judgment of the current context but also embodies the tendency of their communication strategy. Therefore, accurately modeling dialogue intention is crucial for improving the interactivity and contextual fit of generated content during the response generation stage.

[0086] To achieve response generation based on dialogue intent, this invention designs a dialogue intent prediction module in the knowledge reasoning stage. By introducing dialogue intent prediction and embedding mechanisms, the strategic nature and adaptability of the generated response are improved.

[0087] Step 1: Calculation of Dialogue Intent Probability Distribution First, from the contextual representation of knowledge enhancement. Extract the hidden state corresponding to the [CLS] marker. This serves as an aggregated representation of global semantics. Then, a linear transformation and a softmax function are used to predict the probability distribution of the intent category for generating the current utterance:

[0088] ; in, It is the learned weight matrix that maps the context representation to the intent category space; The total number of dialogue intent categories, The dimension representing the context.

[0089] Step 2: Dialogue Intent Prediction Loss Function To optimize the accuracy of intent prediction, this invention employs the cross-entropy loss function, which minimizes the difference between the predicted distribution and the true label based on the probability distribution of the intent category, thereby ensuring the model's accuracy in recognizing dialogue intent. ; in, One-hot encoding of the true dialogue intent label. The model predicts the first The probability of class intent.

[0090] Step 3: Embedded Representation of Dialogue Intent In order to effectively inject the predicted dialogue intent information into the response generation process, the predicted intent category The dialogue intent embedding vector is obtained by mapping a high-dimensional vector representation through the embedding layer in the dialogue intent prediction module and then processing it through the embedding layer. ; in, The dialogue intent embedding vector not only captures the semantic information of the intent category, but also provides clear interaction guidance for the response generation stage; The dimension of the intent embedding is the dimension of the feature vector used to represent the intent category.

[0091] By introducing a dialogue intent prediction and embedding mechanism, this invention can not only identify the interaction goal reflected in the current utterance, but also provide strategic guidance information for the response generation stage, thereby improving the relevance of the response and the naturalness of human-computer interaction.

[0092] Response decoding phase To effectively integrate multi-source semantic information such as emotion, knowledge, and dialogue intent, this invention designs a multi-scale multi-head attention mechanism (MS-MHA) based on the standard Transformer decoder structure, and combines it with emotion latent variables to guide the attention distribution, thereby enhancing the emotional expression ability and semantic adaptability of the generated response.

[0093] Step 1: Dialogue Intent Guidance Initialization To ensure that the generated response statements have clear pragmatic goals, the dialogue intent is embedded in the vector. With the decoder's starting tag vector The vectors are concatenated to form a fused vector, which is then mapped to the dimensions required by the decoder to obtain the initial decoding state, which serves as the starting input to the decoder. ; in, This represents the dialogue intent embedding vector, used to guide the response generation process; It is the embedding of the start tag, representing the semantic state of the decoder at the initial moment; This represents a vector concatenation operation, which concatenates two vectors along the feature dimension. The weight matrix is ​​a learnable matrix used to map the concatenated vector to the dimensions of the decoder's hidden layers. It also controls how the dialogue intent vector and the starting tag vector are passed to the decoder.

[0094] Step 2 Knowledge Enhancement and Context Fusion Due to knowledge-enhanced contextual representation It integrates external knowledge and dialogue context information, and uses them as key and value inputs in the cross-attention mechanism in the decoder to achieve semantic alignment and interaction between the dialogue context and response sequence during the generation process.

[0095] Simultaneously, attention weight vectors for each knowledge reasoning path are introduced. According to the initial decoding state By performing element-wise weighted operations, a keyword-guided knowledge representation matrix is ​​constructed: ; in, This refers to the weight vector of each hop of the knowledge reasoning path obtained from the attention distribution of the knowledge reasoning path. This indicates an element-wise multiplication operation. For context length.

[0096] Step 3: Multi-scale, multi-head attention mechanisms and emotional infusion To enhance the model's adaptability to different semantic scales, a multi-scale, multi-head attention mechanism is introduced into the standard Transformer decoder, forming the context decoding module of the empathic dialogue generation network. This mechanism extracts knowledge-enhanced contextual representations through parallel one-dimensional convolutional operations. The query vector is analyzed using multi-scale semantic features, and dynamic attention regulation is achieved by fusing sentiment latent variables.

[0097] The context decoding module includes a mask self-attention mechanism, a cross-attention mechanism, a multi-scale multi-head attention mechanism, and a feedforward neural network.

[0098] Embed the generated dialogue intent vector As input to the decoder, the dependencies of the dialogue history rounds are modeled through a masked self-attention mechanism to obtain the intent representation of the current dialogue.

[0099] Based on the intent representation of the current dialogue and the global context representation of the dialogue history, the external context is fused through the key and value of the cross-attention mechanism to obtain a global context representation that incorporates external information.

[0100] By employing a multi-scale, multi-head attention mechanism, the query vector, which is generated internally by the decoder and incorporates external information into a global context representation, is input into multiple parallel one-dimensional convolutional modules to extract knowledge and enhance the context representation. Multi-scale semantic features of query vectors.

[0101] Specifically, in order to capture semantic information at different scales, such as Figure 4 As shown, three different kernel sizes (kernel size = 3, 5, 7) are used to convolve the query vector. Perform one-dimensional convolution operations to extract features at different scales: ; ; ; in, It primarily captures local dependencies at the n-gram level, which helps in understanding short-range semantic connections between words; It has strong perception ability and can capture structures of medium span (such as subject, verb, and object). It helps to construct long-distance dependent information and enhances the global understanding of the discourse context.

[0102] By integrating multi-scale semantic features through a feedforward neural network, a unified representation of multi-scale semantic features is generated.

[0103] Subsequently, for the query vector at each scale ( Multi-head attention operations are performed separately. To incorporate emotional information, latent emotional variables are introduced into the attention calculation. As a bias term, the value vector generated by the knowledge representation matrix and the knowledge-enhanced context representation is dynamically adjusted through the multi-scale semantic features of the unified representation:

[0104] ; in, This represents the attention weights of each knowledge reasoning path. Weighted knowledge representation matrix; Value vectors generated for knowledge-enhanced contextual representations; As a global sentiment latent variable, it serves as a bias term to control attention distribution and guide the model to focus on sentiment-related content in the context.

[0105] To integrate multi-scale attention results, this invention introduces a learnable fusion weight vector. And the adaptive fusion coefficients at each scale are obtained through softmax: ; The multi-scale attention fusion is represented as: ; This multi-scale attention fusion representation integrates a context-dependent structure with multiple semantic granularities and introduces sentiment bias modulation, providing a richer, more coherent, and emotionally expressive semantic foundation for response generation.

[0106] After fusing multi-source information such as sentiment enhancement and knowledge reasoning, this invention employs a Pointer Generator Network (PGN) to generate the final response sequence. The PGN combines vocabulary generation with input replication mechanisms, balancing language fluency with knowledge accuracy.

[0107] Multi-scale attention fusion representation With global sentiment latent variables The concatenated data is input into a pointer network to predict the current word. Generation probability distribution: ; in, This indicates that the pointer generation network supports generating or copying input content from a vocabulary, thereby enhancing the entity coverage and empathic expression of the generated content.

[0108] Emotional latent variables This method is used to predict the probability of guiding words. It selects word sequences using greedy algorithms, bundle search, and sampling to generate empathetic responses that contain precisely replicated entities and flow smoothly. This ensures that the generated empathetic responses are emotionally natural, semantically coherent, and consistent with the meaning of the dialogue. Figure 1 To.

[0109] In order to comprehensively optimize the generation quality, emotional expression preservation, and dialogue meaning Figure 1 To achieve multiple objectives such as consistency, this invention also introduces multiple loss functions for joint optimization, which are defined as follows: in, and The emotion classification losses for speakers and listeners are used to constrain the latent emotion variables. The semantic representation space is improved to more accurately reflect the emotional state of both parties and enhance the controllability of emotional expression. The dialogue intent classification loss ensures that the generated intent embedding representation is accurate. Consistency with real labels guides the response generation process to align with the intended dialogue interaction goals; The variational loss of the conditional variational autoencoder based on normalized flow is used to enhance the expressive power of latent variables, effectively alleviate the problem of latent variable vanishing, and improve the model's ability to capture sentiment information. Bag-of-Words Loss effectively alleviates the vanishing latent variable problem and enhances the semantic diversity of generated content by promoting the reconstruction of the response term set by latent variables; various hyperparameters Used to control the relative weights of each loss term during joint training.

[0110] In summary, the empathic dialogue generation method proposed in this invention, which integrates psychological information, provides semantic support for subsequent stages by constructing role context representations, global context representations of dialogue history, and target response representations of the dialogue sequence during the context encoding stage.

[0111] During the emotion enhancement phase, by combining emotion tags with role context representations, fine-grained changes in the emotional interactions between the two parties are accurately captured, enhancing the dynamic expressiveness of emotions. A conditional variational autoencoder based on normalized flow guidance is employed, effectively mitigating the pattern collapse problem in traditional conditional variational autoencoder models through a multi-layer flow transformation mechanism, making the modeling of latent emotion variables more flexible and diverse.

[0112] During the knowledge reasoning stage, high-quality keywords from the dialogue history are selected and their weight allocation is optimized based on contextual dependencies, effectively reducing redundancy and noise interference and ensuring that the system focuses on the knowledge items with the strongest semantic relevance. Then, a multi-hop knowledge reasoning path is constructed to enhance the cognitive depth of the context and the fit of the context, providing accurate and emotionally resonant knowledge support for response generation.

[0113] The probability distribution of the dialogue intent of the current utterance is predicted by using knowledge-enhanced contextual representation, and the intent embedding vector is used as pragmatic target information to guide the decoder to generate a response that conforms to the dialogue objective.

[0114] In the response generation phase, the initial decoding state is embedded based on the dialogue intent, and combined with the keyword-guided knowledge representation matrix and sentiment bias information, a multi-scale semantic feature of dialogue intent, sentiment state, and knowledge path is fused through a multi-scale multi-head attention mechanism. Finally, this fused representation is concatenated with the global sentiment latent variable and input into the pointer generation network to predict the generation probability of the current word, thereby achieving dialogue response generation with empathy, adaptability, and knowledge support.

[0115] Based on this method, the present invention also proposes an empathic dialogue generation system that integrates psychological information, including: The context encoding module is used to obtain the role context representation of the dialogue sequence, the global context representation of the dialogue history, and the target response representation; The emotion enhancement module includes an emotion enhancement context representation generation module, a recognition network, and a normalized flow-based conditional variational autoencoder. The emotion enhancement context representation generation module generates an emotion enhancement context representation by concatenating a role-aware emotion embedding representation with a role context representation. The recognition network identifies the emotional states of the speaker and listener based on the emotion enhancement context representation and the target response representation, obtaining initial emotional latent variables for the speaker and listener. The normalized flow-based conditional variational autoencoder dynamically fuses the initial emotional latent variables of the speaker and listener to obtain a global emotional latent variable that integrates the emotional states of both the speaker and listener. The knowledge enhancement module includes a knowledge reasoning module and a knowledge enhancement module. The knowledge reasoning module is used to construct a set of keywords for the current speaker's dialogue sequence. By retrieving and constructing a set of candidate knowledge tuples corresponding to the keywords, a knowledge subset is determined. When a target entity in the knowledge subset matches a candidate keyword, an initial path connection is formed, and the path connection is gradually expanded to the maximum allowed reasoning depth to determine a set of multi-hop knowledge reasoning paths. The knowledge enhancement module is used to obtain the contextual semantic representation of each knowledge reasoning path and fuse it with the global contextual representation of the dialogue history to obtain a knowledge-enhanced contextual representation. The response generation module includes a speech intent prediction module and a context decoding module. The speech intent prediction module is used to construct a keyword-guided knowledge representation matrix based on the knowledge-enhanced context representation. The context decoding module is used to extract multi-scale semantic features from the query vector of the knowledge-enhanced context representation, using the global sentiment latent variable as a bias term, and dynamically adjusting the value vector generated by the knowledge representation matrix and the knowledge-enhanced context representation through semantic features to obtain a multi-scale attention fusion representation. The multi-scale attention fusion representation is concatenated with the global sentiment latent variable and input into the pointer network. By predicting the probability distribution of the current word, an empathetic response is generated.

[0116] In short, this invention constructs a role context representation of the dialogue sequence, a global context representation of the dialogue history, and a target response representation through a context encoding module; identifies and integrates the emotional states of the speaker and listener through an emotion enhancement module, and improves the emotional expression ability of the generated response through the enhancement of emotional features; enhances the knowledge reasoning ability and semantic relevance in the dialogue context through a knowledge enhancement module, and further guides response generation by predicting dialogue intent; the response generation module enhances the emotional expression and contextual adaptability of the response by introducing a multi-scale multi-head attention mechanism and using global emotional latent variables as bias terms.

[0117] The system can accurately identify and respond to users' emotional states, generating more human and empathetic responses. It also demonstrates strong emotion recognition, knowledge utilization, and dialogue intent matching capabilities in multi-turn dialogues, significantly improving the depth of the dialogue system's emotional and semantic understanding.

[0118] In summary, the empathic dialogue generation method integrating psychological information constructed in this invention has broad application potential and practical significance in intelligent interactive systems. Firstly, in the field of mental health support, it can be used to construct emotional companion robots and intelligent psychological counseling systems, enabling proactive identification and emotional response to users' emotional states. Secondly, in intelligent education scenarios, it can serve as a virtual learning partner with emotional understanding capabilities, providing learners with more empathetic and interactive guidance. Furthermore, in intelligent customer service and social robot systems, it can enhance the naturalness and emotional coherence of user interactions, improving overall user experience and service satisfaction. Overall, this invention not only achieves a technological breakthrough in generating empathic dialogues integrating psychological information but also provides a feasible path for constructing human-computer dialogue systems with cognitive understanding and emotional perception capabilities, laying a solid foundation for the development of intelligent interactive technology.

[0119] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

[0120] Furthermore, unless otherwise stated, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. All references to this specification are incorporated by way of citation to disclose and describe methods relating to those references. In the event of any conflict with any incorporated reference, the content of this specification shall prevail.

Claims

1. A method for generating empathic dialogues that integrates psychological information, characterized in that, Includes the following steps: Obtain the role context representation of the dialogue sequence, the global context representation of the dialogue history, and the target response representation; By constructing and concatenating the emotional embedding representation of role perception with the role context representation, an emotionally enhanced context representation is generated; based on the emotionally enhanced context representation and the target response representation, the global emotional latent variable that integrates the emotional states of the speaker and the listener is obtained by identifying the emotional states of the speaker and the listener and dynamically fusing them. Construct a keyword set for the current speaker's dialogue sequence, and determine a knowledge subset by retrieving and constructing a set of candidate knowledge tuples corresponding to the keywords; when the target entity in the knowledge subset matches the candidate keyword, an initial path connection is formed, and the path connection is gradually expanded to the maximum allowed reasoning depth to determine the set of multi-hop knowledge reasoning paths; obtain the contextual semantic representation of each knowledge reasoning path, and fuse it with the global contextual representation of the dialogue history to obtain a knowledge-enhanced contextual representation; Based on the knowledge-enhanced contextual representation, construct a keyword-guided knowledge representation matrix; Multi-scale semantic features are extracted from the query vector of the knowledge-enhanced context representation. The global sentiment latent variable is used as a bias term. The semantic features are used to dynamically adjust the value vector generated by the knowledge representation matrix and the knowledge-enhanced context representation to obtain a multi-scale attention fusion representation. The multi-scale attention fusion representation is concatenated with global sentiment latent variables and then input into the pointer network. By predicting the probability distribution of the current word, an empathetic response is generated.

2. The method for generating empathic dialogues by integrating psychological information according to claim 1, characterized in that, The acquisition of the role context representation of the dialogue sequence, the global context representation of the dialogue history, and the target response representation is obtained by encoding the input dialogue sequence through a context encoding module, including the following steps: The context encoding module includes a GloVe model, a Transformer-based encoder, and a Transformer-based interactive encoder. Dialogue history Dialogue sequence divided according to roles Conversation sequence with listeners And introduce dialogue history The corresponding target response sequence of the reference response ; On the history of dialogue Target response sequence Speaker dialogue sequence Conversation sequence with listeners Each statement in the text is marked with a special tag [CLS] and then concatenated to obtain the expanded dialogue history. Target response sequence Speaker dialogue sequence Dialogue Sequence with Listeners ; The expanded sequences are mapped to a continuous vector space through the word embedding layer of the GloVe model, thereby obtaining the global contextual semantic representation of the dialogue history. Contextual semantic representation of the target response sequence Contextual semantic representation of speaker dialogue sequence Contextual semantic representation of dialogue sequences with listeners ; Will and Input a Transformer-based encoder to obtain a global context representation of the dialogue history. and target response representation ; Based on pre-trained language models, and Input to a Transformer-based interactive encoder, and As a query, As keys and values, respectively, we obtain the role context representations of the speaker's dialogue sequence and the listener's dialogue sequence. and .

3. The method for generating empathic dialogues by integrating psychological information according to claim 2, characterized in that, The contextual representation of the emotion enhancement is obtained through the emotion enhancement module, specifically including: The emotion enhancement module includes an emotion enhancement context representation generation module; Based on the context of the speaker's and listener's roles and The hidden state vectors corresponding to the [CLS] labels are extracted respectively, and obtained by linear mapping and the Softmax function. and The probability distribution of the sentiment category; according to and The probability distribution of sentiment categories is used, and the sentiment category label corresponding to the maximum value of the probability distribution is mapped to a low-dimensional vector representation through the sentiment embedding layer. Sentiment embedding representations for the speaker and the listener are constructed respectively. and ; The emotion-enhanced context representation generation module will and Representation of the corresponding role context and By splicing the data and introducing a self-attention mechanism, an emotionally enhanced contextual representation is obtained. and .

4. The method for generating empathic dialogues by integrating psychological information according to claim 3, characterized in that, The emotion enhancement module further includes a recognition network and a conditional variational autoencoder based on normalized flow. The recognition network identifies the emotional states of the speaker and listener, obtaining initial latent emotional variables for both. The conditional variational autoencoder based on normalized flow dynamically fuses these initial latent emotional variables to obtain global latent emotional variables that integrate the emotional states of both the speaker and listener. Specifically, this includes: Initialize the latent emotional variables of the speaker and listener separately from the standard normal distribution; Contextual representation based on sentiment enhancement and With target response representation The mean and standard deviation of the speaker's and listener's emotional latent variables are estimated using the recognition network. Based on the mean and standard deviation, a reparameterization technique is introduced to sample the emotional latent variables, obtaining initial emotional latent variables with emotional characteristics. and ; The conditional variational autoencoder based on normalized flow is a conditional variational autoencoder that sequentially introduces a normalized flow mechanism and a dynamic fusion mechanism. A normalized flow mechanism is introduced to analyze the initial emotional latent variables. and By performing multi-level reversible flow transformation, the emotional latent variables of the speaker and listener after transformation are obtained. and ; A dynamic fusion mechanism is introduced to dynamically adjust the emotional latent variables of the speaker and listener after transformation. and The weights in the global affective latent variable are used to obtain the global affective latent variable that integrates the speaker's and listener's emotions. .

5. The method for generating empathic dialogues by integrating psychological information according to claim 1, characterized in that, The keyword set of the current speaker's dialogue sequence is constructed using a semantic deduplication module, specifically including: The semantic deduplication module includes a pre-trained language model, the TF-IDF method, and a scaled dot product attention mechanism. Extract each candidate word The word vector representation is used to calculate the semantic importance score of candidate words through the multilayer perceptron in the pre-trained language model, thus obtaining the candidate word score calculated based on the pre-trained model; and the statistical weight score of candidate words is calculated based on the TF-IDF method, thus obtaining the candidate word score calculated based on the TF-IDF method. The top K words with scores calculated using both the pre-trained model and the TF-IDF method are selected to form a preliminary keyword set and a TF-IDF candidate set, respectively. Lexical normalization and semantic vector fusion are then performed on the preliminary keyword set and the TF-IDF candidate set, respectively, to obtain a fused keyword set. Finally, a deduplication operation is performed on the fused keyword set to remove redundant terms after semantic merging, resulting in a deduplicated keyword set. ; Based on the word vector representations of the acquired candidate words and the global context representation of the dialogue history, the matching degree between each candidate word and the context is calculated through a scaling dot product attention mechanism to obtain the attention weight between the candidate word and the current dialogue context. The attention weights between candidate words and the current dialogue context are fused with candidate word scores calculated based on the pre-trained model and the TF-IDF method to construct a context-aware keyword importance scoring function. The context-aware keyword importance scoring function, the candidate word scores calculated based on the pre-trained model and the TF-IDF method are then weighted and fused to obtain a comprehensive score. Candidate words are ranked according to the comprehensive score, and the top K words with the highest scores are selected to form the final keyword set. That is, the set of keywords in the current speaker's dialogue sequence. .

6. The method for generating empathic dialogues by integrating psychological information according to claim 5, characterized in that, The multi-hop knowledge reasoning path set is determined through the knowledge reasoning module, specifically including: The knowledge reasoning module includes a semantic deduplication module and a knowledge retrieval and filtering module; The knowledge retrieval and filtering module includes a knowledge retrieval module, a context fusion encoder, and a knowledge reasoning path generation module. Keyword set based on the current speaker's dialogue sequence The system retrieves knowledge triples related to the current dialogue context from the external knowledge base of the knowledge retrieval module, obtains their corresponding confidence scores, and constructs a set of candidate knowledge quadruples with confidence scores. By determining a comprehensive scoring function between the target entity and its corresponding context feature vector, the candidate knowledge quadruples are filtered and reordered, and the top K quadruples with the highest scores are selected to form a knowledge subset. ; Extract keywords using a context fusion encoder Semantic feature representation in the current context; The knowledge reasoning path generation module is used to calculate knowledge subsets based on semantic feature representations. target entity in With keywords Semantic similarity; determining the target concept in a quadruple. Whether there is a semantic connection between it and other keywords; when the semantic similarity exceeds a set threshold, that is, the target concept... With a certain keyword If a match is found, a semantic jump path exists, forming a path connection; recursive reasoning is used to expand the path to the set maximum allowed reasoning depth. Generate a set of multi-hop paths The paths are then sorted, and the top K paths are selected as the optimal set of multi-hop knowledge reasoning paths.

7. The method for generating empathic dialogues by integrating psychological information according to claim 6, characterized in that, The knowledge enhancement module obtains the contextual representation of the knowledge enhancement, specifically including: The knowledge enhancement module includes a knowledge reasoning path semantic generation module and a knowledge fusion module; The knowledge reasoning path semantic generation module is used to convert the knowledge reasoning paths in the knowledge reasoning path set into natural language sequences, and input them into a bidirectional gated loop unit to obtain the contextual semantic representation of the knowledge reasoning path set. ; Get the global context representation of the dialogue history The first in The representation of a token ; The scaling dot product attention mechanism is used to compute the first... Knowledge fusion weights corresponding to each position ; The knowledge fusion module is used to integrate... and Contextual representation of knowledge reasoning paths through element-wise weighting Representation of the context of the dialogue Dynamic fusion is performed at various locations to obtain a knowledge-enhanced contextual representation. .

8. The method for generating empathic dialogues by integrating psychological information according to claim 7, characterized in that, The keyword-guided knowledge representation matrix is ​​constructed through the speech intent prediction module, specifically including: From knowledge-enhanced contextual representation Extract the hidden state corresponding to the [CLS] marker. Based on the hidden state The probability distribution of intent categories in the current dialogue is generated through linear transformation and Softmax prediction. The predicted intent category is mapped to a high-dimensional vector representation through the embedding layer in the dialogue intent prediction module, generating a dialogue intent embedding vector. ; Embedding dialogue intent into vectors With the decoder's starting tag vector The vectors are concatenated to form a fused vector, which is then mapped to the dimensions required by the decoder to obtain the initial decoding state. , as the starting input of the decoder; Obtain the attention weight vector for each knowledge reasoning path. Contextual representation that enhances knowledge As the key and value inputs in the cross-attention mechanism of the decoder, based on the initial decoding state A knowledge representation matrix is ​​constructed through element-wise weighted operations. .

9. The method for generating empathic dialogues by integrating psychological information according to claim 8, characterized in that, The multi-scale attention fusion representation is obtained by decoding through the context decoding module, including the following steps: The context decoding module includes a mask self-attention mechanism, a cross attention mechanism, a multi-scale multi-head attention mechanism, and a feedforward neural network; Embed the generated dialogue intent vector As input to the decoder, the dependencies of the dialogue history rounds are modeled through a masked self-attention mechanism to obtain the intent representation of the current dialogue; Based on the intent representation of the current dialogue and the global context representation of the dialogue history, the external context is fused through the key and value of the cross-attention mechanism to obtain a global context representation that incorporates external information; By employing a multi-scale, multi-head attention mechanism, the query vector, which is generated internally by the decoder and incorporates external information into a global context representation, is input into multiple parallel one-dimensional convolutional modules to extract knowledge and enhance the context representation. Multi-scale semantic features of query vectors; By integrating multi-scale semantic features through a feedforward neural network, a unified representation of multi-scale semantic features is generated. Introducing latent emotional variables As a bias term, the value vector generated by the knowledge representation matrix and the knowledge-enhanced context representation is dynamically adjusted through multi-scale semantic features of a unified representation. Attention scores are calculated using Softmax, and the query vector at each scale is adjusted accordingly. (i=1, 2, 3) Perform multi-head attention fusion respectively to obtain multi-scale attention fusion output. .

10. A system for generating empathic dialogues that integrates psychological information, characterized in that, include: The context encoding module is used to obtain the role context representation of the dialogue sequence, the global context representation of the dialogue history, and the target response representation; The emotion enhancement module includes an emotion enhancement context representation generation module, a recognition network, and a normalized flow-based conditional variational autoencoder. The emotion enhancement context representation generation module generates an emotion enhancement context representation by concatenating a role-aware emotion embedding representation with a role context representation. The recognition network identifies the emotional states of the speaker and listener based on the emotion enhancement context representation and the target response representation, obtaining initial emotional latent variables for the speaker and listener. The normalized flow-based conditional variational autoencoder dynamically fuses the initial emotional latent variables of the speaker and listener to obtain a global emotional latent variable that integrates the emotional states of both the speaker and listener. The knowledge enhancement module includes a knowledge reasoning module and a knowledge enhancement module. The knowledge reasoning module is used to construct a set of keywords for the current speaker's dialogue sequence. By retrieving and constructing a set of candidate knowledge tuples corresponding to the keywords, a knowledge subset is determined. When a target entity in the knowledge subset matches a candidate keyword, an initial path connection is formed, and the path connection is gradually expanded to the maximum allowed reasoning depth to determine a set of multi-hop knowledge reasoning paths. The knowledge enhancement module is used to obtain the contextual semantic representation of each knowledge reasoning path and fuse it with the global contextual representation of the dialogue history to obtain a knowledge-enhanced contextual representation. The response generation module includes a speech intent prediction module and a context decoding module. The speech intent prediction module is used to construct a keyword-guided knowledge representation matrix based on the knowledge-enhanced context representation. The context decoding module is used to extract multi-scale semantic features from the query vector of the knowledge-enhanced context representation, using the global sentiment latent variable as a bias term, and dynamically adjusting the value vector generated by the knowledge representation matrix and the knowledge-enhanced context representation through semantic features to obtain a multi-scale attention fusion representation. The multi-scale attention fusion representation is concatenated with the global sentiment latent variable and input into the pointer network. By predicting the probability distribution of the current word, an empathetic response is generated.

Citation Information

Cited By

  • A respiratory disease clinical decision making method, system, device, and medium

    CN122224478A