Man-machine conversation method and system based on retrieval enhancement
By introducing dynamic example search and cognitive understanding modules into the system, combined with a multi-knowledge fusion decoder, the challenges of existing systems in generating context-related and empathetic replies in an open dialogue environment are solved, significantly improving the user's dialogue experience and the system's empathy and context-understanding abilities.
Patent Information
- Application Number
- CN202510093800.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing systems have challenges in generating context-sensitive and empathetic replies, especially in an open dialogue environment, where it is difficult to fully grasp the implicit psychological state of users.
A human-computer dialogue method based on search enhancement is adopted to generate user target responses by performing example search, emotional cognition and multi-knowledge fusion of user discourse. The specific steps include retrieving the user's discourse examples and generating example pair representations; converting the user's dialogue context historical information into a high-dimensional vector representation, and performing emotional cognitive cognitive state representation through a preset common sense transformation model; using examples to perform multi-knowledge fusion and decoding processing to generate user's target response.
It significantly improves the user's dialogue experience, and the generated response is both empathetic and cognitive, which can better understand and respond to users' complex emotional needs.
Smart Images

Figure CN119938861A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a human-computer dialogue method and system based on retrieval enhancement. Background Art
[0002] As the field of dialogue systems continues to develop, Emotional Support Conversation (ESC) has attracted great attention in the dialogue system community. Effective ESC pays more attention to the interlocutor's emotions and helps them resolve negative emotions, which shows that the model regards the interlocutor as a help-seeker. In order to express empathy effectively, ESC needs to fully understand the help-seeker's experience and feelings from historical conversations, aiming to alleviate the interlocutor's negative emotional pressure and provide guidance to overcome the pressure. Integrating emotional support functions into dialogue systems makes them invaluable in various scenarios, significantly enhancing the user experience in areas such as consultation, mental health support, and customer service chat.
[0003] However, existing systems face challenges in generating contextually relevant and empathetic responses, especially in open dialogue environments. Traditional end-to-end generative models often have difficulty in generating contextually appropriate responses, especially in open dialogues. Although some methods have tried to use common sense knowledge to enhance the model's understanding of emotions, they still fall short in fully grasping the user's implicit mental state. Summary of the invention
[0004] The present invention provides a human-computer dialogue method and system based on retrieval enhancement, which solves the technical problem of how to improve the user dialogue experience.
[0005] A first aspect of the present invention provides a human-computer dialogue method based on retrieval enhancement, comprising:
[0006] Perform example retrieval on user utterances and generate example pair representations;
[0007] Convert the user's conversation context history information into a high-dimensional vector representation, and perform emotional cognition on the user's speech through a preset common sense conversion model to generate an emotional cognition state representation;
[0008] The example pair representation, the high-dimensional vector representation and the emotional cognitive state representation are used to perform multi-knowledge fusion and decoding processing to generate the user's target response.
[0009] Optionally, performing example retrieval on the user utterance to generate example pair representations includes:
[0010] Extract multiple question-answer paragraph pairs from the preset emotion support dialogue dataset and build a retrieval library;
[0011] Calculating a similarity score between the user utterance and each of the query-answer paragraph pairs in the search library, and determining a candidate response from the query-answer paragraph pairs according to the similarity score;
[0012] combining the vector representation with each of the candidate responses into example pairs;
[0013] An encoding operation is performed on the example pairs to generate example pair representations.
[0014] Optionally, calculating the similarity score between the user speech and each of the query-answer paragraph pairs in the search library, and determining a candidate response from the query-answer paragraph pairs according to the similarity score, comprises:
[0015] Performing encoding conversion on the user speech to obtain a vector representation;
[0016] Calculating the similarity score between the vector representation and each of the query-answer paragraph pairs in the retrieval library through a pre-trained dense channel retrieval model;
[0017] The query-answer paragraph pairs are sorted in descending order according to the similarity scores, and a preset number of the query-answer paragraph pairs in the top row are selected as candidate responses.
[0018] Optionally, the converting of the user's conversation context history information into a high-dimensional vector representation, and performing emotion recognition on the user's speech through a preset common sense conversion model to generate an emotion recognition state representation, includes:
[0019] Aggregate the user's conversation context history information to generate an aggregate sequence;
[0020] Performing feature encoding on the aggregated sequence to generate a high-dimensional vector representation;
[0021] The user speech is emotionally recognized through a preset common sense conversion model, and the emotional recognition state representation is generated by combining the high-dimensional vector representation.
[0022] Optionally, the performing emotion recognition on the user speech by using a preset common sense conversion model and combining the high-dimensional vector representation to generate an emotion recognition state representation includes:
[0023] Performing emotional cognition on the user's speech through a preset common sense conversion model to generate multiple cognitive states;
[0024] Performing a fusion operation on the plurality of cognitive states to generate a cognitive state sequence;
[0025] Performing encoding operation on the cognitive state sequence to generate cognitive state representation;
[0026] The cognitive state representation and the high-dimensional vector representation are interactively operated to generate an emotional cognitive state representation.
[0027] Optionally, the interactive operation of the cognitive state representation and the high-dimensional vector representation to generate the emotional cognitive state representation includes:
[0028] performing encapsulation operations on the cognitive state representation to generate a composite cognitive representation;
[0029] Performing an optimization operation on the composite cognitive representation and the high-dimensional vector representation to generate a deep cognitive representation;
[0030] Feature extraction and enhancement operations are performed on the deep cognitive representation to generate an emotional cognitive state representation.
[0031] Optionally, the adopting the example pair representation, the high-dimensional vector representation and the emotional cognitive state representation to perform multi-knowledge fusion and decoding processing to generate a user's target response includes:
[0032] Performing a double cross attention operation on the example pair representation, the high-dimensional vector representation, and the emotional cognitive state representation to generate an alignment feature;
[0033] Performing feature weighted aggregation on the alignment features to generate a composite hidden state vector;
[0034] Normalizing the composite hidden state vector;
[0035] A decoding operation is performed on the normalized composite hidden state vector to generate a target response of the user.
[0036] A second aspect of the present invention provides a human-computer dialogue system based on retrieval enhancement, comprising:
[0037] Dynamic example retriever, used to retrieve examples from user utterances and generate example pair representations;
[0038] A cognitive context understanding module is used to convert the user's conversation context history information into a high-dimensional vector representation, and to perform emotional cognition on the user's speech through a preset common sense conversion model to generate an emotional cognitive state representation;
[0039] A multi-source knowledge decoder is used to use the example pair representation, the high-dimensional vector representation and the emotional cognitive state representation to perform multi-knowledge fusion and decoding processing to generate a user's target response.
[0040] A third aspect of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the retrieval-enhanced human-computer dialogue method as described in any one of the above items.
[0041] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed, implements the retrieval-enhanced human-computer dialogue method as described in any one of the above items.
[0042] It can be seen from the above technical solutions that the present invention has the following advantages:
[0043] The present invention first introduces innovative example retrieval to select information-rich and personalized example pairs, and then identifies four cognitive relationships through a preset common sense conversion model to deepen the understanding of the context and implicit psychological state of the help-seeker. Finally, the support decoder integrates multi-source knowledge to ensure that the response generation is both empathetic and cognitively aware, improving the user's conversation experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0045] Figure 1 A flowchart of a human-computer dialogue method based on retrieval enhancement provided by an embodiment of the present invention;
[0046] Figure 2 The D 2 Schematic diagram of the overall architecture of RCU;
[0047] Figure 3 This is a schematic diagram of the dynamic example retriever architecture;
[0048] Figure 4 A structural block diagram of a human-computer dialogue system based on retrieval enhancement provided by an embodiment of the present invention;
[0049] Figure 5 A structural block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0050] The embodiments of the present invention provide a human-computer dialogue method and system based on retrieval enhancement, which are used to solve the technical problem of how to improve the user dialogue experience.
[0051] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0052] In view of the shortcomings of the above-mentioned background technologies, there is an urgent need for a method that can effectively combine dynamic example retrieval and cognitive understanding to improve the quality of emotion support dialogues.
[0053] Current technical means mainly rely on pre-trained language models, which, although they perform well in natural language processing tasks, lack specific optimization for emotional support dialogues. Existing ESC systems fail to fully consider the user's personalized background and contextual information when dealing with the complex and changing emotional needs of help seekers. In addition, the cognitive understanding of emotional dialogues is relatively weak, resulting in a lack of depth and pertinence in response generation. These problems limit the actual application effect of existing ESC systems, especially in key areas such as mental health support and psychological counseling.
[0054] In order to solve the above problems, the present invention proposes an innovative method - D 2 RCU (Dynamic Demonstration Retrieval and Cognitive-Aspect Situation Understanding) aims to significantly improve the quality and effectiveness of emotional support conversations by introducing dynamic example retrieval and cognitive understanding modules.
[0055] See also Figure 1 , Figure 1 A flowchart of the steps of a human-computer dialogue method based on retrieval enhancement provided by an embodiment of the present invention.
[0056] See also Figure 2 , Figure 2 The D 2 The overall architecture of RCU consists of three key components:
[0057] 1) Dynamic Example Retriever: Dynamically retrieve semantically aligned query-paragraph pairs based on the posts and profile of the current help seeker;
[0058] 2) Cognitive Aspects Context Understanding Module: Captures cognitive states through four key cognitive aspects (want, need, intention and influence);
[0059] 3) Multi-source knowledge decoder: combines the encoded example pairs and cognitive states with the dialogue history for response generation.
[0060] The details of the cognitive understanding process are highlighted by the orange box at the bottom. Given a user’s post, we first obtain prior knowledge through dynamic example selection and COMET common sense extraction, then understand the user’s feelings through cognition of the current situation, and finally generate a reply through the knowledge-aware decoder.
[0061] This paper introduces an innovative retrieval mechanism, selects information-rich and personalized example pairs, and designs a cognitive understanding module that uses four cognitive relationships (Effect, Intent, Need, Want) from the ATOMIC knowledge source (Atomic Knowledge Graph) to deepen the understanding of the help-seeker's context and implicit psychological state. Finally, the support decoder integrates multi-source knowledge to ensure that the response generation is both empathetic and cognitively aware.
[0062] A method and system for emotion-supported dialogue based on dynamic example retrieval and cognitive understanding were constructed. By selecting appropriate examples as context, the knowledge in the knowledge graph was integrated into the system and effectively integrated into the deep learning model calculation. A dynamic example selection mechanism based on role information was constructed. By introducing a novel role information utilization method, more information-rich and highly personalized candidate pairs were generated in the dynamic example selection process. A dedicated cognitive understanding module was introduced to explicitly simulate the cognitive understanding of user contexts, emphasizing the importance of cognitive awareness. Through a deep understanding of user contexts, the user's implicit psychological state, including intentions, needs, desires and possible effects, was captured. A response generation method based on retrieval enhancement and cognitive understanding was designed. Combining the advantages of dynamic example selection and cognitive understanding, it not only relies on traditional text matching technology, but also makes full use of cognitive relationships extracted from knowledge graphs to ensure that the generated responses are both empathetic and cognitively aware.
[0063] See also Figure 3 , Figure 3 The dynamic sample retriever architecture.
[0064] The retrieval component (top) leverages a pre-trained dense passage retrieval (DPR) model to identify relevant responses in the training set, while the exemplar component (bottom) compiles the most relevant user-system pairs based on the top 𝑠 results.
[0065] The present invention provides a human-computer dialogue method based on retrieval enhancement, comprising:
[0066] Step 101: perform example retrieval on user utterances and generate example pair representations.
[0067] In an embodiment of the present invention, speech data of user-system interactions are first collected from various channels, such as chat records, customer service conversations, social media comments, etc. These data should contain rich user expressions, covering different topics and situations, and at the same time collect personalized role information related to each user. The personalized role information of the user refers to various types of information related to the role played by the user in a specific situation that can reflect the user's unique characteristics, preferences, behavior patterns, and the collected data are annotated, and the user's speech is associated with the corresponding personalized role information; then, an example library containing a large number of example pairs is constructed, each example pair consisting of a user speech and a corresponding system response, and finally, according to the current user speech, the example pair most relevant to the user is dynamically selected from the example library, and the selected example pairs are appropriately adjusted and combined to generate the final personalized example pair representation.
[0068] Furthermore, step 101 may include the following sub-steps:
[0069] S11. Extract multiple question-answer paragraph pairs from the preset emotion support dialogue dataset and build a retrieval library.
[0070] The preset emotional support conversation dataset refers to the ESConv dataset (Emotional Support Conversation Dataset).
[0071] In the embodiment of the present invention, the strategy text and response text are extracted from the data set to form a retrieval library. Each data in the data set contains a query + a response + a strategy for generating a response (referred to as strategy). That is, the user's question and the system response in each round of dialogue are extracted from the preset ESConv data set, where the system response = strategy text + response text. Figure 3 The form of is: Passage: [strategy, response], forming a series of query-paragraph pairs (i.e., query-answer paragraph pairs), which constitute a sample set containing multiple strategies and responses, providing rich reference materials for subsequent retrieval. The above extracted query-paragraph pairs are organized into a retrieval library for quick access during dynamic example selection.
[0072] It should be noted that the user's personalized role information is used for dynamic example selection to generate personalized candidate example pairs, where the user's personalized role information refers to various types of information that can reflect the user's unique characteristics, preferences, behavior patterns, and roles played in specific situations. This information can describe the user from multiple dimensions, helping the dialogue system to understand user needs more accurately, thereby achieving dynamic example selection and generating personalized example pair representations.
[0073] Furthermore, when performing example retrieval on user utterances, appropriate query-paragraph pairs may be dynamically selected as examples in the retrieval library according to the user utterances.
[0074] S12. Calculate the similarity scores between the user's utterance and each query-answer paragraph pair in the search database, and determine candidate responses from the query-answer paragraph pairs based on the similarity scores.
[0075] In an embodiment of the present invention, a pre-trained dense passage retrieval (DPR) model is used to calculate the similarity score between the vector representation associated with the query user's utterance and the query-answer paragraph pairs in the retrieval library, and the top s most relevant paragraphs are selected as candidate responses.
[0076] Further, S12 may include the following sub-steps:
[0077] S121. Perform encoding conversion on the user speech to obtain a vector representation.
[0078] In the embodiment of the present invention, the user speech includes multiple user query questions. , first by querying the encoder Convert to vector representation .
[0079] S122. Calculate the similarity scores between the vector representation and each query-answer paragraph pair in the retrieval library through the pre-trained dense channel retrieval model.
[0080] In the embodiment of the present invention, the vector representation is calculated by the pre-trained dense channel retrieval model and the query-answer paragraph pairs in the retrieval library are Similarity score :
[0081]
[0082] In the formula, Represents a paragraph encoder.
[0083] S123, sorting the query-answer paragraph pairs in descending order according to the similarity scores, and selecting a preset number of query-answer paragraph pairs at the top as candidate responses.
[0084] In the embodiment of the present invention, the query-answer paragraph pairs found are sorted from high to low according to the obtained similarity scores. The higher the similarity score, the higher the ranking. The top s most relevant query-answer paragraph pairs are selected as candidate responses. , ensuring that the selected example pairs are highly relevant to the current user context.
[0085] S13. Combine the vector representation with each candidate response to form an example pair.
[0086] In this embodiment of the present invention, the selected candidate response The corresponding vector representation Combine into example pairs :
[0087]
[0088] In the formula, Indicates the identification of user input information. Indicates the system response information.
[0089] S14. Perform encoding operation on the example pair to generate example pair representation.
[0090] In the embodiment of the present invention, an encoder is used to convert each example pair Example pair representation converted to hidden state :
[0091]
[0092]
[0093] In the formula, Represents an example pair representation, represents the continuous sequence of example pairs concatenated, Indicates the splicing symbol. Represents the encoder, specifically Figure 3 The purple knowledge encoder, Represents the number of example pairs.
[0094] It should be noted that converting example pairs into hidden state example pair representations is used to mine deep semantic associations in conversation texts, capture semantic information, improve computational efficiency, and facilitate model learning and generalization and multimodal information fusion.
[0095] Step 102: Convert the user's conversation context history information into a high-dimensional vector representation, and perform emotion recognition on the user's speech through a preset common sense conversion model to generate an emotion recognition state representation.
[0096] In an embodiment of the present invention, the COMET model is used to understand the emotional cognitive level and realize the transition from emotion recognition to emotion cognition, specifically encoding the dialogue history and aggregating the interaction information of multiple rounds; using the COMET model to predict the state of the cognitive relationship and encode it; and gradually optimizing the cognitive state representation through the cognitive encoder, cognitive optimizer and cognitive selector to generate the emotional cognitive state representation.
[0097] Further, step 102 may include the following sub-steps:
[0098] S21. Aggregate the user's conversation context history information to generate an aggregate sequence.
[0099] In the embodiment of the present invention, the user's conversation context history information Aggregate to multiply an aggregate sequence:
[0100]
[0101] In the formula, represents an aggregate sequence, Indicates the total number of conversation context history information.
[0102] S22. Feature encode the aggregated sequence to generate a high-dimensional vector representation.
[0103] In an embodiment of the present invention, a dedicated encoder is used to convert the sequence into a high-dimensional vector representation:
[0104]
[0105] In the formula, Represents a high-dimensional vector representation.
[0106] S23. Emotional recognition of user speech is performed through a preset common sense conversion model, and combined with a high-dimensional vector representation, an emotional cognitive state representation is generated.
[0107] The preset common sense transformation model refers to the pre-trained COMET model. Commonsense Transformer is a language model based on the Transformer architecture. It can utilize a large amount of common sense knowledge when processing natural language tasks, and generate text content that is more in line with common sense through learning and reasoning.
[0108] In this embodiment of the present invention, in order to capture the cognitive state, the pre-trained COMET model is used to predict the cognitive state of each cognitive relationship (Effect, Intent, Need, Want), and each aspect is then connected to the current user utterance. The mark indicates .
[0109] Further, S23 may include the following sub-steps:
[0110] S231. Emotional recognition of user speech is performed through a preset common sense conversion model to generate multiple cognitive states.
[0111] In the embodiment of the present invention, for each cognitive relationship, COMET is used to predict the corresponding cognitive state .
[0112] S232: Perform a fusion operation on multiple cognitive states to generate a cognitive state sequence.
[0113] In an embodiment of the present invention, multiple predicted cognitive states are combined into a cognitive state sequence:
[0114]
[0115] S233. Perform encoding operation on the cognitive state sequence to generate cognitive state representation.
[0116] In the embodiment of the present invention, the cognitive state sequence is encoded to obtain the cognitive state representation:
[0117]
[0118] In the formula, represents the cognitive state representation, where , represents the maximum sequence length of cognitive states, Represents the dimension of the hidden layer.
[0119] S234. Interact the cognitive state representation with the high-dimensional vector representation to generate an emotional cognitive state representation.
[0120] The cognitive situation understanding module includes a cognitive understander, which consists of three components: a cognitive encoder, a cognitive optimizer, and a cognitive selector.
[0121] The cognitive state representation obtained from the knowledge encoder is:
[0122]
[0123] Further, S234 may include the following sub-steps:
[0124] S2341. Encapsulate the cognitive state representation to generate a composite cognitive representation.
[0125] In the embodiment of the present invention, the cognitive encoder is used to encapsulate the cognitive state representation to generate a composite cognitive representation, wherein the first user utterance is preceded by The tag is used to represent the semantics of the whole sentence, so that it can represent the semantics of the whole sentence after multiple encodings. The hidden state cognitive state representation corresponding to the label is encapsulated into a composite cognitive representation:
[0126]
[0127]
[0128] In the formula, represents a complex cognitive representation, represents cognitive encoder, Represents the encapsulated information after being encapsulated by the cognitive encoder. , Determined by longer example pairs or cognitive sequences, represents the embedding space, Representation Extraction The hidden state corresponding to the label is the cognitive state representation.
[0129] It should be noted that in natural language processing, The tag is a special tag that is usually used to represent the semantic information of the entire sentence. When you add it before the first sentence After the tag is added, the cognitive encoder will use this tag as a special input when processing the sentence. For example, for the sentence "The weather is really nice today", adding After marking, it becomes " What a nice day today." The cognitive encoder processes this When marking a sentence, The token is encoded together with other tokens in the sentence (such as "today", "weather", etc.). The cognitive encoder encodes the input sentence (including During each encoding process, the encoder updates the hidden state of each tag based on the semantic relationship between the tags, context information, etc. The hidden state corresponding to the tag encapsulates a composite cognitive representation. A composite cognitive representation is a representation that integrates multiple cognitive information. It not only contains the semantic information of a single tag, but also integrates the overall structure of the sentence, contextual information, etc.
[0130] S2342. Optimize the composite cognitive representation and the high-dimensional vector representation to generate a deep cognitive representation.
[0131] In this embodiment of the present invention, in order to enrich the conversation context through cognitive insights, at the token level and This rich representation is optimized by the cognitive optimizer to generate cognitively enhanced context embeddings, which optimizes the composite cognitive representation with the high-dimensional vector representation to generate a deep cognitive representation:
[0132]
[0133]
[0134] In the formula, represents deep cognitive representation, represents the cognitive optimizer, Represents the fused information after the composite cognitive representation and the high-dimensional vector representation are combined.
[0135] S2343. Perform feature extraction and enhancement operations on the deep cognitive representation to generate an emotional cognitive state representation.
[0136] In the embodiment of the present invention, the cognitive selector is used to perform feature extraction and enhancement operations on the deep cognitive representation to generate the emotional cognitive state representation:
[0137] Use sigmoid function to calculate The importance score of each token in is then used to weight the tokens, emphasizing the salient features in the cognitive optimization context. Subsequently, an MLP with ReLU activation abstracts these weighted representations as the emotional cognitive state representation:
[0138]
[0139]
[0140] In the formula, Represents the importance score of each cognitive state calculated by the Sigmoid function, represents the sigmoid function, represents the emotional cognitive state representation after the hidden layer of the cognitive selector, Represents a cognitive selector, specifically a multilayer perceptron.
[0141] Through the process of step S234, the model has a deeper, more abstract, and more meaningful understanding of the cognitive state underlying the help seeker's narrative, thereby being able to generate responses that are not only empathetic but also cognitively consistent with the needs and desires expressed by the help seeker.
[0142] Step 103: Use examples to perform multi-knowledge fusion and decoding processing on the representation, high-dimensional vector representation and emotional cognitive state representation to generate the user's target response.
[0143] In an embodiment of the present invention, multi-source knowledge including example pairs and emotional cognitive understanding is fused to enhance the quality of response generation. In order to make full use of two different knowledge sources - retrieved example pair representations and emotional cognitive state representations - a multi-knowledge fusion decoder is introduced.
[0144] This component ensures that contextual knowledge is coherently integrated into the response generation process, which is critical for designing supportive conversations that are both relevant and cognitively consistent. It consists of the following parts:
[0145] A dual cross-attention operation to align the encoded example pairs and cognitive states with the dialogue history, respectively;
[0146] The features are weighted and aggregated to form a composite hidden state vector;
[0147] Standardization processing to enhance the generalization ability of the model;
[0148] Response generation,generates the target response word by word based on the fused knowledge-enriched hidden states.
[0149] Further, step 103 may include the following sub-steps:
[0150] S31. Perform double cross attention operations on example pair representation, high-dimensional vector representation, and emotional cognitive state representation to generate aligned features.
[0151] In an embodiment of the present invention, the decoder mechanism starts with a double cross-attention operation, which assimilates the encoded example pairs and cognitive states with the dialogue history respectively.
[0152] This process establishes a bidirectional information flow and aligns the knowledge representation with the conversation context.
[0153] Compute history context ( for high-dimensional vector representation) and encoding knowledge sources ( For example pair representation, The affinity score between the emotional cognitive state representation) is used to derive the updated context representation:
[0154]
[0155]
[0156]
[0157]
[0158] In the formula, Represents the Softmax score of the dialogue history and example pair representation, Softmax score representing the dialogue history and emotional cognitive state, represents context-dependent example pairs, represents the context-dependent emotional cognitive state representation, Represents the normalization layer, which is used for alignment.
[0159] Similarly, to align the historical context with the two knowledge entities, a similar computation is performed to derive context-dependent example pair representations and emotional cognitive state representation .
[0160] Alignment features include , , and .
[0161] S32. Perform feature weighted aggregation on the aligned features to generate a composite hidden state vector.
[0162] In the embodiment of the present invention, in order to balance and integrate these alignment features, a weighted aggregation strategy is adopted to obtain a composite hidden state vector as the input of the decoder:
[0163]
[0164]
[0165] In the formula, represents the composite hidden state vector, Represents adaptive parameters, which will be based on the following Adaptive optimization, represents the weight parameter, represents the total number of weight parameters, , the weight coefficient of the adaptive parameter The initialization is the same, but can be optimized during training to reflect the saliency of each feature.
[0166] S33. Normalize the composite hidden state vector.
[0167] In this embodiment of the present invention, in order to promote the consistency of feature dimensions and enhance model generalization, layer normalization is used to Each feature of is standardized:
[0168]
[0169] S34. Perform a decoding operation on the normalized composite hidden state vector to generate a target response of the user.
[0170] In the embodiment of the present invention, a decoder is used to perform a decoding operation on the composite hidden state vector after the normalization process to generate a target response of the user. The generation of The sequence of , depends on the fused knowledge-rich hidden state:
[0171]
[0172] In the formula, represents the target response of the decoder output, represents the embedding of the responses generated so far, Represents the token generated up to the current time step, specifically the token that the Decoder has been generating in response. This symbol is the embedding of all tokens before the current token. Represents a decoder.
[0173] Furthermore, the training target of the decoder is defined by the negative log-likelihood of the correct response. By continuously reducing the negative log-likelihood, the ability of the decoder is trained. The negative log-likelihood is:
[0174]
[0175] In the formula, represents the negative log-likelihood, represents the number of samples, Represents the time step.
[0176] In summary, the present invention provides an emotion support dialogue method and system based on dynamic example retrieval and cognitive understanding. First, character information is used to perform efficient dynamic example selection, generate personalized candidate pairs, and the similarity between the query and the paragraph is calculated through the dense passage retrieval (DPR) model to select the most relevant example pairs.
[0177] Secondly, the COMET model is introduced to predict and encode the cognitive relationship state (Effect, Intent, Need, Want) of the user context to capture the user's implicit psychological state.
[0178] Then, a double cross-attention mechanism is used to fuse multi-source knowledge, including example pairs, cognitive states, and dialogue history, to form a composite hidden state vector, and a weighted aggregation strategy is used to integrate information from different sources. Finally, combined with normalization processing and decoder application, the target response is generated word by word based on the fused knowledge-enriched hidden state, ensuring that the generated response is both empathetic and cognitively aware.
[0179] This invention solves the shortcomings of the emotional support dialogue system in understanding and responding to the complex emotional needs of users to a great extent, significantly improves the system's empathy and situational understanding capabilities, and especially performs well in generating high-quality, personalized support responses. In addition, this invention also has certain reference significance for other tasks involving natural language processing that require understanding and generating compassionate and cognitively deep text.
[0180] See also Figure 4 , Figure 4 A structural block diagram of a human-computer dialogue system based on retrieval enhancement provided by an embodiment of the present invention.
[0181] The present invention provides a human-computer dialogue system based on retrieval enhancement, comprising:
[0182] Dynamic example retriever, used to retrieve examples from user utterances and generate example pair representations;
[0183] The cognitive context understanding module is used to convert the user's conversation context history information into a high-dimensional vector representation, and to perform emotional cognition on the user's speech through a preset common sense conversion model to generate an emotional cognitive state representation;
[0184] The multi-source knowledge decoder is used to perform multi-knowledge fusion and decoding processing using example pair representation, high-dimensional vector representation and emotional cognitive state representation to generate the user's target response.
[0185] See also Figure 5 , Figure 5 A structural block diagram of a computer device provided in an embodiment of the present invention.
[0186] An electronic device according to an embodiment of the present invention includes: a memory 301 and a processor 302, wherein the memory 302 stores a computer program; when the computer program is executed by the processor 302, the processor 302 executes the human-computer dialogue method based on retrieval enhancement as in any of the above embodiments.
[0187] The memory 301 may be an electronic memory such as a flash memory, an EEPROM (electrically erasable programmable read-only memory), an EPROM, a hard disk or a ROM. The memory 301 has a storage space 303 for a program code 313 for executing any method step in the above method. For example, the storage space 303 for the program code may include individual program codes 313 for implementing the various steps in the above method, respectively. These program codes may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards or floppy disks. The program code may be compressed, for example, in an appropriate form. When these codes are run by a computing processing device, the computing processing device performs the various steps in the method described above. These program codes may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards or floppy disks. The program code may be compressed, for example, in an appropriate form. When these codes are executed by a computing and processing device, the computing and processing device is caused to execute the various steps of the above-described human-computer dialogue method based on retrieval enhancement.
[0188] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the human-computer dialogue method based on retrieval enhancement as described in any of the above embodiments is implemented.
[0189] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0190] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0191] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0192] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0193] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0194] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A human-computer dialogue method based on retrieval enhancement, characterized in that: include: Perform example retrieval on user utterances and generate example pair representations; Convert the user's conversation context history information into a high-dimensional vector representation, and perform emotional cognition on the user's speech through a preset common sense conversion model to generate an emotional cognition state representation; The example pair representation, the high-dimensional vector representation and the emotional cognitive state representation are used to perform multi-knowledge fusion and decoding processing to generate the user's target response.
2. The human-computer dialogue method based on retrieval enhancement according to claim 1 is characterized in that: The step of retrieving examples from user utterances and generating example pair representations includes: Extract multiple question-answer paragraph pairs from the preset emotion support dialogue dataset and build a retrieval library; Calculating a similarity score between the user utterance and each of the query-answer paragraph pairs in the search library, and determining a candidate response from the query-answer paragraph pairs according to the similarity score; combining the vector representation with each of the candidate responses into example pairs; An encoding operation is performed on the example pairs to generate example pair representations.
3. The human-computer dialogue method based on retrieval enhancement according to claim 2 is characterized in that: The calculating the similarity score between the user speech and each of the query-answer paragraph pairs in the search library, and determining a candidate response from the query-answer paragraph pairs according to the similarity score, comprises: Performing encoding conversion on the user speech to obtain a vector representation; Calculating the similarity score between the vector representation and each of the query-answer paragraph pairs in the retrieval library through a pre-trained dense channel retrieval model; The query-answer paragraph pairs are sorted in descending order according to the similarity scores, and a preset number of the query-answer paragraph pairs in the top row are selected as candidate responses.
4. The human-computer dialogue method based on retrieval enhancement according to claim 1 is characterized in that: The method converts the user's conversation context history information into a high-dimensional vector representation, and performs emotion recognition on the user's speech through a preset common sense conversion model to generate an emotion recognition state representation, including: Aggregate the user's conversation context history information to generate an aggregate sequence; Performing feature encoding on the aggregated sequence to generate a high-dimensional vector representation; The user speech is emotionally recognized through a preset common sense conversion model, and the emotional recognition state representation is generated by combining the high-dimensional vector representation.
5. The human-computer dialogue method based on retrieval enhancement according to claim 4 is characterized in that: The method of performing emotion recognition on the user speech by using a preset common sense conversion model and combining the high-dimensional vector representation to generate an emotion recognition state representation includes: Performing emotional cognition on the user's speech through a preset common sense conversion model to generate multiple cognitive states; Performing a fusion operation on the plurality of cognitive states to generate a cognitive state sequence; Performing encoding operation on the cognitive state sequence to generate cognitive state representation; The cognitive state representation and the high-dimensional vector representation are interactively operated to generate an emotional cognitive state representation.
6. The human-computer dialogue method based on retrieval enhancement according to claim 5 is characterized in that: The interactive operation of the cognitive state representation and the high-dimensional vector representation to generate the emotional cognitive state representation includes: performing encapsulation operations on the cognitive state representation to generate a composite cognitive representation; Performing an optimization operation on the composite cognitive representation and the high-dimensional vector representation to generate a deep cognitive representation; Feature extraction and enhancement operations are performed on the deep cognitive representation to generate an emotional cognitive state representation.
7. The human-computer dialogue method based on retrieval enhancement according to claim 1 is characterized in that: The step of using the example pair representation, the high-dimensional vector representation, and the emotional cognitive state representation to perform multi-knowledge fusion and decoding processing to generate a user's target response includes: Performing a double cross attention operation on the example pair representation, the high-dimensional vector representation, and the emotional cognitive state representation to generate an alignment feature; Performing feature weighted aggregation on the alignment features to generate a composite hidden state vector; Normalizing the composite hidden state vector; A decoding operation is performed on the normalized composite hidden state vector to generate a target response of the user.
8. A human-computer dialogue system based on retrieval enhancement, based on the human-computer dialogue method based on retrieval enhancement according to any one of claims 1 to 7, characterized in that: include: Dynamic example retriever, used to retrieve examples from user utterances and generate example pair representations; A cognitive context understanding module is used to convert the user's conversation context history information into a high-dimensional vector representation, and to perform emotional cognition on the user's speech through a preset common sense conversion model to generate an emotional cognitive state representation; A multi-source knowledge decoder is used to use the example pair representation, the high-dimensional vector representation and the emotional cognitive state representation to perform multi-knowledge fusion and decoding processing to generate a user's target response.
9. An electronic device, characterized in that: It comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the retrieval-enhanced human-computer dialogue method as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the human-computer dialogue method based on retrieval enhancement as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Emotion-sharing dialogue generation method based on context awareness and emotion reasoning
CN117892736A
Cited By
Double-granularity text retrieval method based on atomic sentences
CN120821825A