Emotional support dialogue generation method integrating sensory reasoning and mental memory enhancement

Through the combination of the emotional mind theory framework and large language model, the response mode is dynamically selected, which solves the problem of insufficient emotional understanding and response ability of the emotional support dialogue system, and achieves efficient and personalized emotional support dialogue generation, improving user experience and mental health.

CN120371949APending Publication Date: 2025-07-25NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510395438.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing emotional support dialogue system has insufficient emotional understanding and response capabilities, cannot accurately capture the user's emotional state, lacks empathy in response, generates replies that are stiff and have poor dialogue coherence, and cannot meet the actual needs of users.

Method used

An emotionally supported dialogue generation method that integrates sensory rational reasoning and mental memory enhancement, extracts personal cues and factual cues through the emotional mind theory framework, dynamically selects an emotionally-dominated or rationally-dominated response model, combines a Sentence-BERT encoder and a large language model to generate targeted responses, and ensures semantic coherence through an autoregressive decoder.

Benefits of technology

It achieves high-quality, efficient and personalized generation of emotional support dialogues, improves users' emotional experience and mental health, ensures the fit and emotional coordination of responses, and avoids repeated suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371949A_ABST
    Figure CN120371949A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to an emotional support dialogue generation method integrating sensory reasoning and mental memory enhancement, which comprises the following steps: performing structured psychological state extraction on dialogue content input by a user, and generating personal clues and fact clues by utilizing an emotional mental theory framework; dynamically selecting an emotional dominant or rational dominant response mode through a personal fact classifier, and determining a generated emotional state; encoding the historical memory and the emotional state by adopting a Sensory-BERT encoder, calculating semantic similarity between the historical memory and the emotional state, and dynamically filtering redundant information through a forgetting elimination mechanism based on the semantic similarity to remove outdated memory so as to obtain updated memory; inputting the emotional state and the updated memory into a large language model to generate a targeted response; semantic coherence is guaranteed through a mask self-attention mechanism, and repeated suggestions are avoided in multiple rounds of dialogues. According to the method, high-quality, high-efficiency and high-individuation generation of the emotional support dialogue is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of artificial intelligence, and particularly to an emotional support dialogue generation method that integrates perceptual reasoning and enhanced mental memory. Background Art

[0002] With the rapid development of the digital age, emotional support dialogue has become an important application in the field of mental health. When facing emotional stress and psychological distress, users often obtain emotional comfort and support by chatting with virtual assistants or chatbots. Research shows that deeply understanding the user's mental state and providing personalized emotional support can significantly improve user satisfaction and mental health levels. However, traditional emotional dialogue support technologies still have significant technical bottlenecks in automatically generating dialogue responses that fit the user's needs, are empathetic, and have emotional coordination. There is an urgent need for an innovative solution that takes into account emotional understanding, cognitive logic, and real-time performance.

[0003] Although the mainstream emotional support dialogue systems on the current market have basic dialogue generation capabilities, their emotional understanding and response capabilities mainly rely on preset rules or simple keyword matching. The "In-depth Research and Development Prospect Forecast Report on the Chinese AI Emotional Companion Market from 2025 to 2031" points out that the emotional support functions of current emotional support dialogue systems cannot meet the requirements. Specifically, they have insufficient emotional understanding ability and cannot accurately capture the user's emotional state; the responses lack empathy, and the generated replies are rigid and lack emotional resonance; the dialogue coherence is poor, and they cannot dynamically adjust the response strategy according to the progress of the dialogue. Such problems seriously restrict the practicality and user experience of emotional support dialogue systems and cannot meet the actual needs of users. There is an urgent need for a technological breakthrough.

[0004] Currently, the emotional support dialogue technologies proposed in academia and industry mainly include two categories: one is the emotional support method based on a single factor, whose technical principle is to generate responses by analyzing the user's emotions or background information, but it has defects. This technology ignores the complexity and dynamics of the user's mental state, resulting in responses lacking depth and personalization. The other is the emotional support method based on multiple factors, whose technical principle is to comprehensively consider multiple factors such as the user's emotions, background, and cognition to generate responses, but it also has defects. This technology lacks the ability to deeply reason about the user's mental state and dynamically adjust, resulting in responses that may not match the user's needs. In addition, current emotional support dialogue technologies also have other technical bottlenecks, such as insufficient emotional reasoning ability and inability to accurately predict the user's mental state; the memory update mechanism is imperfect and cannot dynamically adjust the memory content according to the progress of the dialogue; the real-time performance is insufficient, and the response generation takes a long time, affecting the fluency of the dialogue. Summary of the Invention

[0005] To solve the above technical problems, an embodiment of the present application proposes an emotional support dialogue generation method that integrates perceptual reasoning and mental memory enhancement, aiming to deeply integrate the theory of mind with large language models, solve the contradiction between emotional understanding and response generation, and thus achieve the high-quality, high-efficiency, and highly personalized generation of emotional support dialogues, significantly improving the user's emotional experience and mental health level, and having broad commercial prospects and social value.

[0006] In a first aspect, an embodiment of the present application proposes an emotional support dialogue generation method that integrates perceptual reasoning and mental memory enhancement. The method includes: extracting the structured mental state from the dialogue content input by the user, and generating personal clues and factual clues using the emotional theory of mind framework; based on the personal clues and factual clues, dynamically selecting an emotion-dominated or reason-dominated response mode through a personal fact classifier to determine the generated emotional state; encoding the historical memory and emotional state using a Sentence-BERT encoder, calculating the semantic similarity between the two, and dynamically filtering redundant information through a forgetting elimination mechanism to remove outdated memories based on the semantic similarity to obtain the updated memory; inputting the emotional state and the updated memory into a large language model, and generating a targeted response by the large language model in combination with the emotional state and the updated memory; decoding the targeted response using an autoregressive decoder, and ensuring semantic coherence through a masked self-attention mechanism to avoid repeating suggestions in multi-turn dialogues.

[0007] Optionally, extracting the structured mental state from the dialogue content input by the user and generating personal clues and factual clues using the emotional theory of mind framework includes:

[0008] Obtaining the dialogue content U input by the user t , where the subscript t indicates that the current dialogue is the t-th round of dialogue;

[0009] Inputting U t into the emotional theory of mind framework, and the emotional theory of mind framework performs structured mental state extraction on U t to generate personal clues and factual clues

[0010] The extraction process of the emotional theory of mind framework is divided into two stages. In the first stage, the large language model is defined as a psychotherapist through role-based prompts to constrain the output format and ensure the structuring of the clues. In the second stage, semantic similarity examples are dynamically matched based on retrieval enhancement technology to improve context relevance;

[0011] The extraction process of the emotional theory of mind framework is represented by the formula:

[0012]

[0013] Among them, φ EToM (·) represents the emotional theory of mind framework.

[0014] Optionally, based on personal cues and factual cues, a sentiment-dominated or reason-dominated response pattern is dynamically selected through a personal fact classifier, which is expressed by the formula:

[0015]

[0016] Among them, PFC(·) represents the personal fact classifier, Dis t represents the decision result output by the personal fact classifier, and Dis t takes values of Personal, Factual, or Both. Personal means giving priority to personal cues Factual means giving priority to factual cues Both means considering both personal cues and factual cues

[0017] The generated emotional state O t is determined by Dis t , and O t is expressed by the formula:

[0018]

[0019] Among them, when the user's mood fluctuates violently, the personal fact classifier preferentially selects personal restriction to generate the emotional state O t to avoid interference from rational information.

[0020] Optionally, the Sentence-BERT encoder is used to encode the historical memory and the emotional state, and the semantic similarity between the two is calculated, which is achieved through the following formula:

[0021]

[0022] Among them, m t-1 represents the historical memory, SBERT(·) represents the Sentence-BERT encoder, ‖·‖ represents taking the norm, and Sim(m t-1 , O t ) represents the calculated semantic similarity.

[0023] Optionally, based on the semantic similarity, redundant information is dynamically filtered through a forgetting elimination mechanism to remove outdated memories, and the updated memory is obtained, which is achieved through the following formula:

[0024] m t = m t-1-γORM{η t [Sim(m t-1 ,O t )],(1 - λ1)};

[0025] Among them, η t (·) represents a threshold function, γ is an adjustment coefficient adaptively adjusted according to the emotional intensity, λ1 is the first preset threshold, and ORM(·) represents a forgetting elimination mechanism.

[0026] Optionally, the emotional state and the updated memory are input into a large language model, and the large language model generates a targeted response by combining the emotional state and the updated memory, which is achieved through the following formula:

[0027] R t = F gen (C t ,O t ,m t );

[0028] C t = {U1,R1,U2,R2,…,U t-1 ,R t-1 ,U t};

[0029] Among them, F gen (·) represents the large language model, C t represents the historical dialogue context, and R t represents the targeted response to U t .

[0030] Optionally, an autoregressive decoder is used to decode the targeted response, and the masked self-attention mechanism is used to ensure semantic coherence to avoid repeated suggestions in multi-round conversations, which is achieved through the following formula:

[0031]

[0032] M t = η t [Sim(M t+1 ,O t ),λ2]

[0033] Among them, M t represents the EToM memory at the end of the t-th round of conversation, η t (·) represents a threshold function, λ2 is the second preset threshold, and Sim(M t+1 ,O t ) represents the similarity between M t+1 and O tThe similarity between, θ represents the parameters of the large language model before update, U represents the user input dialogue content, R represents the targeted response, D is the training dataset, p(R t |U,R,M <t ; θ) represents the probability that the large language model generates a response given the U and R of this round, and the M of the previous rounds. E (U,R,M)~D (·) represents taking the expectation, and θ * represents the parameters of the updated large language model.

[0034] An emotional support dialogue generation method that integrates sensible reasoning and mental memory enhancement proposed in this application first extracts the structured psychological state from the user input dialogue content, and generates personal clues and factual clues using the emotional mental theory framework. Then, through the personal fact classifier, it dynamically selects the emotional or rational dominant response mode to determine the generated emotional state. Next, it encodes the historical memory and emotional state, calculates the semantic similarity to provide a basis for subsequent memory update. Based on the semantic similarity, it dynamically filters redundant information through the forgetting elimination mechanism to remove outdated memories and ensure the timeliness of dialogue information. Then, it inputs the emotional state and the updated memory into the large language model to generate a targeted response. Finally, it uses an autoregressive decoder for decoding, and ensures the semantic coherence of the response through the masked self-attention mechanism. The global memory is updated after each round of dialogue to ensure long-term consistency and retain key support strategies. In emotional support dialogues, the psychological state and emotional needs of users are complex and changeable. By extracting personal clues and factual clues through the emotional mental theory framework, it can accurately reason about the psychological state of users, thereby generating more empathetic and personalized responses. The priority between emotional information and rational information needs to be dynamically adjusted. By selecting the emotional dominant or rational dominant response mode through the personal fact classifier, it ensures the fitness and emotional coordination of the response. In addition, there may be redundant or outdated information in historical memories. By dynamically filtering and updating memories through the forgetting elimination mechanism, it ensures the accuracy and relevance of the response. It should also be noted that by combining the user's current emotions and historical facts, generating targeted responses, and maintaining emotional consistency in multi-round dialogues, avoiding repeated suggestions, it effectively improves the coherence of the dialogue and the user experience.

[0035] Second aspect, an embodiment of the present application proposes an emotional support dialogue generation system that integrates perceptual reasoning and mental memory enhancement. The system includes: an acquisition module for acquiring the dialogue content input by the user; a structured mental state extraction module for extracting the structured mental state from the dialogue content input by the user and generating personal clues and factual clues using the emotional mental theory framework; a dynamic linear classification module for dynamically selecting an emotion-dominated or reason-dominated response mode based on the personal clues and factual clues through a personal fact classifier and determining the generated emotional state; an outdated filtering module for encoding the historical memory and emotional state using a Sentence-BERT encoder, calculating the semantic similarity between the two, and dynamically filtering redundant information through a forgetting elimination mechanism based on the semantic similarity to remove outdated memories and obtain updated memories; a targeted response generation module for inputting the emotional state and updated memories into a large language model, and the large language model generates a targeted response by combining the emotional state and updated memories; a response decoding module for decoding the targeted response using an autoregressive decoder and ensuring semantic coherence through a masked self-attention mechanism to avoid repeated suggestions in multi-round conversations.

[0036] Third aspect, an embodiment of the present application proposes an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute an emotional support dialogue generation method that integrates perceptual reasoning and mental memory enhancement as described above.

[0037] Fourth aspect, an embodiment of the present application proposes a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it can implement an emotional support dialogue generation method that integrates perceptual reasoning and mental memory enhancement as described above.

[0038] It can be understood that the beneficial effects of the above second aspect to fourth aspect can refer to the relevant descriptions in the above first aspect, and will not be repeated here. Description of the Drawings

[0039] To more clearly illustrate the embodiments of the present application or the technical solutions in the related art, the following will briefly introduce the drawings required for the description of the embodiments of the present application or the related technology. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. The drawings described here are only used to explain the present application and are not used to limit the present application.

[0040] Figure 1 It is a flowchart of an emotional support dialogue generation method that integrates perceptual reasoning and mental memory enhancement provided in an embodiment of the present application;

[0041] Figure 2 It is a visualization diagram of an emotional support dialogue generation method that integrates perceptual reasoning and mental memory enhancement provided in an embodiment of the present application;

[0042] Figure 3 It is a schematic diagram of the overall factors of EToM provided in an embodiment of the present application;

[0043] Figure 4 It is a schematic diagram of psychological reasoning in a multi-turn dialogue provided in an embodiment of the present application;

[0044] Figure 5 It is a schematic diagram of the overall components of the EToM generator prompt provided in an embodiment of the present application;

[0045] Figure 6 It is a schematic diagram of an update strategy provided in an embodiment of the present application;

[0046] Figure 7 It is a comparison chart of the outputs of MIA and other models provided in an embodiment of the present application;

[0047] Figure 8 It is a schematic structural diagram of an emotional support dialogue generation system that integrates perceptual reasoning and mental memory enhancement provided in another embodiment of the present application;

[0048] Figure 9 It is a schematic structural diagram of an electronic device provided in another embodiment of the present application. Detailed implementation manners

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will introduce and elaborate on each embodiment of the present application in conjunction with the accompanying drawings. Those of ordinary skill in the art can understand that in each embodiment of the present application, many technical details are proposed to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented. The following division of each embodiment is for convenience of description and should not constitute any limitation to the specific implementation manner of the present application. Each embodiment can be combined and cross-referenced with each other on the premise of not being contradictory.

[0050] An embodiment of the present application proposes an emotional support dialogue generation method that integrates perceptual reasoning and mental memory enhancement, which is applied to an electronic device. Here, the electronic device can be a terminal or a server. In this embodiment and the following embodiments, the server is taken as an example for illustration. Next, the implementation details of the emotional support dialogue generation method that integrates perceptual reasoning and mental memory enhancement proposed in this embodiment will be specifically described. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this solution.

[0051] The specific process of the emotional support dialogue generation method that integrates perceptual reasoning and mental memory enhancement proposed in this embodiment can be as Figure 1 shown, and its visualized representation is as Figure 2 shown. This method includes the following steps:

[0052] Step 101: Extract the structured mental state from the dialogue content input by the user, and generate personal clues and factual clues using the Emotional Theory of Mind (EToM) framework.

[0053] In specific implementation, the server obtains the dialogue content input by the user in real time. After obtaining the dialogue content input by the user, it first extracts the structured mental state from the dialogue content input by the user, and generates personal clues and factual clues using the Emotional Theory of Mind (EToM) framework. The Emotional Theory of Mind (EToM) framework can be abbreviated as the EToM framework.

[0054] In an example, denote the dialogue content input by the user as U t , and the subscript t at the lower right indicates that the current dialogue is the t-th round of dialogue. In addition to obtaining U t , the server also synchronously obtains the historical dialogue context C t , where C t contains the dialogue content input by the user and the targeted responses generated by the large language model in the previous t - 1 rounds. It is also possible to include U t in C t , that is, C t = {U1, R1, U2, R2, …, U t-1 , R t-1 , U t}. The server inputs U t into the EToM framework, and the EToM framework extracts the structured mental state from U t , thereby generating personal clues and factual clues The extraction process of the EToM framework is divided into two stages. In the first stage, the large language model is defined as a psychotherapist through role-based prompts, and the output format is constrained to ensure the structuring of the clues. In the second stage, based on the retrieval enhancement technology, semantically similar examples are dynamically matched to enhance the context relevance.

[0055] In one example, the extraction process of the EToM framework is expressed by the formula as follows:

[0056]

[0057] where φ EToM (·) represents the EToM framework.

[0058] It can be understood that in an emotional support conversation, the user's mental state and emotional needs are complex and changeable. By extracting personal clues and factual clues through the EToM framework, the user's mental state can be accurately inferred, so as to generate more empathetic and personalized responses.

[0059] In one example, the overall factors affecting the EToM framework can be as Figure 3 shown.

[0060] Step 102: Based on personal clues and factual clues, dynamically select an emotion-dominated or reason-dominated response mode through a personal fact classifier to determine the generated emotional state.

[0061] In a specific implementation, after the server extracts personal clues and factual clues, it can dynamically select an emotion-dominated or reason-dominated response mode through the personal fact classifier PFC to determine the generated emotional state.

[0062] In one example, the server dynamically selects an emotion-dominated or reason-dominated response mode based on personal clues and factual clues through a personal fact classifier, which can be expressed by the formula as follows:

[0063]

[0064] where PFC(·) represents the personal fact classifier, and Dis t represents the decision result output by the personal fact classifier. The value of Dis t is Personal, Factual, or Both. Personal means giving priority to personal clues Factual means giving priority to factual clues Both means considering both personal clues and factual clues

[0065] In one example, the generated emotional state O t is determined by Dis t and O t is expressed by the formula as follows:

[0066]

[0067] It should be noted that when the user's emotions fluctuate violently, the personal fact classifier preferentially selects personal contraction. Generate emotional state O t , and avoid interference from rational information.

[0068] In one example, psychological reasoning in a multi-turn conversation can be as Figure 4 shown. The prompt contains definitions of sensibility and rationality, descriptions of psychological techniques, and examples for different situations, ensuring that the classifier can make reasonable decisions based on the context.

[0069] It can be understood that the priority between sensible information and rational information needs to be dynamically adjusted. By using the personal fact classifier PFC to select a sensible-dominated or rational-dominated response mode, the appropriateness and emotional coordination of the response can be ensured.

[0070] Step 103: Use the Sentence-BERT encoder to encode the historical memory and emotional state, calculate the semantic similarity between the two, and based on the semantic similarity, dynamically filter redundant information through the forgetting elimination mechanism to remove outdated memories, obtaining the updated memory.

[0071] In a specific implementation, after obtaining the emotional state, the server can use the Sentence-BERT encoder to encode the historical memory and emotional state, calculate the semantic similarity between the two, and based on the semantic similarity, dynamically filter redundant information through the forgetting elimination mechanism to remove outdated memories, obtaining the updated memory.

[0072] In one example, when the server uses the Sentence-BERT encoder to encode the historical memory and emotional state and calculate the semantic similarity between the two, it is essentially calculating the cosine similarity between the two. This calculation process can be achieved through the following formula:

[0073]

[0074] where m t-1 represents the historical memory, SBERT(·) represents the Sentence-BERT encoder, ‖·‖ represents taking the norm, and Sim(m t-1 , O t ) represents the calculated semantic similarity.

[0075] In one example, the forgetting elimination mechanism is simply referred to as ORM. The server uses the ORM-based memory updater to update the memory to ensure that the information in the conversation remains relevant. ORM then selects whether to retain or discard outdated memory entries according to the similarity threshold.

[0076] In one example, based on semantic similarity, redundant information is dynamically filtered through a forgetting elimination mechanism to remove outdated memories, and the updated memories can be achieved through the following formula:

[0077] m t = m t-1 - γORM{η t [Sim(m t-1 , O t )], (1 - λ1)};

[0078] Among them, η t (·) represents a threshold function, γ is an adjustment coefficient adaptively adjusted according to the emotional intensity, λ1 is the first preset threshold, and ORM(·) represents the forgetting elimination mechanism.

[0079] It can be understood that there may be redundant or outdated information in the historical memories in the dialogue. By dynamically filtering and updating the memories through the forgetting elimination mechanism, the accuracy and relevance of the response can be ensured.

[0080] Step 104, input the emotional state and the updated memories into the large language model, and the large language model generates a targeted response by combining the emotional state and the updated memories.

[0081] In a specific implementation, after the server completes the update of the memories, it can input the emotional state and the updated memories into the large language model, and the large language model generates a targeted response for this round by combining the emotional state and the updated memories.

[0082] In one example, the large language model generates a targeted response by combining the emotional state and the updated memories, which can be achieved through the following formula:

[0083] R t = F gen (C t , O t , m t );

[0084] C t = {U1, R1, U2, R2, …, U t-1 , R t-1 , U t};

[0085] Among them, F gen (·) represents the large language model, C t represents the historical dialogue context, and R t represents the targeted response to U t .

[0086] In one example, the overall components of the EToM generator prompt are as Figure 5As shown, the large language model (LLM) plays a role in generating responses.

[0087] Step 105: Decode the targeted response using an autoregressive decoder, and ensure semantic coherence through the masked self-attention mechanism to avoid repeated suggestions in multi-turn conversations.

[0088] In a specific implementation, emotional support dialogue generation is a multi-turn process. Therefore, the server needs to decode the targeted response using an autoregressive decoder and ensure semantic coherence through the masked self-attention mechanism to avoid repeated suggestions in multi-turn conversations.

[0089] In one example, the progressive response decoding process includes model update and global memory update, and the update process can be implemented through the following formula:

[0090]

[0091] M t = η t [Sim(M t+1 , O t ), λ2]

[0092] where M t represents the EToM memory at the end of the t-th round of conversation, η t (·) represents the threshold function, λ2 is the second preset threshold, Sim(M t+1 , O t ) represents the similarity between M t+1 and O t , θ represents the parameters of the large language model before update, U represents the user input dialogue content, R represents the targeted response, D is the training dataset, p(R t |U, R, M <t ; θ) represents the probability that the large language model generates a response given this round of U and R, and the previous round of M, E (U,R,M)~D (·) represents taking the expectation, and θ * represents the parameters of the updated large language model.

[0093] An emotional support dialogue generation method that integrates perceptual reasoning and mental memory enhancement proposed in this embodiment first extracts the structured mental state from the dialogue content input by the user, and generates personal clues and factual clues using the emotional mental theory framework. Then, the personal fact classifier dynamically selects the emotional or rational dominant response mode to determine the generated emotional state. Next, the historical memory and emotional state are encoded, and the semantic similarity is calculated to provide a basis for subsequent memory update. Based on the semantic similarity, redundant information is dynamically filtered through the forgetting elimination mechanism to remove outdated memories, ensuring the timeliness of dialogue information. Subsequently, the emotional state and the updated memory are input into the large language model to generate targeted responses. Finally, the autoregressive decoder is used for decoding, and the semantic coherence of the response is guaranteed through the masked self-attention mechanism. The global memory is updated after each round of dialogue to ensure long-term consistency and retain key support strategies. In emotional support dialogues, the mental state and emotional needs of users are complex and variable. By extracting personal clues and factual clues through the emotional mental theory framework, the mental state of users can be accurately inferred, thus generating more empathetic and personalized responses. The priority between emotional information and rational information needs to be dynamically adjusted. By selecting the emotional dominant or rational dominant response mode through the personal fact classifier, the fittingness and emotional coordination of the response are ensured. In addition, there may be redundant or outdated information in the historical memory. By dynamically filtering and updating the memory through the forgetting elimination mechanism, the accuracy and relevance of the response are ensured. It should also be noted that by combining the user's current emotion and historical facts, targeted responses are generated, and emotional consistency is maintained in multiple rounds of dialogue, avoiding repeated suggestions, effectively improving the coherence of the dialogue and the user experience.

[0094] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step, or some steps can be decomposed into multiple steps. As long as the same logical relationship is included, it is within the protection scope of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but not changing the core design of its algorithm and process, are all within the protection scope of this application.

[0095] In one embodiment, the emotional support dialogue generation method that integrates perceptual reasoning and mental memory enhancement proposed in this application can be implemented in a model, which is called MIA, as Figure 7 shown. Compared with other models, the responses generated by MIA are more in line with user needs.

[0096] Another embodiment of this application proposes an emotional support dialogue generation system that integrates perceptual reasoning and mental memory enhancement. The details of the emotional support dialogue generation system that integrates perceptual reasoning and mental memory enhancement proposed in this embodiment will be specifically described below. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this example.

[0097] Figure 8 It is a schematic diagram of an emotional support dialogue generation system that integrates perceptual reasoning and enhanced mental memory proposed in this embodiment. The system includes: an acquisition module 201, a structured psychological state extraction module 202, a dynamic linear classification module 203, an obsolete filtering module 204, a targeted response generation module 205, and a response decoding module 206.

[0098] The acquisition module 201 is used to acquire the dialogue content input by the user.

[0099] The structured psychological state extraction module 202 is used to extract the structured psychological state from the dialogue content input by the user, and generate personal clues and factual clues using the emotional mental theory framework.

[0100] The dynamic linear classification module 203 is used to dynamically select an emotion-dominated or reason-dominated response mode based on personal clues and factual clues through a personal fact classifier, and determine the generated emotional state.

[0101] The obsolete filtering module 204 is used to encode the historical memory and emotional state using a Sentence-BERT encoder, calculate the semantic similarity between the two, and dynamically filter redundant information through a forgetting elimination mechanism based on the semantic similarity to remove obsolete memories, obtaining updated memories.

[0102] The targeted response generation module 205 is used to input the emotional state and updated memories into a large language model, and the large language model generates a targeted response by combining the emotional state and updated memories.

[0103] The response decoding module 206 is used to decode the targeted response using an autoregressive decoder, and ensure semantic coherence through a masked self-attention mechanism to avoid repeated suggestions in multi-turn conversations.

[0104] It is not difficult to find that this embodiment is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above method embodiment. The relevant technical details and technical effects mentioned in the above embodiments are still valid in this embodiment. To avoid repetition, they are not elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiments.

[0105] It is worth mentioning that each module and component involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovative part of this application, units that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.

[0106] Another embodiment of the present application provides an electronic device, whose specific structure may be as Figure 9 shown, including: at least one processor 301; and a memory 302 communicatively connected to the at least one processor 301; wherein, the memory 302 stores instructions executable by the at least one processor 301, and when the instructions are executed by the at least one processor 301, the at least one processor 601 is enabled to execute an emotional support dialogue generation method that combines perceptual reasoning and mental memory enhancement as described in the above method embodiments.

[0107] Among them, the memory and the processor are connected by a bus. The bus includes any number of interconnected buses and bridges, and the bus can connect various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a component or multiple components, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium.

[0108] The processor is responsible for managing the bus and general processing, and can also provide various functions, including but not limited to timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store data used by the processor when executing operations.

[0109] Another embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor, can implement an emotional support dialogue generation method that combines perceptual reasoning and mental memory enhancement as described in the above method embodiments.

[0110] That is, those skilled in the art can understand that all or part of the steps in the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium and includes several instructions for enabling a device (such as a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs, etc., which can store program codes.

[0111] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application. In actual applications, various changes can be made in form and details without departing from the spirit and scope of the present application. For those of ordinary skill in the technical field, several improvements and refinements can be made without departing from the principle of the present application, and these improvements and refinements are also regarded as the protection scope of the present application.

Claims

1. An emotional support dialogue generation method that integrates perceptual reasoning and mental memory enhancement, characterized in that, The method includes: Performing structured mental state extraction on the conversation content input by the user, and generating personal clues and factual clues using the emotional mind theory framework; Based on the personal clues and factual clues, dynamically select an emotion-dominated or reason-dominated response mode through a personal fact classifier to determine the generated emotional state; Use the Sentence-BERT encoder to encode the historical memory and emotional state, calculate the semantic similarity between the two, and based on the semantic similarity, dynamically filter redundant information through the forgetting elimination mechanism to remove outdated memories and obtain the updated memory; Input the emotional state and the updated memory into a large language model, and the large language model generates a targeted response by combining the emotional state and the updated memory; Use an autoregressive decoder to decode the targeted response, and ensure semantic coherence through a masked self-attention mechanism to avoid repeated suggestions in multi-turn conversations.

2. A method for generating an emotional support dialogue that integrates perceptual reasoning and enhanced mental memory according to claim 1, characterized in that Performing structured mental state extraction on the conversation content input by the user, and generating personal clues and factual clues using the emotional mind theory framework, including: Obtain the conversation content U input by the user t , where the subscript t in the lower right corner indicates that the current conversation is the t-th round of conversation; Input U t into the affective theory of mind framework, which extracts structured mental states from U t to generate personal cues and factual cues The extraction process of the emotional mind theory framework is divided into two stages. In the first stage, the large language model is defined as a psychotherapist through role-based prompts to constrain the output format and ensure the structuring of the clues. In the second stage, semantic similarity examples are dynamically matched based on retrieval enhancement technology to improve context relevance; The extraction process of the emotional mind theory framework is represented by the formula: Among them, φ EToM (·) represents the emotional theory of mind framework.

3. The emotional support dialogue generation method integrating perceptual reasoning and mental memory enhancement according to claim 2, characterized in that Based on the personal clues and factual clues, dynamically select an emotion-dominated or reason-dominated response mode through a personal fact classifier, which is represented by the formula: Among them, PFC(·) represents the Personal Fact Classifier, and Dis t represents the decision result output by the Personal Fact Classifier. The value of Dis t is Personal, Factual or Both. Personal means giving priority to personal clues Factual means giving priority to factual clues Both means considering both personal clues and factual clues Generated emotional state O t Determined by Dis t O t Expressed by the formula as: Among them, when the user's mood fluctuates violently, the personal fact classifier preferentially selects personal restriction Generate emotional state O t , avoiding interference from rational information.

4. A method for generating an emotional support dialogue that integrates perceptual reasoning and enhanced mental memory according to claim 3, characterized in that Use the Sentence-BERT encoder to encode the historical memory and emotional state, and calculate the semantic similarity between the two, which is achieved through the following formula: Among them, m t-1 represents historical memory, SBERT(·) represents the Sentence-BERT encoder, ‖·‖ represents taking the norm, and Sim(m t-1 , O t ) represents the calculated semantic similarity.

5. A method for generating an emotional support dialogue that integrates perceptual reasoning and enhanced mental memory according to claim 4, characterized in that Based on the semantic similarity, dynamically filter redundant information through the forgetting elimination mechanism to remove outdated memories and obtain the updated memory, which is achieved through the following formula: m t = m t-1 -γORM{η t [Sim(m t-1 , O t )], (1 - λ1)}; Among them, η t (·) represents a threshold function, γ is an adjustment coefficient adaptively adjusted according to the emotional intensity, λ1 is a first preset threshold, and ORM(·) represents a forgetting elimination mechanism.

6. The emotional support dialogue generation method integrating perceptual reasoning and mental memory enhancement according to claim 5, characterized in that Input the emotional state and the updated memory into a large language model, and the large language model generates a targeted response by combining the emotional state and the updated memory, which is achieved through the following formula: R t = F gen (C t , O t , m t ); C t = {U1, R1, U2, R2, …, U t-1 , R t-1 , U t}; Among them, F gen (·) represents a large language model, C t represents the historical conversation context, R t represents the targeted response to U t of 7. A method for generating an emotional support dialogue that integrates perceptual reasoning and enhanced mental memory according to any one of claims 1 to 6, characterized in that, Use an autoregressive decoder to decode the targeted response, and ensure semantic coherence through a masked self-attention mechanism to avoid repeated suggestions in multi-turn conversations, which is achieved through the following formula: M t = η t [Sim(M t+1 , O t ), λ2] Among them, M t represents the EToM memory at the end of the t-th round of conversation, η t (·) represents the threshold function, λ2 is the second preset threshold, Sim(M t+1 , O t ) represents the similarity between M t+1 and O t , θ represents the parameters of the large language model before update, U represents the user input conversation content, R represents the targeted response, D is the training data set, p(R t |U, R, M <t ; θ) represents the probability that the large language model generates a response given the U and R of this round and the M of the previous rounds, E (U,R,M)~D (·) represents taking the expectation, θ * represents the parameters of the updated large language model.

8. An emotional support dialogue generation system that integrates perceptual reasoning and mental memory enhancement, characterized in that, The system includes: An acquisition module for acquiring the conversation content input by the user; A structured mental state extraction module for performing structured mental state extraction on the conversation content input by the user and generating personal clues and factual clues using the emotional mind theory framework; A dynamic linear classification module for dynamically selecting an emotion-dominated or reason-dominated response mode based on the personal clues and factual clues through a personal fact classifier to determine the generated emotional state; An outdated filtering module for using the Sentence-BERT encoder to encode the historical memory and emotional state, calculating the semantic similarity between the two, and dynamically filtering redundant information based on the semantic similarity through the forgetting elimination mechanism to remove outdated memories and obtain the updated memory; A targeted response generation module, configured to input the emotional state and the updated memory into a large language model, and the large language model generates a targeted response by combining the emotional state and the updated memory; A response decoding module, configured to decode the targeted response by using an autoregressive decoder, and ensure semantic coherence through a masked self-attention mechanism, so as to avoid repeated suggestions in multi-turn conversations.

9. An electronic device, characterized in that, Comprising: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute a method for generating an emotional support dialogue that integrates sensible reasoning and mental memory enhancement as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it can implement a method for generating an emotional support dialogue that integrates sensible reasoning and mental memory enhancement as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Hierarchical memory and context awareness retrieval method of role large model and related products

    CN121743515A

  • Psychological state authenticity detection and intervention system based on EEG and psychological orthogonal basis group mapping

    CN122392952A