Emotional dialogue generation method and device, electronic equipment and storage medium

By constructing an emotional dialogue generation method and utilizing cognitive strategy reinforcement learning algorithms and a hybrid reward mechanism, we have achieved accurate identification and professional intervention of users' cognitive distortions, solved the problem of inaccurate cognitive diagnosis in existing technologies, and improved the effectiveness of psychological counseling.

CN122264086APending Publication Date: 2026-06-23HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
Filing Date
2026-03-09
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing technologies lack the ability to deeply understand and diagnose cognitive distortions in psychological counseling, resulting in poor cognitive intervention effects and even exacerbating the cognitive distortions of those seeking help.

Method used

A method for generating emotional dialogues is constructed by building an agent with an emotional dialogue dataset and a pre-trained large language model. By utilizing cognitive strategy reinforcement learning algorithms and a hybrid reward mechanism, it can accurately identify and professionally intervene in user cognitive distortions and generate target response text.

Benefits of technology

It enables accurate identification and professional intervention of cognitive distortions in those seeking help, improves users' negative emotions, alleviates cognitive distortions, and enhances the accuracy and safety of cognitive intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122264086A_ABST
    Figure CN122264086A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, and discloses an emotional conversation generation method and device, electronic equipment and a storage medium, the method comprising the following steps: constructing an emotional conversation dataset and constructing a plurality of intelligent agents based on a pre-trained large language model, which respectively play the roles of a seeker and a consultant; obtaining current conversation text received by a first intelligent agent corresponding to the seeker, mapping a target cognitive distortion classification system of the current conversation text based on the emotional conversation dataset; determining the cognitive bias of the current user based on the target cognitive distortion classification system; determining a target intervention strategy of the current conversation text based on the cognitive bias by using a second intelligent agent corresponding to the consultant, a preset hybrid reward mechanism and a cognitive strategy reinforcement learning algorithm; and optimizing reply text in the target intervention strategy to generate a target reply text, so that the cognitive diagnosis and cognitive intervention strategy accuracy for the seeker are improved, and accurate identification and professional intervention of cognitive distortion of the seeker are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to methods, apparatus, electronic devices, and storage media for generating emotional dialogues. Background Technology

[0002] With the development of artificial intelligence technology, using large language models (LLMs) to provide mental health support has become an important trend in the intersection of natural language processing (NLP) and healthcare. This LLM-based system aims to provide widely accessible emotional support dialogues to those seeking help, alleviating their psychological distress.

[0003] The solutions in related technologies mainly rely on rule-based systems or traditional retrieval-based dialogue models. These methods usually have fixed dialogue templates or match answers from a limited response database, making it difficult to handle complex and ever-changing psychological counseling scenarios. They lack a deep understanding of users' emotions and the ability to deeply model and diagnose the core psychological mechanism of "cognitive distortion." In psychological counseling scenarios, they only focus on identifying and comforting users' explicit emotional labels, lacking the reasoning and diagnosis of implicit logical fallacies in the client. They cannot help the client reconstruct their thinking from a cognitive level, resulting in low accuracy in cognitive diagnosis and poor cognitive intervention effects, which in turn leads to poor help effects or even aggravates the client's cognitive distortion. Summary of the Invention

[0004] This invention provides a method, apparatus, electronic device, and storage medium for generating emotional dialogues, in order to solve the problem that low accuracy in cognitive diagnosis of those seeking help leads to poor cognitive intervention, resulting in poor assistance to those seeking help or even exacerbating their cognitive distortions.

[0005] In a first aspect, the present invention provides a method for generating emotional dialogue, the method comprising: We constructed an emotional dialogue dataset and built multiple intelligent agents based on a pre-trained large language model, which respectively played the roles of help seeker and consultant. Obtain the current dialogue text received by the first intelligent agent corresponding to the person seeking help, and map the target cognitive distortion classification system of the current dialogue text based on the emotional dialogue dataset; Based on the target cognitive distortion classification system, determine the current user's cognitive bias; Based on cognitive bias, the second intelligent agent corresponding to the consultant determines the target intervention strategy for the current dialogue text using a pre-set hybrid reward mechanism and cognitive strategy reinforcement learning algorithm. Based on preset conditional constraint functions, the target response text in the target intervention strategy is optimized to generate the target response text.

[0006] This invention maps the current dialogue text received by the first intelligent agent corresponding to the user seeking help into a target cognitive distortion classification system, thereby determining the user's cognitive bias. Then, a second intelligent agent corresponding to the consultant uses a preset hybrid reward mechanism and a cognitive strategy reinforcement learning algorithm to determine the target intervention strategy for the current dialogue text. After optimizing the target intervention strategy, a corresponding target response text is generated. This constructs a complete technology chain from cognitive distortion diagnosis to optimal strategy decision-making and intervention response generation, enabling the emotional dialogue method to leap from traditional surface-level emotional resonance to a professional-grade intelligent agent with deep cognitive repair capabilities. It ensures that the final output target response text of the model maintains logical consistency with its internal cognitive bias diagnosis results, achieving accurate identification and professional intervention of the user's cognitive distortion, which is beneficial for improving the user's negative emotions and alleviating the user's cognitive distortion psychology.

[0007] In one optional implementation, a target cognitive distortion classification system is mapped from the current dialogue text based on the emotion dialogue dataset, including: The text features are obtained by parsing the current dialogue text; Based on the mapping relationship between dialogue text and cognitive label tuples concatenated from dialogue text in the emotion dialogue dataset, the target cognitive label tuples of the mapped text features are obtained. A target cognition distortion classification system is determined based on target cognition label tuples.

[0008] This invention parses the current dialogue text to obtain text features, and based on the mapping relationship between the dialogue text and the cognitive tag tuples concatenated with the dialogue text, and the target cognitive tag tuple mapping the text features, it infers the logical fallacies hidden in the user's current text dialogue and diagnoses the user's current cognition.

[0009] In one optional implementation, the cognitive bias of the current user is determined based on a target cognitive distortion classification system, including: A pre-defined cognitive distortion classification system is established; wherein, the pre-defined cognitive distortion classification system includes the correspondence between the cognitive distortion classification system and different types of cognitive distortion, different intensities of cognitive distortion, and different levels of cognitive security risk; Based on a pre-defined cognitive distortion classification system, the target cognitive distortion type, target cognitive distortion intensity, and target cognitive security risk level are abstracted from the target cognitive distortion classification system. The cognitive bias of the current user is determined based on the type of target cognitive distortion, the intensity of target cognitive distortion, and the level of target cognitive security risk.

[0010] This invention determines the cognitive bias of the current user by abstracting the target cognitive distortion type, target cognitive distortion intensity, and target cognitive safety risk level from the target cognitive distortion classification system of the current dialogue text based on a preset cognitive distortion classification system. This allows for flexible adjustment of intervention methods according to the specific type, intensity, and risk level of cognitive distortion, enabling the intelligent agent to achieve stable, professional, and safe cognitive intervention in complex and ever-changing consultation scenarios.

[0011] In one optional implementation, a second agent corresponding to the consultant determines the target intervention strategy for the current dialogue text based on cognitive bias using a preset hybrid reward mechanism and a cognitive strategy reinforcement learning algorithm, including: Initialize the dialogue learning environment of the second agent, which is an interactive loop that provides intervention strategies based on the current dialogue text; The second intelligent agent is used to perform reinforcement learning on the dialogue learning environment based on the cognitive policy reinforcement learning algorithm to determine the response policy value function of the current dialogue text; The comprehensive reward function acquired at the current time step is analyzed using a second intelligent agent to analyze cognitive biases. Based on the preset symptom improvement reward function, preset rule matching reward function, and preset safety constraint reward function in the hybrid reward mechanism, the comprehensive reward function is solved. Based on the solution results of the response strategy value function and the comprehensive reward function, the deep Q-network is iteratively trained to obtain the target intervention strategy, which includes at least the Socratic questioning strategy and the cognitive reconstruction strategy.

[0012] This invention uses a cognitive strategy reinforcement learning algorithm to perform reinforcement learning on the dialogue learning environment, determining the response strategy value function of the current dialogue text; and solves the comprehensive reward function by pre-setting a symptom improvement reward function, a pre-setting rule matching reward function, and a pre-setting safety constraint reward function; thereby obtaining the optimal intervention strategy through continuous iterative training and improving the evaluation accuracy of the model.

[0013] In an optional implementation, before generating the target response text by optimizing the response text in the target intervention strategy based on a preset condition constraint function, the method further includes: Based on the target intervention strategy, the initial response text for the current text is inferred using a pre-trained optimal strategy model; A pre-defined teacher model is established, and the initial response text is used as a prompt for the pre-defined teacher model to enhance the initial response text and generate the response text for the target intervention strategy.

[0014] This invention enhances the initial response text of the current text inferred from the target intervention strategy to generate an enhanced response text of the target intervention strategy, thereby improving the model's ability to provide high-quality cognitive intervention responses. It ensures that the external response text output by the model maintains strict logical consistency with the cognitive bias of its internal diagnostic judgment, thus improving the interpretability of the model.

[0015] In one optional implementation, the target response text is generated by optimizing the response text in the target intervention strategy based on a preset conditional constraint function, including: Determine the preset conditional mask loss function and the preset optimization objective function in the preset conditional constraint function; Using cognitive bias as an output condition constraint, the preset conditional mask loss function is solved; The target intervention strategy is used as a constraint on the response text to solve the preset optimization objective function; Based on the solution results of the preset conditional mask loss function and the preset optimization objective function, the target response text is generated.

[0016] This invention uses a preset conditional mask loss function and a preset optimization objective function in the preset conditional constraint function to treat cognitive bias as an output conditional constraint, so that the model can accurately output the type, intensity and risk level of cognitive distortion, thereby improving the model's diagnostic ability. It also uses the target intervention strategy as a response file constraint, so that the model can generate an intervention response based on the optimal strategy, thereby improving the model's professional cognitive diagnostic ability and ensuring that the external response text output by the model and the cognitive bias of its internal diagnostic judgment maintain strict logical consistency, thereby improving the interpretability of the model.

[0017] In one alternative implementation, the method further includes: The process involves updating the current dialogue text received by the first agent, updating the cognitive bias based on the updated current dialogue text, and then executing the step of determining the target intervention strategy for the current dialogue text by the second agent based on the cognitive bias, using a hybrid reward mechanism and a cognitive policy reinforcement learning algorithm.

[0018] This invention achieves cognitive bias diagnosis for each round of user input by updating the current dialogue text and updating cognitive bias based on the updated current dialogue text, and updates the corresponding target intervention strategy based on the updated cognitive bias diagnosis results, thereby improving the flexibility of the model.

[0019] Secondly, the present invention provides an emotional dialogue generation device, the device comprising: The module is used to build an emotional dialogue dataset and to build multiple intelligent agents based on a pre-trained large language model, which play the roles of help seeker and consultant respectively. The dialogue acquisition module is used to acquire the current dialogue text received by the first intelligent agent corresponding to the helper, and to map the target cognitive distortion classification system of the current dialogue text based on the emotional dialogue dataset. The cognition determination module is used to determine the current user's cognitive bias based on the target cognitive distortion classification system; The strategy determination module is used to determine the target intervention strategy for the current dialogue text by utilizing the second intelligent agent corresponding to the consultant based on cognitive bias and by using a preset hybrid reward mechanism and cognitive strategy reinforcement learning algorithm. The generation module is used to generate target response text based on preset conditional constraint functions and optimized response text in the target intervention strategy.

[0020] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the method described in the first aspect or any corresponding embodiment thereof.

[0021] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to perform the method described in the first aspect or any corresponding embodiment thereof. Attached Figure Description

[0022] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of an application scenario according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the first type of emotional dialogue generation method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of a second process for generating emotional dialogue according to an embodiment of the present invention; Figure 4 This is a flowchart illustrating the reasoning process of the emotional dialogue generation method according to an embodiment of the present invention; Figure 5 This is a structural block diagram of an emotional dialogue generation device according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0026] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0027] As an optional application scenario of this invention, such as Figure 1 As shown, application 101 is installed in terminal device 110, and user 130 can interact with application 101 through terminal device 110 and / or access device of terminal device 110.

[0028] For example, application 101 can be any application that provides question-and-answer related services. For instance, application 101 could be a question-and-answer interactive application, such as a text-to-text application, an image-to-text application, etc. Figure 1 In the application scenario shown, if application 101 is active, the terminal device 110 can display the interface 102 of application 101. The interface 102 may include various pages that application 101 can provide, such as interactive pages, settings pages, query pages, etc.

[0029] In some embodiments, terminal device 110 is communicatively connected to server 120 to provide services to application 101. Terminal device 110 may be a mobile terminal, fixed terminal, or portable terminal, etc., including but not limited to mobile phones, desktop computers, laptop computers, multimedia tablets, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 may also support any type of interface, and server 120 may be various types of computing systems or servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0030] It should be noted that, Figure 1 This is merely an example of an application scenario and does not limit the scope of protection of this invention.

[0031] The embodiments of the present invention will now be described with reference to the accompanying drawings. It should be understood that the pages shown in the drawings are merely examples, and various page designs are possible in practice. The various graphic elements on the page may have different arrangements and different visual representations, one or more elements may be omitted or replaced, and one or more other elements may also be present; no limitations are imposed in the embodiments of the present invention. Furthermore, the embodiments are described primarily with respect to terminal device 110 in this context. It should be understood that the actions described relative to terminal device 110 can be performed by application 101 on terminal device 110, or can be performed by application 101 in conjunction with its server (e.g., server 120).

[0032] Traditional technical solutions primarily rely on rule-based systems or traditional retrieval-based dialogue models. These methods typically pre-set fixed dialogue templates or match answers from a limited response database, making them ill-suited for complex and ever-changing psychological counseling scenarios. While they can provide some emotional feedback, they lack a deep understanding and flexibility regarding the user's emotions.

[0033] In recent years, with breakthroughs in LLM technology, mainstream technical solutions have begun to shift towards leveraging the powerful generative and contextual understanding capabilities of LLM to construct intelligent agents for psychological counseling. Existing cutting-edge technical solutions mainly fall into two categories: the first is specialized models based on fine-tuning, such as SoulChat, CPsyCoun, and ChatCounselor. These solutions enable general LLM models to possess a certain degree of empathic dialogue ability by performing supervised fine-tuning (SFT) or preference learning (such as DPO) on specific psychological counseling datasets. They typically learn from the counselor's past responses as "standard answers," focusing on mimicking the language style and surface empathy patterns of human counselors.

[0034] The second category is simulation methods based on role-playing or cue engineering, such as PsyDT and AnnaAgent. These schemes guide a general-purpose LLM to act as a therapist by constructing a prompt containing few-shot examples or memory modules, generating responses based on the dialogue history. Some works (such as CSO) attempt to introduce planning mechanisms such as Monte Carlo Tree Search (MCTS) to optimize the selection of response strategies. Overall, existing technologies mainly focus on improving the fluency and empathy of model responses, attempting to provide users with emotional comfort in natural dialogue.

[0035] While solutions in related technologies have made some progress in simulating everyday emotional dialogues, they still have significant methodological flaws and mechanistic bottlenecks when dealing with professional psychological counseling tasks.

[0036] First, current technologies generally lack the ability to deeply model and diagnose the core psychological mechanism of "cognitive distortion." Most current LLM (Less-Low Mood) counseling systems remain at the "surface empathy" stage, meaning the models primarily focus on identifying and comforting users' explicit emotional labels (such as sadness and anxiety), while neglecting the root cause of psychological distress emphasized in Cognitive Behavioral Therapy (CBT): irrational cognitive distortions (such as "catastrophizing" and "black-and-white thinking"). Due to the lack of reasoning and diagnostic steps regarding the implicit logical fallacies of the client, existing models often only provide superficial comfort and fail to help the client reconstruct their thinking at the cognitive level, resulting in interventions that only address the symptoms and not the root cause.

[0037] Secondly, existing solutions suffer from opacity and safety risks in the selection of intervention strategies. Most models directly output responses through end-to-end generation or retrieve strategies based on coarse rules, lacking an explicit and interpretable strategy decision-making process. This approach makes it difficult for models to flexibly adjust intervention methods according to the specific type, intensity, and risk level (such as suicide risk) of cognitive distortion. For example, when faced with high-risk, severe cognitive distortion, existing models may still employ mild empathy strategies, or even generate inappropriate suggestions due to hallucination issues (such as reinforcing the user's erroneous cognition), which poses a significant safety hazard in psychological counseling scenarios.

[0038] Finally, from the perspective of data utilization and learning mechanisms, existing methods suffer from the limitation of "misaligned imitation." Current mainstream fine-tuning data (SFT / DPO data) largely originates from raw consultation records without rigorous cognitive annotation. The quality of human counselor responses in these records varies greatly, and they may not all accurately address cognitive distortions. Existing models directly imitate these responses as Ground Truth, preventing them from learning the professional logical chain of "diagnosis-strategy-intervention." Instead, they may learn inefficient or even erroneous consultation patterns. This passive learning approach, lacking "strategy awareness," makes it difficult for existing intelligent agents to achieve stable, professional, and safe cognitive intervention in complex and ever-changing consultation scenarios.

[0039] The purpose of this invention is to overcome the shortcomings of the prior art and provide a Cognitive Policy-driven Large Language Model (CoPoLLM) method for emotional support dialogue, namely, an emotional dialogue generation method. This method aims to break through the limitations of existing large-scale psychological counseling models that only remain at the surface level of empathetic responses. By constructing a highly interpretable and safe "diagnosis-strategy-intervention" closed-loop mechanism, it can achieve accurate identification and professional intervention of cognitive distortions in the client.

[0040] According to an embodiment of the present invention, an embodiment of an emotional dialogue generation method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0041] This embodiment provides a method for generating emotional dialogue, which can be used in the aforementioned mobile terminals, such as mobile phones and computers. Figure 2 This is a schematic diagram of the first type of emotional dialogue generation method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Construct an emotional dialogue dataset and construct multiple intelligent agents based on a pre-trained large language model, which respectively play the roles of the person seeking help and the consultant.

[0042] It should be noted that the emotional dialogue dataset is based on the theory of cognitive behavioral therapy (CBT) and aims at the core goal of "cognitive diagnosis and cognitive intervention". It performs refined modeling of the cognitive distortions implied in the expressions of those seeking help. Through the emotional dialogue dataset system, standardized cognitive labels that can be used for subsequent reasoning, intervention strategy matching and risk assessment can be abstracted from the original natural language, providing a structured and computable prior cognitive knowledge base for subsequent cognitive reasoning and intervention strategies.

[0043] The system first collects raw dialogue data from multiple publicly available emotional support dialogue corpora and constructs a cognitive bias classification system within the framework of cognitive behavioral therapy. Through an expert annotation process, each round of help-seeker expressions is assigned structured labels such as cognitive distortion type, distortion intensity, and safety risk level, ultimately forming an emotional dialogue dataset that provides high-quality cognitive priors for subsequent decision-making and training.

[0044] In this context, an intelligent agent is a proxy capable of perceiving the environment and taking actions to achieve a specific goal. In the optional application scenarios of this embodiment, multiple intelligent agents of a pre-trained large language model are proxies capable of perceiving the current dialogue learning environment and providing intervention strategies based on the dialogue text to achieve the goal of generating response text. This intelligent agent can be software, hardware, or a system, possessing autonomy, adaptability, and interactive capabilities. The intelligent agent can perceive changes in data input within the dialogue learning environment—that is, multiple different received dialogue texts—and make judgments and decisions based on its learned knowledge and algorithms, thereby providing intervention strategies to achieve the goal of accurately identifying and professionally intervening in the cognitive distortions of the person seeking help.

[0045] It should be noted that the intelligent agent provided in this embodiment acts as both the seeker and the consultant. The seeker receives dialogue text input by the user, which can be a textual piece of language. The consultant provides corresponding intervention strategies based on the received dialogue text to achieve the goal of accurately identifying and professionally intervening in the seeker's cognitive distortions. The intervention strategies provided by the consultant are based on irrational logic such as "catastrophic thinking" and "emotional reasoning" implicit in the dialogue text, rather than the surface meaning reflected in the dialogue text. Furthermore, the intervention strategies in this embodiment are strategy-oriented and can be determined as the optimal intervention strategies based on cognitive type, intensity, and risk level.

[0046] Step S202: Obtain the current dialogue text received by the first intelligent agent corresponding to the person seeking help, and map the target cognitive distortion classification system of the current dialogue text based on the emotional dialogue dataset.

[0047] It should be noted that the target cognition distortion classification system is used to characterize the systematic cognitive biases present in the current dialogue text.

[0048] Cognitive distortion refers to a type of thinking error that causes difficulties in the user's information processing, ultimately leading to psychological barriers. Examples include "all or nothing" thinking, overgeneralization, mental filtering, denigrating positive things, jumping to conclusions, exaggeration and understatement, emotional reasoning, "should" statements, labeling and inappropriate labeling, and attribution.

[0049] Step S203: Based on the target cognitive distortion classification system, determine the current user's cognitive bias.

[0050] It should be noted that cognitive bias refers to the phenomenon that when an individual perceives themselves, others, or the external environment, the perceptual results are distorted due to their own or situational factors.

[0051] Step S204: Based on cognitive bias, the second agent corresponding to the consultant determines the target intervention strategy for the current dialogue text using a preset hybrid reward mechanism and a cognitive strategy reinforcement learning algorithm.

[0052] It should be noted that the second intelligent agent corresponding to the consultant simulates the client's thinking patterns, language habits, emotional reactions, and cognitive biases, and infers a dialogue pattern that is more easily accepted by the client, thereby determining the target intervention strategy.

[0053] The pre-set hybrid reward mechanism includes a cognitive type reward mechanism, a cognitive intensity reward mechanism, and a cognitive safety risk level reward mechanism. It is used to characterize whether the current dialogue causes the seeker's emotions to be relieved, whether it reduces the density of negative words implied in the language, whether it maintains the willingness to have a dialogue, and whether the seeker shows more flexible cognition in subsequent dialogues and whether it produces positive changes at the behavioral level.

[0054] Among them, the cognitive strategy reinforcement learning algorithm learns and reinforces which dialogue intervention strategies can maximize mixed rewards under the current state of cognitive bias. These cognitive strategies include, but are not limited to, Socratic questioning, cognitive restructuring guidance, and empathic responses.

[0055] Step S205: Based on the preset condition constraint function, optimize the response text in the target intervention strategy to generate the target response text.

[0056] It should be noted that by using preset conditional constraint functions to force the model to learn to accurately output the type, intensity, and risk level of cognitive distortion, the model's diagnostic capabilities are enhanced. Furthermore, by forcing the model to generate intervention responses based on the optimal target intervention strategy, emotional dialogue is elevated from traditional surface-level emotional resonance to a professional-grade intelligent agent with deep cognitive repair capabilities.

[0057] The emotional dialogue generation method provided in this embodiment maps the current dialogue text received by the first intelligent agent corresponding to the helper to the target cognitive distortion classification system of the current text, thereby determining the current user's cognitive bias. Then, the second intelligent agent corresponding to the consultant uses a preset hybrid reward mechanism and cognitive strategy reinforcement learning algorithm to determine the target intervention strategy of the current dialogue text. After optimizing the target intervention strategy, the corresponding target response text is generated. This constructs a complete technology chain from cognitive distortion diagnosis to optimal strategy decision-making and intervention response generation, enabling the emotional dialogue method to leap from traditional surface emotional resonance to a professional-grade intelligent agent with deep cognitive repair capabilities. It ensures that the target response text output by the model maintains logical consistency with its internal cognitive bias diagnosis results, achieving accurate identification and professional intervention of the helper's cognitive distortion, which is beneficial to improving the user's negative emotions.

[0058] This embodiment provides a method for generating emotional dialogue, which can be used in the aforementioned mobile terminals, such as mobile phones and computers. Figure 3 This is a schematic diagram of a second process for generating emotional dialogue according to an embodiment of the present invention, such as... Figure 3 As shown, the process includes the following steps: Step S301 involves constructing an emotional dialogue dataset and building multiple agents based on a pre-trained large language model, each playing the roles of a seeker of help and a counselor. For details, please refer to [link to details]. Figure 2 Step S201 of the illustrated embodiment will not be described again here.

[0059] Step S302: Obtain the current dialogue text received by the first intelligent agent corresponding to the person seeking help, and map the target cognitive distortion classification system of the current dialogue text based on the emotional dialogue dataset.

[0060] Specifically, step S302 includes: Step S3021: Parse the current dialogue text to obtain text features.

[0061] It should be noted that semantic analysis of the current dialogue text yields text features, which are used to characterize the density of negative words contained in the current dialogue text, as well as extreme words that can characterize black-and-white cognitive bias; words and thought processes that characterize the exaggeration and downplaying of catastrophic cognitive bias, such as whether a small problem is exaggerated into the worst outcome; words and thought processes that characterize normalization cognitive bias, such as whether all negative external emotions are attributed to oneself; and words and thought processes that characterize emotional reasoning, such as whether feelings are used to replace facts.

[0062] Step S3022: Based on the mapping relationship between the dialogue text and the cognitive label tuples concatenated from the dialogue text in the emotion dialogue dataset, map the target cognitive label tuples of the text features.

[0063] It should be noted that the cognitive label tuple represents the type of cognitive distortion, the intensity of cognitive distortion, and the level of cognitive safety risk. This cognitive label tuple is represented as... ;in, Indicates the type of cognitive distortion, used to characterize the pattern of systematic cognitive bias present in the expression (current dialogue text); Indicates the intensity of cognitive distortion, used to quantify the degree to which this cognitive bias affects an individual's judgment and behavior; This indicates the level of cognitive safety risk, used to assess whether the expression (current dialogue text) involves potential self-harm, suicide, or other high-risk psychological issues. This three-dimensional cognitive label tuple endows cognitive biases not only with category attributes but also with orderable and decision-making risk semantics, providing a foundation for subsequent tiered intervention and safety constraints.

[0064] Specifically, initialize the reinforcement learning environment (dialogue learning environment). In this environment, time steps are defined. The state space below and action space We define the state representation for reinforcement learning as follows: In order to capture the deep semantics in dialogue, we construct a unified text state representation. This representation includes not only the current text uttered by the person seeking help (the current dialogue text) It also concatenates its corresponding cognitive tag tuple description, as shown in the following formula:

[0065] in, This indicates a vector or text concatenation operation. It is a function that converts labels into natural language descriptions; This represents a function that converts the intensity of cognitive distortion into a natural language description; This represents a function that converts cognitive safety risk levels into natural language descriptions; through a state encoder. Map text states to continuous low-dimensional state vectors .

[0066] The action space is defined as: a set of k standard CBT intervention strategies, such as utilizing gray areas, evidence examination, and de-catastrophizing. (Consultant agent) The goal is to select the optimal action. .

[0067] Step S3023: Determine the target cognition distortion classification system based on the target cognition label tuple.

[0068] It should be noted that the target cognitive distortion classification system is determined by identifying the cognitive label tuples corresponding to the current dialogue text.

[0069] The emotional dialogue generation method provided in this embodiment obtains text features by parsing the current dialogue text, and diagnoses the logical fallacies hidden in the user's current text dialogue and the user's current cognition based on the mapping relationship between the dialogue text and the cognitive tag tuples concatenated with the dialogue text and the target cognitive tag tuple that maps the text features.

[0070] Step S303: Based on the target cognitive distortion classification system, determine the current user's cognitive bias.

[0071] Specifically, step S303 includes: Step S3031: Determine a preset cognitive distortion classification system; wherein, the preset cognitive distortion classification system includes the correspondence between the cognitive distortion classification system and different cognitive distortion types, different cognitive distortion intensities, and different cognitive safety risk levels.

[0072] Specifically, in constructing the different cognitive distortion types in the pre-defined cognitive distortion classification system, this embodiment adopts the classic CBT cognitive bias classification system proposed by Beck, and selects the eight most representative core cognitive distortion types to cover the most common logical biases in emotional support and psychological counseling scenarios. Their definitions and descriptions are shown in Table 1.

[0073] Table 1 Types of Cognitive Distortion

[0074] Specifically, to differentiate the impact of the same cognitive bias on psychological states at different degrees of severity within the pre-defined cognitive distortion classification system, this embodiment further introduces a cognitive distortion intensity dimension i, which is divided into three levels. This dimension reflects the stability and frequency of cognitive biases and an individual's ability to self-reflect on their thinking patterns. Its grading criteria are shown in Table 2.

[0075] Table 2 Cognitive Distortion Intensity

[0076] Specifically, for different cognitive safety risk levels in the pre-defined cognitive distortion classification system, and considering the potentially high-risk psychological states in emotional support dialogues, this embodiment introduces a safety risk level dimension *r* to identify key semantic signals related to life safety. The definition of this dimension follows risk assessment standards in clinical psychology, and its classification is shown in Table 3.

[0077] Table 3 Cognitive Safety Risk Levels

[0078] The sentiment-supported dialogue dataset is constructed based on multiple publicly available sentiment-supported dialogue corpora and obtains three-dimensional cognitive labels through expert annotation. Before entering the system, the data has undergone consistency checks and conflict resolution to ensure that each dialogue sample (the current dialogue text) has clear... Tag combination. Through the above design, a cognitive bias data foundation that combines psychological theoretical basis with engineering usability is provided for the emotional dialogue generation method, so that subsequent cognitive reasoning, strategy matching and security control processes are all based on interpretable and verifiable cognitive structures.

[0079] Step S3032: Based on the preset cognitive distortion classification system, abstract the target cognitive distortion type, target cognitive distortion intensity, and target cognitive security risk level in the target cognitive distortion classification system.

[0080] It should be noted that the type of distorted goal cognition refers to the thought patterns that are identified as playing a dominant role in the distress of the person seeking help in the current dialogue text and that require priority intervention. For example, the current dialogue text is: "Because I am too stupid (personalization), I will definitely fail this exam (catastrophizing), and my life is over (overgeneralization)."

[0081] The intensity of the target cognitive distortion is quantified by the degree to which the type of cognitive distortion occupies the client's mind and its emotional impact. The current dialogue text is "Because I'm too stupid (personalization), I'm sure I'll fail this exam (catastrophizing), my life is over (overgeneralization)." The corresponding cognitive distortion is clearly identifiable and has a negative impact on emotions and behavior, but the individual still retains a certain degree of reflection and regulation ability.

[0082] The target cognitive safety risk level is used to assess whether this cognitively distorted way of thinking might lead the person seeking help to engage in behavior that harms themselves or others. The current dialogue text is "Because I'm too stupid (personalization), I'm definitely going to fail this exam (catastrophizing), my life is over (overgeneralization)." The corresponding cognitive safety risk level is relatively mild, mainly an adaptation problem, and the individual's overall functioning and coping resources remain intact.

[0083] Step S3033: Determine the current user's cognitive bias based on the target cognitive distortion type, the target cognitive distortion intensity, and the target cognitive security risk level.

[0084] The emotional dialogue generation method provided in this embodiment determines the current user's cognitive bias by abstracting the target cognitive distortion type, target cognitive distortion intensity, and target cognitive safety risk level from the target cognitive distortion classification system of the current dialogue text based on a preset cognitive distortion classification system. This allows for flexible adjustment of intervention methods according to the specific type, intensity, and risk level of cognitive distortion, enabling the intelligent agent to achieve stable, professional, and safe cognitive intervention in complex and ever-changing consultation scenarios.

[0085] Step S304: Based on cognitive bias, the second agent corresponding to the consultant determines the target intervention strategy for the current dialogue text using a preset hybrid reward mechanism and a cognitive strategy reinforcement learning algorithm.

[0086] Specifically, step S304 includes: Step S3041: Initialize the dialogue learning environment of the second agent. The dialogue learning environment is an interactive loop that provides intervention strategies based on the current dialogue text.

[0087] It should be noted that the dialogue learning environment of the second agent is initialized so that the second agent can provide intervention strategies based on the current dialogue text through the mapping relationship between text features and cognitive bias diagnosis results.

[0088] Step S3042: Use the second agent to perform reinforcement learning on the dialogue learning environment based on the cognitive policy reinforcement learning algorithm to determine the response policy value function of the current dialogue text.

[0089] Among them, the cognitive policy reinforcement learning algorithm is used to determine the response policy function of the current dialogue text by utilizing a deep Q-network and a preset decision policy function.

[0090] Specifically, in the current dialogue learning environment simulating the interaction loop, the consultant agent uses a deep Q-network (DQN) to approximate the action value function (response strategy value function). Its pre-set decision-making strategy adopts -greedy strategy:

[0091] in, This represents the solution result of the action value function; Indicates the current time step The state is represented by low-dimensional features generated by the encoder after information such as the current dialogue text, cognitive distortion type, intensity, and risk level are mapped. This refers to specific intervention actions (strategies) in the action space, such as strategies like "Socratic questioning" and "evidence inspection." This represents the current parameters of the DQN network used by the agent; Step S3043: Analyze the comprehensive reward function obtained at the current time step using the second agent to analyze the cognitive bias.

[0092] It should be noted that, in order to solve the problem of inaccurate evaluation caused by the illusion of large models, this embodiment innovatively designs a hybrid reward function of "rule guidance and model correction".

[0093] Step S3044: Solve the comprehensive reward function based on the preset symptom improvement reward function, preset rule matching reward function, and preset safety constraint reward function in the hybrid reward mechanism.

[0094] It should be noted that this embodiment also includes constructing an evaluator agent, which uses the preset symptom improvement reward function, preset rule matching reward function and preset safety constraint reward function in the hybrid reward mechanism to solve the comprehensive reward function.

[0095] Specifically, in time step The reward obtained by the agent at that time The calculation formula is as follows:

[0096] in, It is a symptom improvement reward (preset symptom improvement reward function), determined by the evaluator agent. Scoring is based on the degree of reduction in cognitive distortion in the preceding and following statements made by the first intelligent agent corresponding to the requester. This represents the dialogue text at time t. This represents the dialogue text at time t+1.

[0097] It is a rule-matching reward (preset rule-matching reward function), indicating that the currently selected strategy is verified based on the CBT manual. Does it apply to the current type of cognitive distortion? .

[0098] It is a safety constraint reward (preset safety constraint reward function), when the risk level When the value is high, the model is forced to choose a conservative safety strategy; otherwise, a heavy penalty is imposed. , , These are the weighting coefficients for each part. Indicates the currently selected strategy; This indicates the level of perceived security risk.

[0099] Step S3045: Based on the solution results of the response strategy value function and the comprehensive reward function, the deep Q network is iteratively trained to obtain the target intervention strategy, wherein the target intervention strategy includes at least a Socratic questioning strategy and a cognitive reconstruction strategy.

[0100] Specifically, based on the aforementioned rewards, the model parameters of the DQN network are updated by minimizing the temporal difference error. Target Q value The calculation formula is as follows:

[0101] in, As a discount factor, These are the reference model parameters for the target network, used to measure the model shift of the DQN network before and after training. Through continuous iterative training, the optimal policy network is obtained. The target intervention strategy is obtained from the optimal strategy network; Indicates the next state Given all available candidate intervention strategies, the system iterates through the action space to find the optimal subsequent strategy that maximizes the value function when calculating the target value. It represents the state at the current time step t+1, and is a low-dimensional feature representation generated by the encoder after mapping information such as the dialogue text, cognitive distortion type, intensity, and risk level at time t+1.

[0102] The emotional dialogue generation method provided in this embodiment uses a cognitive strategy reinforcement learning algorithm to perform reinforcement learning on the dialogue learning environment, thereby determining the response strategy value function of the current dialogue text; and solves the comprehensive reward function by pre-setting a symptom improvement reward function, a pre-setting rule matching reward function, and a pre-setting safety constraint reward function; thus, the optimal intervention strategy is obtained through continuous iterative training, improving the evaluation accuracy of the model.

[0103] Step S305: Based on the target intervention strategy, the initial response text for the current text is inferred using the pre-trained optimal strategy model.

[0104] Specifically, to address the issue of a lack of high-quality cognitive intervention responses in the raw data, a pre-trained optimal strategy is utilized. (Targeted intervention strategy) Perform data augmentation. For each statement made by the person seeking help in the dataset... First use Deducing the optimal intervention action (Initial response text).

[0105] Step S306: Determine the preset teacher model, and use the initial response text as a prompt for the preset teacher model to enhance the initial response text and generate the response text of the target intervention strategy.

[0106] This action (initial response text) The input is fed into a pre-defined teacher model (such as GPT-4o) to generate high-quality response text that conforms to this strategy. Each sample pair is represented as ,here For the context of the dialogue, It contains cognitive label sequences (cognitive label tuples). Response text aligned with strategy .

[0107] The emotional dialogue generation method provided in this embodiment enhances the initial response text of the current text inferred from the target intervention strategy to generate an enhanced response text of the target intervention strategy. This improves the model's ability to provide high-quality cognitive intervention responses, ensures that the external response text output by the model maintains strict logical consistency with the cognitive bias of its internal diagnostic judgment, and improves the interpretability of the model.

[0108] Step S307: Based on the preset condition constraint function, optimize the response text in the target intervention strategy to generate the target response text.

[0109] Specifically, step S307 includes: Step a1: Determine the preset conditional mask loss function and the preset optimization objective function in the preset conditional constraint function.

[0110] It should be noted that the dual-stream conditional optimization algorithm is used to fine-tune the parameters of the target large language model. To prevent the learning gradient of the diagnostic task from being overwhelmed by the generation task, a pre-defined conditional mask loss function is designed. The pre-defined optimization objective function combines the diagnostic and intervention flows.

[0111] Step a2: Using cognitive bias as an output condition constraint, solve the preset conditional mask loss function.

[0112] Specifically, this embodiment designs a target masking mechanism. Among them, the conditional mask loss function is defined (preset conditional mask loss function). as follows:

[0113] Where, if the mark Belongs to the target sequence ,but Otherwise, it is 0. The parameters of the large language model LLM to be trained. It represents the mathematical expectation, which in the loss function represents the average loss calculated over the entire training dataset or the current batch of samples.

[0114] Step a3: Using the target intervention strategy as a constraint on the response text, solve the preset optimization objective function.

[0115] Specifically, the preset optimization objective function is as follows:

[0116] Among them, through The forced model learns to accurately output the type, intensity, and risk level of cognitive distortion, thereby improving the model's diagnostic capabilities; through The forced model generates intervention responses based on the optimal strategy; The parameters of the large language model LLM to be trained; The cognitive label tuple at time t contains the type of cognitive distortion, the intensity of cognitive distortion, and the level of cognitive safety risk. This represents the response text after data augmentation of the initial response text; This indicates the dialogue context, including the current dialogue text and its corresponding response text.

[0117] Step a4: Based on the solution results of the preset conditional mask loss function and the preset optimization objective function, generate the target response text.

[0118] This embodiment constructs a two-stream conditional optimization algorithm and a target masking mechanism. Addressing the issue of large models easily forgetting diagnostic logic during fine-tuning, a two-stream loss function is designed to forcibly decouple and simultaneously optimize the "cognitive label prediction stream" and the "intervention text generation stream" during training. This design ensures that the model's final output of external response text (i.e., intervention rhetoric) maintains strict logical consistency with its internal diagnostic judgments (cognitive distortion type, risk level), achieving true "consistency between words and actions" and high interpretability.

[0119] The emotional dialogue generation method provided in this embodiment uses a preset conditional mask loss function and a preset optimization objective function in the preset conditional constraint function to take cognitive bias as an output conditional constraint, so that the model can accurately output the type, intensity and risk level of cognitive distortion, thereby improving the model's diagnostic ability. It also takes the target intervention strategy as a response file constraint, so that the model can generate intervention responses based on the optimal strategy, thereby improving the model's professional cognitive diagnostic ability and ensuring that the external response text output by the model and the cognitive bias of its internal diagnostic judgment maintain strict logical consistency, thereby improving the interpretability of the model.

[0120] Step S308: Update the current utterance text received by the first agent, update the cognitive bias according to the updated current utterance text, and execute the step of determining the target intervention strategy of the current dialogue text by the second agent based on the cognitive bias, using a hybrid reward mechanism and a cognitive strategy reinforcement learning algorithm.

[0121] It should be noted that the emotional dialogue dataset is updated in each round of dialogue. Each input from the person seeking help is evaluated to update cognitive biases, and the target intervention strategy is updated based on these updated cognitive biases, thereby generating the target response text.

[0122] The emotional dialogue generation method provided in this embodiment updates the current dialogue text and updates cognitive biases based on the updated current dialogue text, thereby enabling cognitive bias diagnosis for each round of user input and updating the corresponding target intervention strategy based on the updated cognitive bias diagnosis results, thus improving the flexibility of the model.

[0123] Combination Figure 4 This section describes an application example of this embodiment.

[0124] Specifically, in this embodiment, by constructing a fine-grained cognitive distortion classification system and dataset, the model is able to accurately identify irrational logic such as "catastrophic thinking" and "emotional reasoning" implied in the help seeker's words, rather than merely reflecting emotions.

[0125] By establishing a strategy-oriented intervention mechanism and utilizing a multi-agent reinforcement learning environment to explore the optimal intervention strategy, the model can autonomously decide whether to adopt professional CBT strategies such as "Socratic questioning" or "cognitive reconstruction" based on the type, intensity, and risk level of cognitive distortion, thus avoiding blind generation.

[0126] Achieve end-to-end alignment for dual-stream condition optimization; through an innovative dual-stream condition optimization algorithm, deeply integrate the advanced strategic knowledge acquired through reinforcement learning with the generative capabilities of a large language model, thereby generating high-quality intervention responses that conform to the professional logic of psychological counseling and possess natural language fluency, while strictly controlling safety risks during the intervention process.

[0127] Specifically, in the optional application scenarios of this embodiment, the system architecture of this method mainly includes three core modules: Cognitive Bias Data Foundation (CBDF), Cognitive Policy Reinforcement Learning Engine (CPRL), and Dual-Stream Conditional Optimization (DSCO).

[0128] The Cognitive Bias Data Foundation Module (CBDF) is used to build, maintain, and continuously update the CogBiasESC cognitive bias emotion support dialogue dataset, providing a structured and computable prior cognitive knowledge foundation for various cognitive reasoning and intervention modules in the CoPoLLM system.

[0129] Unlike existing emotional support dialogue datasets that primarily focus on emotional states or empathic responses, this module, based on Cognitive Behavioral Therapy (CBT), aims at the core goal of "cognitive diagnosis and cognitive intervention," and provides a refined model of the cognitive distortions implicit in the expressions of those seeking help. Through this module, the system can abstract standardized cognitive labels from raw natural language, which can be used for subsequent reasoning, intervention strategy matching, and risk assessment.

[0130] Defining a three-dimensional cognitive labeling system: To achieve accurate characterization of cognitive biases, this embodiment defines a three-dimensional cognitive label tuple in the CBDF module: .in, This indicates the type of cognitive distortion, used to characterize patterns of systematic cognitive biases present in expression. Indicates the intensity of distortion, used to quantify the degree to which this cognitive bias affects an individual's judgment and behavior. This indicates the level of safety risk, used to assess whether the expression involves potential self-harm, suicide, or other high-risk psychological issues. This three-dimensional structure endows cognitive biases not only with categorical attributes but also with orderable and decision-making risk semantics, providing a foundation for subsequent modules to implement tiered interventions and safety constraints.

[0131] Cognitive Policy Reinforcement Learning Engine (CPRL): This engine is the core of cognitive decision-making in this embodiment, responsible for autonomously exploring the optimal intervention strategy in a multi-agent simulation environment. It comprises three agents: a consultant agent (…). ), the help-seeking intelligent agent ( ) and evaluator agent ( This engine utilizes a Deep Q-Network (DQN) to learn the state-to-action mapping policy. .

[0132] It's important to note that in a reinforcement learning framework, an agent learns a policy by interacting with its environment to maximize its total reward. At each time step, the agent selects an action based on its current state, and the environment provides the next state and an immediate reward based on this action. The goal of DQN is to learn a policy—a mapping from states to actions—to maximize future cumulative rewards.

[0133] Two-Stream Conditional Optimization (DSCO): This module is responsible for transferring the discrete policy knowledge learned by CPRL to the generative large language model. It employs a two-stream loss function based on a masking mechanism, enabling the model to simultaneously learn both "cognitive diagnosis" and "policy execution" capabilities.

[0134] In an optional application embodiment of this example, the specific implementation process may include the following: Step 1: Constructing a multi-agent reinforcement learning environment and state space modeling; First, initialize the reinforcement learning environment. In this environment, time steps are defined. The state space below and action space We define the state representation of reinforcement learning as follows: In order to capture the deep semantics in the dialogue, a unified text state representation is constructed in this embodiment. This representation includes not only the current textual discourse of the person seeking help. It also concatenates its corresponding cognitive tag tuple description. The specific formula is as follows:

[0135] in, This indicates a vector or text concatenation operation. This is a function that converts labels into natural language descriptions. Then, it is passed through a state encoder. Map text states to continuous low-dimensional state vectors .

[0136] The action space is defined as: a set of k standard CBT intervention strategies, such as utilizing gray areas, evidence examination, and de-catastrophizing. (Consultant agent) The goal is to select the optimal action. .

[0137] Step 2: Implement cognitive strategy reinforcement learning and hybrid reward calculation; In the simulated interaction loop, the consultant agent uses a deep Q-network (DQN) to approximate the action value function. Its decision-making strategy adopts -greedy strategy:

[0138] To address the inaccurate evaluation problem caused by the large model illusion, this embodiment innovatively designs a hybrid reward function of "rule-guided, model-corrected". At time step... The reward obtained by the agent at that time The calculation formula is as follows:

[0139] in, It is a reward for symptom improvement, given by the assessor's intelligent agent. Scoring is based on the degree to which the cognitive distortion of the person seeking help is reduced in their statements before and after the request.

[0140] It's a rule-matching reward, based on the CBT manual to validate the currently selected strategy. Does it apply to the current type of cognitive distortion? .

[0141] It is a safety constraint reward, when the risk level When the value is high, the mandatory constraint model selects a conservative safety strategy; otherwise, a high penalty is imposed. , , These are the weighting coefficients for each part.

[0142] Based on the aforementioned rewards, the DQN network parameters are updated by minimizing the temporal difference error. Target Q value The calculation formula is as follows:

[0143] in, As a discount factor, The parameters of the target network are used to obtain the optimal policy network through iterative training. .

[0144] Step 3: Data augmentation based on the optimal strategy; To address the issue of a lack of high-quality cognitive intervention responses in the raw data, a pre-trained optimal strategy is utilized. Perform data augmentation. For each requester's statement in the dataset... First use Deducing the optimal intervention action Then, perform this action. The prompt is fed into the teacher model (such as GPT-4o) to generate high-quality response text that conforms to this strategy. Therefore, the augmentation dataset CogBiasESC-PRO was constructed, where each sample pair is represented as... ,here For the context of the dialogue, Includes cognitive label sequences Responses aligned with strategy .

[0145] Step 4: Model training based on Two-Stream Conditional Optimization (DSCO); Finally, the parameters of the target large language model are fine-tuned using a two-stream conditional optimization algorithm. To prevent the generation task from overwhelming the learning gradient of the diagnostic task, a target masking mechanism is designed. Define the conditional mask loss function. as follows:

[0146] Where, if the mark Belongs to the target sequence ,but Otherwise, it is 0. The parameters to be trained in the LLM model; the final optimization objective. It combines diagnostic and interventional flows:

[0147] The first term of the formula The forced model learns to accurately output the type, intensity, and risk level of cognitive distortion, thereby improving the model's diagnostic capabilities. (The latter item...) The forced model generates intervention responses based on the optimal strategy.

[0148] Through the above steps, a psychological counseling big language model CoPoLLM was successfully trained, which has both professional cognitive diagnostic capabilities and can generate safe and effective intervention responses.

[0149] The following section provides a detailed explanation of the implementation process of the emotional dialogue generation method in a specific application scenario of an online psychological counseling support system.

[0150] Scenario Setting: Assume the user of this system is a college student named "Xiaolin" who is experiencing academic setbacks. The system loads the weights of the CogBiasESC-PRO augmented dataset constructed in this embodiment and has completed the initial configuration for the educational stress scenario.

[0151] The following examples illustrate the optional application implementations of this method, including: multi-dimensional state perception and cognitive diagnosis steps, policy decision-making steps based on the CPRL engine, intervention generation steps under dual-stream condition optimization, and dynamic adaptation and system evolution steps. In the multidimensional state perception and cognitive diagnosis step, the system first receives Xiaolin's current utterance text: "I bombed this math exam. I feel like my life is completely over, and I'll never find a good job." The cognitive bias data foundation module immediately analyzes this text. The system doesn't just capture the emotion of "frustration," but identifies deep-seated cognitive logical fallacies. The cognitive label tuple generated internally by the system is specifically instantiated as follows: Types of cognitive distortion: identified as “catastrophic thinking” (i.e., exaggerating a negative application into a hopeless catastrophe) and “generalization” (using words like “forever”).

[0152] Intensity: Determined as "severe" (absolute words such as "completely finished" and "forever" were detected).

[0153] Risk level: classified as "low risk" (no self-harm or suicidal intent detected).

[0154] In the policy decision-making step based on the CPRL engine, based on the above state, the system's cognitive policy reinforcement learning engine activates the consultant agent. In action space A, the agent calculates the Q-value of each policy. At this point, the agent finds that if it chooses a simple "emotional response" policy (such as "I am also saddened to hear this news"), although it can obtain some symptom improvement reward, it cannot obtain a high rule-matching reward, because the CBT rule states that cognitive intervention is necessary for severe catastrophic thinking. Therefore, the agent ultimately chooses the policy action with the highest Q-value: "evidence checking" (i.e., guiding the user to find evidence that does not support their catastrophic conclusion). At the same time, since the risk level is low, the safety constraint reward does not trigger the blocking mechanism, and this policy is allowed to be executed.

[0155] In the intervention generation step under dual-stream conditional optimization, the system enters the generation stage and uses parameters trained by the dual-stream conditional optimization module for inference. Internally, the system constructs prompts containing diagnostic information and selected strategies, which are then input into the LLM (Local Level Model). With the help of the target masking mechanism, the model accurately translates the strategy intent into natural language. The system ultimately outputs the following intervention response: “I hear your anxiety; failing the exam is indeed frustrating. But let’s pause and look at the facts (evidence checking strategy): Can failing one exam really determine whether you can find a job in a completely different field decades later? Are there any counterexamples that prove the two are not absolutely related?” This response not only contains empathy but, more importantly, strictly executes the strategy logic of “evidence checking.”

[0156] In the dynamic adaptation and system evolution steps, after Xiaolin received the response, the system provided the next round of feedback: "Hmm...maybe you're right, many successful people have failed courses. But I'm still worried about the makeup exam next week." The system performed another diagnostic loop and found that Xiaolin's cognitive distortion intensity significantly decreased from "severe" to "mild." The distortion type changed from "catastrophizing" to the specific "anticipatory anxiety." System state evolution: Given the decrease in intensity, the CPRL engine dynamically adjusted its strategy. In the next round of action selection, the strategy weight shifted from "strong cognitive restructuring" to the more moderate "suggestions and responses," suggesting that Xiaolin develop a specific study plan. Through the above entire process, this embodiment demonstrates how to start by identifying the specific cognitive distortion of "catastrophizing," and through strategic reasoning, ultimately achieve precise intervention and successfully reduce the user's anxiety intensity in a complete closed-loop process.

[0157] For example, such as Figure 4 The diagram shows the process reasoning of the emotional dialogue generation method.

[0158] In this embodiment, the system is divided into three interconnected stages, which sequentially complete the construction of cognitive bias data, the learning of cognitive intervention strategies, and the training of collaborative language models. This forms a closed-loop cognitive intervention system based on cognitive behavioral therapy, with reinforcement learning as the decision-making core and a large language model as the generation carrier.

[0159] In Phase One (Data Construction and Labeling), the system first collects raw dialogue data from multiple publicly available emotional support dialogue corpora and constructs a cognitive bias classification system within the framework of cognitive behavioral therapy. Through an expert labeling process, each round of help-seeker expressions is assigned structured labels such as cognitive distortion type, distortion intensity, and safety risk level, ultimately forming the CogBiasESC dataset, providing high-quality cognitive priors for subsequent decision-making and training.

[0160] In Phase Two (Cognitive Policy Reinforcement Learning Engine, CPRL), the system maps the dialogue context and its corresponding cognitive bias labels to reinforcement learning states, and constructs a cognitive policy selection network based on a DQN structure. This phase stabilizes policy updates through a combination of dual DQN and KL constraints, and defines various cognitive intervention operations conforming to the CBT principle in the action space. The evaluation agent calculates reward signals from multiple dimensions, including intervention effectiveness improvement, policy matching degree, and risk control, guiding the policy network to continuously iterate and optimize in multiple rounds of human-computer interaction, thereby learning the optimal intervention decision under different cognitive bias conditions.

[0161] In Phase 3 (CoPoLLM, Training and Optimization), the system systematically integrates the cognitive strategies optimized through reinforcement learning with the large language model's generation capabilities. Through a two-stream conditional optimization mechanism, cognitive bias information and dialogue context are used as conditional inputs to guide the model to simultaneously satisfy multiple objectives such as emotional support, cognitive correction, and safety constraints when generating responses, ultimately resulting in a CoPoLLM collaborative language model with cognitive perception and intervention capabilities.

[0162] Through the above three-stage technical approach design, an end-to-end closed-loop modeling from raw emotional dialogue data to high-quality cognitive intervention responses was achieved, providing a systematic solution for safe, interpretable, and effective intervention in complex psychological support scenarios.

[0163] Experimental results show that, compared with state-of-the-art baseline models, the CoPoLLM framework proposed in this embodiment significantly outperforms existing methods in key metrics such as diagnostic accuracy of cognitive distortions, matching degree of intervention strategies, and risk control capabilities. In particular, CoPoLLM demonstrates superior performance compared to traditional imitation learning models in the diagnosis of complex cognitive distortion scenarios.

[0164] Secondly, it enhances the safety and controllability of the large-scale psychological counseling model. By introducing an explicit safety risk reward function during the reinforcement learning stage, the model can adopt conservative and safe intervention strategies in high-risk situations (such as when identifying a high risk level of "suicidal ideation" or "severe self-denial"). This effectively avoids the dangerous responses that may lead to inducement or reinforce erroneous cognitions that existing end-to-end generative models might produce, building a safety barrier for the practical application of AI-based psychological counseling.

[0165] Finally, it enhances the interpretability of emotional support dialogues. Traditional ESC models are often a "black box," where users and regulators cannot know the logic behind a comforting statement generated by the model. CoPoLLM, however, clearly outputs its structured diagnostic results of the client's current cognitive state and the basis for its choice of intervention strategy. This highly transparent decision-making process allows mental health professionals to effectively review and supervise the model's behavior, greatly enhancing the trust in human-machine collaborative therapy.

[0166] This embodiment also provides an emotional dialogue generation device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0167] This embodiment provides an emotional dialogue generation device, such as... Figure 5 As shown, it includes: Module 501 is used to build an emotional dialogue dataset and to build multiple intelligent agents based on a pre-trained large language model, which respectively play the roles of help seeker and consultant. The dialogue acquisition module 502 is used to acquire the current dialogue text received by the first intelligent agent corresponding to the helper, and to map the target cognitive distortion classification system of the current dialogue text based on the emotional dialogue dataset. The cognition determination module 503 is used to determine the current user's cognitive bias based on the target cognitive distortion classification system; The strategy determination module 504 is used to determine the target intervention strategy of the current dialogue text by utilizing the second intelligent agent corresponding to the consultant based on cognitive bias and by using a preset hybrid reward mechanism and cognitive strategy reinforcement learning algorithm. The generation module 505 is used to generate target response text based on preset condition constraint functions and optimize the response text in the target intervention strategy.

[0168] In some alternative implementations, the dialogue acquisition module 502 includes: The parsing unit is used to parse the current dialogue text to obtain text features.

[0169] The mapping unit is used to map the target cognitive label tuple of text features based on the mapping relationship between dialogue text and cognitive label tuples concatenated from dialogue text in the emotion dialogue dataset.

[0170] The classification system determination unit is used to determine the target cognitive distortion classification system based on the target cognitive label tuple.

[0171] In some alternative implementations, the cognition determination module 503 includes: The preset classification system determination unit is used to determine the preset cognitive distortion classification system; wherein, the preset cognitive distortion classification system includes the correspondence between the cognitive distortion classification system and different cognitive distortion types, different cognitive distortion intensities and different cognitive security risk levels.

[0172] The extraction unit is used to abstract the target cognitive distortion type, target cognitive distortion intensity, and target cognitive security risk level from the target cognitive distortion classification system based on the preset cognitive distortion classification system.

[0173] The cognitive bias determination unit is used to determine the current user's cognitive bias based on the target cognitive distortion type, the target cognitive distortion intensity, and the target cognitive security risk level.

[0174] In some alternative implementations, the strategy determination module 504 includes: The initialization unit is used to initialize the dialogue learning environment of the second agent. The dialogue learning environment is an interactive loop that provides intervention strategies based on the current dialogue text.

[0175] The first determining unit is used to perform reinforcement learning on the dialogue learning environment using a second intelligent agent based on a cognitive policy reinforcement learning algorithm, and to determine the response policy value function of the current dialogue text.

[0176] The analysis unit is used to analyze the comprehensive reward function obtained at the current time step by the second agent to analyze the cognitive bias.

[0177] The solution unit is used to solve the comprehensive reward function based on the preset symptom improvement reward function, preset rule matching reward function, and preset safety constraint reward function in the hybrid reward mechanism.

[0178] The intervention strategy determination unit is used to iteratively train the deep Q-network based on the solution results of the response strategy value function and the comprehensive reward function to obtain the target intervention strategy, wherein the target intervention strategy includes at least a Socratic questioning strategy and a cognitive reconstruction strategy.

[0179] In some alternative embodiments, the device further includes: The inference module is used to infer the initial response text of the current text based on the target intervention strategy and using a pre-trained optimal strategy model.

[0180] The enhancement module is used to determine the preset teacher model and use the initial response text as a prompt for the preset teacher model to enhance the initial response text and generate the response text of the target intervention strategy.

[0181] In some alternative implementations, the generation module 505 includes: The second determining unit is used to determine the preset condition mask loss function and the preset optimization objective function in the preset condition constraint function.

[0182] The loss function solving unit is used to solve the preset conditional mask loss function by taking cognitive bias as the output condition constraint.

[0183] The objective function solving unit is used to solve the preset optimization objective function by taking the target intervention strategy as a constraint on the response text.

[0184] The generation unit is used to generate the target response text based on the solution results of the preset conditional mask loss function and the preset optimization objective function.

[0185] In some alternative embodiments, the device further includes: The update module is used to update the current dialogue text received by the first agent, update the cognitive bias according to the updated current dialogue text, and execute the step of determining the target intervention strategy of the current dialogue text by the second agent based on the cognitive bias, using a hybrid reward mechanism and a cognitive policy reinforcement learning algorithm.

[0186] The emotional dialogue generation apparatus provided in this embodiment of the invention can execute the emotional dialogue generation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.

[0187] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.

[0188] The following is a detailed reference. Figure 6 This diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from memory 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of the electronic device. The processor 601, ROM 602, and RAM 603 are interconnected via a bus 6604. An input / output (I / O) interface 605 is also connected to bus 604.

[0189] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0190] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a memory 608, or installed from a ROM 602. When the computer program is executed by the processor 601, it performs the functions defined in the emotional dialogue generation method of the embodiments of the present invention.

[0191] Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present invention.

[0192] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the emotional dialogue generation method shown in the above embodiments is implemented.

[0193] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0194] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for generating emotional dialogue, characterized in that, The method includes: We constructed an emotional dialogue dataset and built multiple intelligent agents based on a pre-trained large language model, which respectively played the roles of help seeker and consultant. Obtain the current dialogue text received by the first intelligent agent corresponding to the person seeking help, and map the target cognitive distortion classification system of the current dialogue text based on the emotional dialogue dataset; Based on the aforementioned target cognitive distortion classification system, the cognitive bias of the current user is determined; Based on the cognitive bias, the second intelligent agent corresponding to the consultant determines the target intervention strategy for the current dialogue text using a preset hybrid reward mechanism and a cognitive strategy reinforcement learning algorithm. Based on a preset condition constraint function, the response text in the target intervention strategy is optimized to generate the target response text.

2. The method according to claim 1, characterized in that, The target cognitive distortion classification system based on the emotional dialogue dataset that maps the current dialogue text includes: The current dialogue text is parsed to obtain text features; Based on the mapping relationship between dialogue text and cognitive label tuples concatenated from dialogue text in the emotional dialogue dataset, the target cognitive label tuples of the text features are mapped. The target cognitive distortion classification system is determined based on the target cognitive label tuple.

3. The method according to claim 1, characterized in that, The determination of the current user's cognitive bias based on the target cognitive distortion classification system includes: A preset cognitive distortion classification system is determined; wherein, the preset cognitive distortion classification system includes the correspondence between the cognitive distortion classification system and different types of cognitive distortion, different intensities of cognitive distortion, and different levels of cognitive security risk; Based on a pre-defined cognitive distortion classification system, the target cognitive distortion type, target cognitive distortion intensity, and target cognitive security risk level in the target cognitive distortion classification system are abstracted. The cognitive bias of the current user is determined based on the target cognitive distortion type, the target cognitive distortion intensity, and the target cognitive security risk level.

4. The method according to claim 1, characterized in that, The step of utilizing the second intelligent agent corresponding to the consultant to determine the target intervention strategy for the current dialogue text based on the cognitive bias, using a preset hybrid reward mechanism and a cognitive strategy reinforcement learning algorithm, includes: Initialize the dialogue learning environment of the second agent, wherein the dialogue learning environment is an interactive loop that provides intervention strategies based on the current dialogue text; The second intelligent agent performs reinforcement learning on the dialogue learning environment based on the cognitive policy reinforcement learning algorithm to determine the response policy value function of the current dialogue text; The second agent is used to analyze the comprehensive reward function acquired at the current time step for the cognitive bias; Based on the preset symptom improvement reward function, preset rule matching reward function, and preset safety constraint reward function in the hybrid reward mechanism, the comprehensive reward function is solved. Based on the solution results of the response strategy value function and the comprehensive reward function, the deep Q network is iteratively trained to obtain the target intervention strategy, wherein the target intervention strategy includes at least a Socratic questioning strategy and a cognitive reconstruction strategy.

5. The method according to claim 1, characterized in that, Before generating the target response text in the target intervention strategy based on a preset condition constraint function, the method further includes: Based on the target intervention strategy, the initial response text of the current text is inferred using a pre-trained optimal strategy model; A preset teacher model is determined, and the initial response text is used as a prompt for the preset teacher model to enhance the initial response text and generate a response text for the target intervention strategy.

6. The method according to claim 5, characterized in that, The step of optimizing the response text in the target intervention strategy based on a preset condition constraint function to generate the target response text includes: Determine the preset conditional mask loss function and the preset optimization objective function in the preset conditional constraint function; The cognitive bias is used as an output condition constraint to solve the preset conditional mask loss function. The target intervention strategy is used as a constraint on the response text to solve the preset optimization objective function; The target response text is generated based on the solution results of the preset conditional mask loss function and the preset optimization objective function.

7. The method according to claim 1, characterized in that, The method further includes: The first agent updates the current dialogue text received, updates the cognitive bias based on the updated current dialogue text, and executes the step of the second agent determining the target intervention strategy for the current dialogue text based on the cognitive bias using a hybrid reward mechanism and a cognitive policy reinforcement learning algorithm.

8. An emotional dialogue generation device, characterized in that, The device includes: The module is used to build an emotional dialogue dataset and to build multiple intelligent agents based on a pre-trained large language model, which play the roles of help seeker and consultant respectively. The dialogue acquisition module is used to acquire the current dialogue text received by the first intelligent agent corresponding to the helper, and to map the target cognitive distortion classification system of the current dialogue text based on the emotional dialogue dataset. The cognition determination module is used to determine the current user's cognitive bias based on the target cognitive distortion classification system; The strategy determination module is used to determine the target intervention strategy of the current dialogue text by utilizing the second intelligent agent corresponding to the consultant based on the cognitive bias, and by using a preset hybrid reward mechanism and a cognitive strategy reinforcement learning algorithm. The generation module is used to generate target response text by optimizing the response text in the target intervention strategy based on preset conditional constraint functions.

9. An electronic device, characterized in that, include: A memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to perform the method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the method of any one of claims 1 to 7.