A Psychological Companion Dialogue Method and Storage Medium Based on Multi-Agent Collaboration
Through multi-agent collaboration and multi-modal analysis technology, the problems of insufficient strategy diversity and personalization in the existing AI emotional dialogue system are solved, and high-quality and personalized multi-round dialogue content generation are achieved, improving user experience.
Patent Information
- Application Number
- CN202411633681.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-11-14
AI Technical Summary
The existing AI emotional dialogue methods have insufficient strategies and are not targeted when generating replies, making it difficult to optimize in real time based on user feedback, resulting in insufficient flexibility and personalization of reply content.
Multi-agent collaborative work and multi-modal analysis technology are adopted to generate high-quality multi-round dialogue content through shared memory modules, strategy planning modules and dialogue content generation modules, combined with online reinforcement learning methods.
It improves the flexibility and personalization of reply, enhances the diversity and fit of conversation content, and can optimize strategies in real time based on user feedback to improve user experience.
Smart Images

Figure CN119599029B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language understanding in the human-computer dialogue system of artificial intelligence technology, and specifically relates to a psychological companionship dialogue method and storage medium based on multi-agent collaboration. Background Technique
[0002] In recent years, the mental health problems of contemporary teenagers have attracted wide attention. With the increase of social pressure and the rapid development of information technology, the factors affecting the mental health of teenagers have become more and more complex and diverse, including academic pressure, interpersonal relationships, family environment, etc. Teenagers with mental problems often do not want to actively talk to a psychologist or seek help. Therefore, it has become an urgent need to use advanced artificial intelligence technology to monitor the mental state on campus and provide intelligent psychological counseling services for at-risk objects. With the development of the field of artificial intelligence, there have already been multiple dialogue applications combining deep learning technology with the psychological field, as follows:
[0003] Patent CN 111564202B, "Psychological Counseling Method, Psychological Counseling Terminal and Storage Medium Based on Human-Computer Dialogue", mainly uses a human-computer dialogue system and an emotion recognition model to conduct psychological counseling for users. This method uses the target voice and dialogue information input by the user to identify the dialogue emotion from two dimensions: emotional valence and emotional arousal, and obtains and executes a response plan matching the user through a psychological conversation technology model. However, the acquisition of the response plan adopts the method of matching preset target event conversation templates and communication steps in the database, so the response strategy is not flexible and targeted enough.
[0004] Patent CN 117932041 A, "Emotion Support Dialogue Generation Method, System and Device Based on Chain of Thought Reasoning", mainly uses chain of thought reasoning to assist in the generation of emotion support dialogues. This method obtains the user's emotional state information and response strategy by constructing an emotion chain of thought and a strategy chain of thought. However, the emotional information provided by the dialogue text is limited, and more accurate analysis results need to be combined with tone, expression, etc. This method only involves the dialogue history information when conducting emotion chain of thought reasoning, and does not perform real-time analysis on the user's speech intonation and facial expression. In addition, this method only uses prompt words to drive the large model to generate responses, although it reduces the training cost, it is difficult to guarantee the execution effect of the strategy.
[0005] Patent CN 114999610 A, "Method for Constructing a Dialogue System with Emotion Perception and Support Based on Deep Learning", mainly uses a deep learning model to construct a dialogue system with emotion perception and support. This method generates an adaptive dialogue strategy by introducing a large model, and replies to the user's dialogue or makes recommendations according to the generated strategy. However, this method can only make a one-time strategy selection, and only one strategy can be selected in a single round. Moreover, it does not combine the user's psychological state information for dialogue generation, resulting in insufficient emotional support for users. In addition, the large model used for strategy prediction in this method is an offline version obtained by pre-training and cannot be updated in real time according to the dialogue feedback, resulting in not being well-suited to the individual needs of users in a real dialogue environment.
[0006] Patent CN117972075A, "Emotional Dialogue Generation Method with Collaboration between Mind and Language Agents", mainly uses the collaborative operation of multiple agents such as state analysis, strategy planning, and language generation to improve the system's ability to conduct emotional conversations with users, and ensures the coherence and personality consistency of the replies by adding role-setting prompt words. In addition, this method also uses similarity retrieval to match the training corpus to generate reply strategies. However, due to the relatively fixed training corpus, the strategy generation cannot be innovated according to the actual situation, resulting in insufficient diversity and richness of the strategy group. In addition, when generating the reply text, this method requires a single model to complete the generation tasks of multiple strategies, resulting in difficulty in achieving the desired effect when executing more complex reply strategies and difficulty in optimizing the subsequent model according to the feedback.
[0007] Most current AI emotional dialogue methods use similarity matching or simple reasoning for strategy planning, and basically drive a large language model by prompt words to generate reply text. Although this reduces the training and operation costs, it also significantly reduces the diversity, pertinence, and effectiveness of emotional conversations. Therefore, this patent proposes to construct a multi-agent system, which collaborates with the strategy planning agent, each dialogue state update agent, and each dialogue generation agent through memory sharing among the agents to generate replies, conduct high-quality multi-round conversations with users, and provide them with psychological counseling and emotional support. At the same time, multi-modal evaluation is used to create a user psychological portrait, thereby significantly improving the accuracy of user emotion analysis and the dialogue ability in the scenario of psychological counseling. In addition, an online reinforcement learning method is introduced to train the agents for reply strategy planning, improving the flexibility and richness of the replies. Summary of the Invention
[0008] The purpose of the present invention is to solve the problem of how to construct a psychological companionship dialogue system that can provide personalized, high naturalness, and accurate emotion recognition by integrating the collaborative work of multi-agents and multi-modal analysis technology, thereby effectively improving the user experience and meeting their psychological support needs.
[0009] To achieve the above object, the present invention adopts the following technical means:
[0010] A psychological companionship dialogue method based on multi-agent collaboration, comprising the following steps:
[0011] Step 1: Conduct multi-modal emotion assessment on the user input to obtain a description of the user's emotional state;
[0012] Step 2: According to the user's question and the description of the user's emotional state, retrieve relevant dialogue states and external knowledge from the shared memory module;
[0013] Step 3: According to the retrieved dialogue states and external knowledge, the strategy planning module generates a reply strategy to obtain a strategy group;
[0014] Step 4: The dialogue content generation module calls the corresponding vertical domain agent according to the reply strategy generated in Step 3 to generate specific reply content;
[0015] Step 5: Store the user's question and the model's reply in the short-term memory of the shared memory module in this round and update the long-term memory.
[0016] In the above solution, Step 1 includes the following steps:
[0017] Step 1.1: Encode the text input by the user using BERT to obtain a semantic vector;
[0018] Step 1.2: Extract features from the user's input voice and image to obtain an emotion feature vector;
[0019] Step 1.3: Combine the semantic vector in Step 1.1 and the emotion feature vector obtained in Step 2.2, and through a multi-modal emotion assessment model, obtain a description of the user's emotional state;
[0020] In the above solution, Step 2 includes the following steps:
[0021] Step 2.1: Use the BERT encoder to encode the user's question and the result of Step 1 to obtain a retrieval vector. At this time, the retrieval vector will be used to query the vector database in the shared memory module, which stores the vector representations of the dialogue history and external knowledge;
[0022] Step 2.2: Based on the result of Step 2.1, perform semantic retrieval on the vector database to obtain relevant dialogue state information and external knowledge;
[0023] Step 2.3: At the same time, based on the user input, perform literal retrieval on the relational database and the graph database to obtain dialogue state information with the same keywords as the user input;
[0024] Step 2.4: Perform weighted sorting on the results of Steps 2.2 and 2.3, and select the top k most relevant texts;
[0025] In the above solution, the implementation of the shared memory module includes the following steps:
[0026] Step A1: Initialize the shared memory module, including setting up short-term memory and long-term memory;
[0027] Step A2: Store the context background and conversation history in the short-term memory;
[0028] Step A3: Store the external knowledge documents and conversation status in the long-term memory;
[0029] Step A4: When the conversation is in progress, call the content of the short-term memory and splice it with the input;
[0030] Step A5: After the reply generation is completed, write the user input and system output of the current round into the conversation history;
[0031] Step A6: The conversation status includes the conversation history and the processed multi-level conversation memory;
[0032] Step A7: The multi-level conversation memory includes the original conversation at the lowest level, the paragraph summary at the second level, the event summary at the third level, the emotion summary at the fourth level, and the user profile at the highest level;
[0033] Step A8: After each round of conversation generation, call the relevant agent to update the conversation status at each level;
[0034] Step A9: After the update is completed, call the data caching and management module for processing, including updating the graph database, relational database, and vector database.
[0035] In the above solution, Step 3 includes the following steps:
[0036] Step 3.1: The strategy planning module generates multiple possible reply strategies based on the input conversation status;
[0037] Step 3.2: Input the results of Step 3.1 into the conversation content generation module to generate specific reply content, and the reply content of the optimal strategy is directly used as the model reply for this round of conversation;
[0038] Step 3.3: Transmit the specific reply content generated in Step 3.2 back to the strategy planning module, use the preference model to sort the results of Step 3.1, determine the optimal strategy according to the sorting to form a strategy group, and the strategy planning module performs strategy iteration and parameter update.
[0039] In the above solution, Step 3.1 specifically includes the following steps:
[0040] Step 3.1.1, Policy Planning Module π t Generate multiple policies based on the input dialogue state. Among them, the optimal policy is denoted as The complete policy sequence is denoted as The policy planning module is abbreviated as the agent;
[0041] Step 3.1.2, The dialogue state is called from the shared memory, including the dialogue history, dialogue event summary, user mental state, user emotion, and user profile;
[0042] Each output policy needs to include the name of the agent called and the information passed to it. The complete input prompt template is:
[0043] {
[0044] User question in this round: [User question], user's current emotion is [Observed mood];
[0045] Current dialogue situation: [memory summary], relevant knowledge: [Useful knowledge];
[0046] The agents of this system are: [Agent Name: {Usage, Input Parameters},…]
[0047] You are an intelligent scheduling agent. You need to plan how to use one or more [AgentNames] to respond to the user's question in this round based on the above information, and provide emotional support to relieve their negative emotions. Please output your response plan in the following format, in order, and the plan should include at least one policy;
[0048] Used agent:
[0049] Task description:
[0050] }
[0051] In the above solution, step 3.3 specifically includes the following steps:
[0052] 3.3.1, Use the preference model to sort the results of the policy sequence generated in step 3.1 to determine the optimal policy. The preference model is based on the Bradley-Terry model, and its preference score calculation formula is:
[0053] P(a 1 >a 2 |x, a 1 , a 2 ) = σ(r * (x,a1 ) - r * (x, a 2 ))
[0054] where r * (x, a) is the included reward function, σ(·) is the Sigmoid function, x represents the input dialogue state, a represents the action, a 1 , a 2 represents different actions; a 1 > a 2 means that action a 1 is considered better or more preferred than action a 2 ;
[0055] Step 3.3.2. Supplement the preference ranking sequence into the policy preference dataset and update the parameters of the policy planning module through direct preference optimization. The optimized loss function is:[[]]
[0056]
[0057] where a w = a win , a l = a lose , a w represents the action that leads to win or victory, a l represents the action that leads to loss or defeat;
[0058] θ represents the learned policy parameter, π θ represents a parameterized policy, where θ is a set of parameters;
[0059] π0 represents a reference policy;
[0060] σ(·) is the sigmoid function, which maps real values to the interval (0, 1)(0, 1);
[0061] η is a hyperparameter;
[0062] π0(a * |x) is the probability of taking action a * in state x;
[0063] Step 3.3.3. Select the optimal new policy π t+1 by maximizing the expected return while considering the diversity of the policy. The process of selecting the optimal policy is:[[]]
[0064]
[0065] where argmax πDenotes selecting the strategy that can maximize the expected return among all possible strategies π;
[0066] Denotes taking the expectation over all possible dialogue states x;
[0067] Denotes taking the expectation over all possible actions a given the strategy π and the dialogue state x;
[0068] η is a hyperparameter used to balance strategy selection and strategy diversity;
[0069] Denotes the strategy π t The KL divergence of the strategy π relative to the initial strategy π0, used to measure the difference between the strategy π t And π0.
[0070] In the above solution, step 4 includes the following steps:
[0071] Step 4.1, The scheduling agent calls the required vertical domain agent according to the reply strategy;
[0072] Step 4.2, The vertical domain agent selected in step 4.1 generates the reply content of a specific strategy according to the incoming relevant information;
[0073] Step 4.3, The scheduling agent calls the vertical domain agent to continue generating the reply content of the next strategy, and splices it after the result of step 4.2;
[0074] Step 4.4, Loop steps 4.1 to 4.3 until all strategies in the strategy group are used, and obtain the complete model reply for this round.
[0075] In the above solution, step 5 includes the following steps:
[0076] Step 5.1, Store the user's question and the model reply for this round in the dialogue history record of the short-term memory;
[0077] Step 5.2, Update the dialogue state in the long-term memory according to the update result of step 5.1, including dialogue history summary, event summary, emotion summary and user portrait;
[0078] Step 5.3, Store the update result of step 5.2 in the vector database, relational database and graph database.
[0079] The present invention also provides a storage medium. When a processor executes the program in the storage medium, it implements the described method for psychological companionship dialogue based on multi-agent collaboration.
[0080] Because the present invention adopts the above technical means, it has the following beneficial effects:
[0081] By introducing the methods of multi-agent collaboration and human-feedback-based reinforcement learning, the present invention solves the problems of unstable quality, difficulty in optimization, and lack of "humanity" in the responses generated by a single dialogue large model, making the responses generated by the psychological companionship system more relevant to the conversation content, enabling targeted optimization of the response models of individual generation strategies, and allowing the system to adjust to a better conversation style according to user feedback during operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 Block diagram of a psychological companionship dialogue system based on multi-agent collaboration;
[0083] Figure 2 Shared memory module;
[0084] Figure 3 Dialogue state;
[0085] Figure 4 Knowledge retrieval module;
[0086] Figure 5 Conceptual diagram of the policy planning module;
[0087] Figure 6 Frame of the t-th round of policy planning agent operation (black arrow) and training (orange arrow);
[0088] Figure 7 Dialogue content generation module
[0089] Figure 8 Call the vertical domain agent to generate responses. DETAILED DESCRIPTION OF THE INVENTION
[0090] The following will give a detailed description of the embodiments of the present invention. Although the present invention will be described and explained in conjunction with some specific embodiments, it should be noted that the present invention is not limited to these embodiments only. On the contrary, any modifications or equivalent substitutions made to the present invention shall be covered within the scope of the claims of the present invention.
[0091] In addition, for a better illustration of the present invention, numerous specific details are given in the following detailed description. Those skilled in the art will understand that the present invention can be implemented without these specific details.
[0092] The present invention relates to a dialogue system based on multi-agent collaboration in the field of psychological companionship, as Figure 1 shown, including a shared memory module, an information retrieval module, a policy planning module, and a dialogue content generation module. The following will introduce them in detail according to the modules:
[0093] I. Shared memory module
[0094] Technical problems considered: Information transmission is required for multi-agent collaboration; context information needs to be invoked for dialogue responses; a large amount of diverse data needs to be stored after the system is put into use.
[0095] Solutions:
[0096] Store in an orderly manner according to long-term and short-term memories and different data types to improve the retrieval speed;
[0097] All memories generated during the operation of the agents are stored in this module. When an agent needs to invoke during operation, it can directly access all memories, preventing information imbalance among agents and eliminating the transmission step.
[0098] Introduce multiple layers of dialogue states to hierarchically summarize and extract information from the original dialogue; use three data structures to store multi-level states to improve the retrieval efficiency.
[0099] The function of this module is to save the dialogue history, external knowledge, and useful information generated during the system operation for each agent to invoke. The shared memory is divided into short-term memory and long-term memory. The content of the short-term memory includes the context background and the dialogue history record, both of which are stored in text form. The context background is usually set by the user when starting a dialogue, such as role, scene preset, problem background, etc. The context background generally does not change during multiple rounds of dialogue to improve the model's understanding of the dialogue scenario, thereby enhancing the language quality and topic relevance of the system's response. The dialogue history record stores the complete original dialogue, providing historical information to each agent to make the model's response smoothly connect with the previous text. If the current dialogue round is the t-th round, the dialogue history record is the dialogue data of the previous t - 1 rounds, where each round includes the user's question, the system's response, and the corresponding strategy group. During the dialogue, the system will call the content of the short-term memory and splice it with the input; after the response generation is completed, the user input and the system output of this round will be written into the dialogue history record.
[0100] The long-term memory mainly involves external knowledge documents and dialogue states. External knowledge includes scientific common sense, psychological field knowledge, etc. After the knowledge document data is segmented, the corresponding vector representation is obtained through embedding encoding and stored in the vector database. During the dialogue generation process, each agent will call relevant documents from the vector database as needed. The dialogue state includes the dialogue history record and the processed multi-level dialogue memory.
[0101] Regarding the multi-level dialogue memory:
[0102] The lowest level is the original dialogue;
[0103] The paragraph summary of the second layer is obtained by summarizing multiple rounds of dialogue by the "dialogue summary agent" (one user question + one model response is one round);
[0104] The third layer is the event summary, which includes three parts: entity extraction, event extraction, and time extraction. It is mainly obtained by analyzing the content of the summary in the lower layer using the "entity extraction agent" and the "event analysis agent".
[0105] The fourth layer is the emotion summary, which includes two parts: emotion analysis and mental state. They are obtained by using the "emotion analysis agent" and the "multimodal psychological assessment agent" to analyze events respectively. The specific content is: the user's attitude towards a certain topic or event, the current mental state, etc.
[0106] The highest-level dialogue state is the user profile, which summarizes information about the user's personality traits, thinking patterns, topics of interest, viewpoints, etc. from the lower-level memory. It is a relatively comprehensive overview of the user image. The user profile not only includes the description of the current dialogue user but also the relevant information of the entities that appear in the dialogue.
[0107] The dialogue state will be stored in the vector database, relational database, and graph database in three forms. When the system conducts policy planning and dialogue content generation, it will retrieve and call the content of these three databases as needed. After each round of dialogue generation, relevant agents will be called to cooperate in updating the dialogue state of each layer in the order from bottom to top. After the update is completed, the data cache and management module will be called for processing: 1) Using the "user profile" as the center point of the graph, connect relevant events, emotions, and specific memories and store them in the graph database; 2) Establish a user list in the relational database, create multiple data tables for each user to store the dialogue state in layers, and the content of each layer is the index of the next layer; 3) Encode the content of each layer (including the original dialogue) using BERT and store it in the vector database.
[0108] Information Retrieval Module
[0109] The large language model that constitutes the agent uses corpus for fixed scenarios or tasks during pre-training and fine-tuning. Therefore, when chatting with users and encountering topics, scientific questions, or real-time data questions outside the corpus domain, the model's responses perform poorly and are prone to fabricating answers. To improve the situation of model hallucinations and enhance its ability to correctly understand different topics and complete tasks in different domains, this system designs an information retrieval module. This module queries based on the dialogue input, first generates semantic vectors through an encoder, and then retrieves relevant text information from the vectorized information to assist the agent in responding to user questions.
[0110] This module is divided into two parts: semantic retrieval and literal retrieval. Semantic retrieval is based on the retrieval vector encoded by BERT for the user input, and vector retrieval is performed on the vector database to obtain conversation state information related to the user input content (including the original conversation, event summary, relevant characters, user emotions, psychological state, portrait) and external knowledge. To ensure the effectiveness of search matching, external documents and conversation states also need to be embedded and encoded using the BERT encoder. Literal retrieval is based on the retrieval vector obtained by BM25 encoding for the user input, and keyword retrieval is performed on the content of the relational database and the graph database encoded using the same method to obtain conversation state information with the same keywords as the user input. Then, weighted sorting is performed according to the similarity results of the two searches, and the top k texts with the highest similarity are selected. Finally, the selected document vectors are converted into natural language texts and passed to the intelligent agent to fill in the corresponding parts of the prompt template to assist it in performing tasks.
[0111] Policy Planning Module
[0112] Existing problems:
[0113] 1. The existing policy generation module directly uses the large language model generate content or Markov chain predict methods to generate policy groups; the former has too poor interpretability and is difficult to optimize; the latter generated policy groups are difficult to be targeted for different conversation scenarios and are easily singular.
[0114] 2. The text generated by the existing policy generation module has very general human-like performance (not very like the spoken conversations in daily human chats). If only adjusting the prompt to improve this point, a large number of experiments are required, and the effect is very small.
[0115] The strategy planning module is the core module of this system. Its main function is to generate strategies for the system to reply to users and guide the invocation and output of each agent. The initial input of the module is a piece of text, which mainly includes three parts: the natural language text spliced from the dialogue history and relevant memories (referring to the long-term memories related to this dialogue except the original dialogue history, such as events / characters / emotions involved in the previous text. "Splicing" means splitting the input content and filling it into the prompt), the external knowledge retrieved, and the description of the user's emotional state. The strategy planning large model needs to break down the user's question in the current round, and combine the chat history, user emotions, and relevant psychological knowledge to determine the method of the system's reply (such as empathy, cognitive behavioral therapy, recommendation, etc., corresponding to the names of the agents to be executed) and the content to be included (such as replying to specific scientific questions, soothing the user's emotions, etc.). The "reply strategy" output by this module is a text with a fixed format, which describes the specific content of each task in the strategy group and the agents planned to be invoked, and lists them in the execution order. The output result will be passed to the dialogue content generation module for the generation of specific reply content.
[0116] Currently, the existing emotional support multi-turn dialogue systems mainly use Markov chain processes or multi-class neural networks for strategy planning. However, the generated strategies are simply represented by keywords, and the subsequent generation effect of the reply content is not good, and it is difficult to improve the humanity of the generated content. To improve this situation, this patent uses the method of online reinforcement learning to construct a strategy planning agent, that is, the online RLHF (Reinforcement Learning from Human Feedback) method. By introducing human preference feedback (implemented using a preference model) during the training process, the agent continuously learns how to generate replies that better meet human preferences and improve the ability to provide appropriate emotional support and solutions to users. The overall framework is as Figure 6 shown. Assuming that the current is the t-th round of dialogue, the specific operation and update steps are as follows:
[0117] 1. First, the strategy planning agent (π t ) generates multiple strategies based on the input dialogue state. Among them, the strategy judged to be the optimal one is denoted as The complete strategy sequence is denoted as The dialogue state is called from the shared memory, including the dialogue history, dialogue event summary, user mental state, user emotions, and user portrait. Each output strategy needs to include the name of the agent to be invoked and the information passed to it. The complete input prompt template is:
[0118] This round of user question: [User question], the user's current mood is [Observed mood]. Current conversation situation: [memory summary], relevant knowledge: [Useful knowledge]. The agents in this system are: [Agent Name: {Usage, Input Parameters},...]
[0119] You are an intelligent scheduling agent. You need to plan how to use one or more [AgentNames] to respond to this round of user questions based on the above information, and provide emotional support to relieve their negative emotions. Please output your reply plan in the following format, in order (the plan should include at least one strategy).
[0120] Used agent:
[0121] Task description:
[0122] 2. Next, the strategy sequence Π t Enter the dialogue content generation module to generate specific responses. Each strategy generates an optimal response as a representative (the strategy planning agent (large model) has been pre-trained, and the optimal response is directly given by the model). The optimal strategy 's response is directly output, which is the model's response for this round of dialogue. Then, all responses (including responses) are sent back to this module and sorted using the preference model. The preference model is based on the Bradley-Terry (BT) model and is supervised and fine-tuned using a pre-collected response preference dataset, and is not updated in the system. The model generates a preference score for each input response, and then sorts them from largest to smallest according to this score to simulate sorting responses according to human preferences. The preference model satisfies P(a 1 >a 2 |x, a 1 , a 2 ) = σ(r * (x,a 1 ) - r * (x, a 2 ))). Where r * (x,a) is the included reward function, and σ(·) is the Sigmoid function. In this system, x represents the input dialogue state, and a represents the action (corresponding to the name of the dialogue generation agent). The loss function optimized during the training of the preference model is: Where a w is the action selected as the best response, a l is the action not selected as the best response, Denote the expectation as $E$, $x$ represents the input dialogue state, $\log\sigma(r θ (x,a w ) - r θ (x,a l )) measures the preference degree of the agent to select $a w as the best response rather than $a l . is the strategy preference dataset.
[0123] 3. The preference ranking sequence is added as new data to the strategy preference dataset $D$. Update the parameters of the policy planning agent through Direct Preference Optimization (DPO). The input $\pi_0$ of this process is the initial policy agent, obtained through supervised fine-tuning. The optimized loss function is:
[0124]
[0125] where $a w = a win , $a l = a lose . Parameter update is performed through MLE. The process of selecting the optimal new policy is:
[0126]
[0127] where $\arg\max_{ π denotes selecting the policy that can maximize the expected return among all possible policies $\pi$;
[0128] denotes taking the expectation over all possible dialogue states $x$;
[0129] denotes taking the expectation over all possible actions $a$ given the policy $\pi$ and the dialogue state $x$;
[0130] $\eta$ is a hyperparameter used to balance policy selection and policy diversity;
[0131] denotes the KL divergence of the policy $\pi t relative to the initial policy $\pi_0$, used to measure the difference between the policy $\pi t and $\pi_0$.
[0132] Obtain the policy planning agent $\pi t+1 for the $(t + 1)$-th round. The dialogue content generation module
[0133] The main advantages of this module are:
[0134] 1. Utilize multiple agents to collaborate for natural language text generation in specific scenarios.
[0135] 2. To improve the generation effect of a single agent, the agent with a single policy can be optimized according to the test responses. If it is necessary to add or adjust the policy, it can be efficiently optimized to improve the fit between the complete response and the user's question and the policy.
[0136] Currently, using a prompt to drive a large model to generate a complete answer containing multiple response strategies has mediocre effects, and it can only be optimized by continuously adjusting the prompt and observing the output situation, with very low optimization efficiency. Therefore, we consider training a specialized large model for each response strategy to separately generate the response content for each strategy and compose a complete reply.
[0137] If these response contents are directly concatenated, it may result in a lack of coherence between the first and the second sentences. Therefore, we use a scheduling agent to schedule the specialized agents; and during the generation process, the semi-finished responses are fed into the specialized agents, and the prompt is used to drive them to generate coherent responses that conform to the required strategies based on the semi-finished responses.
[0138] In recent years, the fine-tuning technology of large language models has become increasingly mature, and fine-tuning models based on closed-source large language models such as GPT have been applied in various fields. Among them, guiding the model to complete tasks in a specified manner through a prompt template is the simplest and most common method. However, when the prompt template faces multiple tasks in specific scenarios, the generated effects are very limited and it is difficult to optimize. Therefore, to improve the dialogue quality of the psychological companionship system, this patent proposes to construct a dialogue content generation module with multi-agent collaboration, where each agent specializes in a subtask and only generates the response content for one strategy.
[0139] In this module, the scheduling agent will call the required vertical domain agents according to the given policy to have a high-quality dialogue with the user, aiming to provide psychological counseling and emotional support to the user. The vertical domain agents of this system include empathy agents, CBT (Cognitive Behavior Therapy) agents, recommendation agents, psychological counseling agents, common question answering agents, chatting and companionship agents, etc. During the training process, by collecting the corpora of each task and using the QLoRA (Quantized Low-Rank Adapter) method to fine-tune the benchmark large model, agents for each task are obtained. Figure 7 Shows the overall framework of this module, and the specific implementation uses the LangGraph module.
[0140] At the beginning of each round of conversation, the strategy planning agent usually first calls the multi-modal emotion assessment to obtain the real-time user emotion and update the memory related to the user's mental state. Then it directly calls the vertical domain agent according to the reply. Agent A first completes the prompt template by itself according to the incoming relevant information, and then the corresponding large language model (LLM) uses the chain of thought or the method of calling external tools to generate a reply in combination with the result of knowledge retrieval. Since each vertical domain agent only focuses on one strategy, a complete reply usually requires the cooperation of multiple agents to generate. For example, if the reply strategy is "empathy - CBT - recommendation", the scheduling agent will first call the empathy agent to generate the empathy reply content; then it will pass the original conversation information and the empathy reply to the CBT agent together to generate the CBT reply, which is spliced after the empathy reply; then it will pass the current reply as part of the input to the recommendation agent, and use a specific Prompt to guide it to generate the recommendation content, and splice to get the complete model reply for this round. The model reply will be stored in the "conversation history record" of the short-term memory together with the user's question for this round; at the same time, it will be passed to the strategy planning module and the conversation state update module respectively for the iteration of the strategy planning agent and the update of the long-term memory. In the above, the specific Prompt is Figure 8 the Prompt Template in it. This template is a generalized version for the vertical domain agents. In each agent, the prompt will be filled with some content. For example, the prompt of the recommendation agent is to supplement the name, function, received parameters, etc. of this agent on the generalized version.
[0141] The invention relates to a psychological companionship conversation method and a storage medium based on multi-agent collaboration, aiming to build a psychological companionship conversation system that can provide personalized, highly natural and accurate emotion recognition by integrating multi-agent collaborative work and multi-modal analysis technology. The following is an analysis of the main features and advantages of this technical solution:
[0142] Multi-modal emotion assessment: By combining multi-modal emotion assessment of text, voice and image, it can more accurately identify the user's emotional state and provide more considerate psychological support for the conversation.
[0143] Shared memory module: By using the distinction between short-term and long-term memory, it effectively stores and retrieves conversation history, external knowledge and user portraits, improving the coherence and personalization of the conversation.
[0144] Strategy planning module: Adopt the online reinforcement learning method to train the agent for reply strategy planning, enhancing the flexibility and diversity of the reply, and improving the pertinence and humanization of the conversation content.
[0145] Dialogue content generation module: By the collaborative work of multiple vertical domain agents, the generation effect of a single agent is improved, and the agent with a single strategy can be optimized according to the test responses, thereby enhancing the fit between the complete response and the user's question and strategy.
[0146] Construction of user profile: Through multi-level summarization of the dialogue history (including the original dialogue, paragraph summary, event summary, emotion summary, and user profile), a comprehensive user profile is constructed, which helps the system better understand the user's needs and provide personalized services.
[0147] Information retrieval module: By combining semantic retrieval and literal retrieval, the ability to accurately understand and answer the user's questions is improved, especially when dealing with topics outside the domain or scientific questions.
[0148] Real-time optimization and iteration: The system continuously optimizes the strategy planning agent through online learning, enabling it to adjust the dialogue style according to user feedback, and improving the adaptability and effectiveness of the dialogue system.
[0149] In summary, through multi-agent collaboration and multi-modal analysis technologies, the present invention effectively solves some problems in the existing psychological companion dialogue system, such as insufficient flexibility and diversity of response strategies, and low fit between the dialogue content and the user's needs, thereby enhancing the user experience and the effect of psychological support.
Claims
1. A psychological companionship dialogue method based on multi-agent collaboration, characterized in that It includes the following steps: Step 1: Conduct multi-modal emotion assessment on the user input to obtain a description of the user's emotional state; Step 2: Retrieve relevant dialogue states and external knowledge from the shared memory module according to the user's question and the description of the user's emotional state; Step 2.1: Use the BERT encoder to encode the user's question and the result of Step 1 to obtain retrieval vectors. At this time, the retrieval vectors will be used to query the vector database in the shared memory module, which stores the vector representations of the dialogue history and external knowledge; Step 2.2: Based on the result of Step 2.1, perform semantic retrieval on the vector database to obtain relevant dialogue state information and external knowledge; Step 2.3: At the same time, based on the user input, perform literal retrieval on the relational database and the graph database to obtain dialogue state information with the same keywords as the user input; Step 2.4: Perform weighted sorting on the results of Step 2.2 and Step 2.3, and select the top k most relevant texts; Step 3: According to the retrieved dialogue states and external knowledge, the strategy planning module generates reply strategies to obtain a strategy group; Step 4: The dialogue content generation module calls the corresponding vertical domain agent according to the reply strategy generated in Step 3 to generate specific reply content; Step 5: Store the user's question and the model's reply in the short-term memory of the shared memory module in this round, and update the long-term memory.
2. The method for a psychological companionship dialogue based on multi-agent collaboration according to claim 1, wherein Step 1 includes the following steps: Step 1.1: Perform BERT encoding on the text input by the user to obtain semantic vectors; Step 1.2: Extract features from the voice and images input by the user to obtain emotion feature vectors; Step 1.3: Combine the semantic vectors in Step 1.1 and the emotion feature vectors obtained in Step 2.2, and through the multi-modal emotion assessment model, obtain a description of the user's emotional state.
3. A method for psychological companionship dialogue based on multi-agent collaboration according to claim 1, characterized in that Among them, the implementation of the shared memory module includes the following steps: Step A1: Initialize the shared memory module, including the settings of short-term memory and long-term memory; Step A2: Store the context background and dialogue history records in the short-term memory; Step A3: Store external knowledge documents and dialogue states in the long-term memory; Step A4: When the dialogue is in progress, call the content of the short-term memory and splice it with the input; Step A5: After the reply generation is completed, write the user input and the system output in the current round into the dialogue history record; Step A6: The dialogue state includes the dialogue history record and the processed multi-level dialogue memory; Step A7: The multi-level dialogue memory includes the original dialogue at the lowest level, the paragraph summary at the second level, the event summary at the third level, the emotion summary at the fourth level, and the user portrait at the highest level; Step A8: After each round of dialogue generation, call the relevant agent to update the dialogue state at each level; Step A9: After the update is completed, call the data cache and management module for processing, including the update of the graph database, the relational database, and the vector database.
4. A method for psychological companionship dialogue based on multi-agent collaboration according to claim 1, characterized in that, Step 3 includes the following steps: Step 3.1: The strategy planning module generates multiple possible reply strategies according to the input dialogue state; Step 3.2: Input the result of Step 3.1 into the dialogue content generation module to generate specific response content. The response content of the optimal strategy is directly used as the model response for this round of dialogue. Step 3.3: Transmit the generated specific response content of Step 3.2 back to the strategy planning module. Use the preference model to sort the results of Step 3.1, determine the optimal strategy based on the sorting, form a strategy group, and the strategy planning module performs strategy iteration and parameter update.
5. The method for psychological companionship dialogue based on multi-agent collaboration according to claim 4, wherein Step 3.1 specifically includes the following steps: Step 3.1.1, Policy Planning Module Generate multiple policies based on the input dialogue state, where the policy judged to be the optimal one is denoted as , and the complete policy sequence is denoted as , and the policy planning module is abbreviated as the agent; Step 3.1.2: Invoke the dialogue state from the shared memory, including dialogue history, dialogue event summary, user mental state, user emotion, and user profile. Each output strategy needs to include the name of the agent called and the information passed to it. The complete input prompt template is: { User question in this round: [User question], user's current emotion is [Observed mood]; Current dialogue situation: [memory summary], relevant knowledge: [Useful knowledge]; The agents of this system are: [Agent Name: {Usage, Input Parameters},⋯] You are an intelligent scheduling agent. You need to plan how to use one or more [AgentNames] to respond to the user's question in this round according to the above information, and provide emotional support to relieve their negative emotions. Please output your response plan in the following format, in order, and the plan should include at least one strategy; Used agent: Task description: }。 6. The method for a psychological companionship dialogue based on multi-agent collaboration according to claim 4, wherein Step 3.3 specifically includes the following steps: 3.3.1: Use the preference model to sort the results of the strategy sequence generated in Step 3.1 to determine the optimal strategy. The preference model is based on the Bradley-Terry model, and its preference score calculation formula is: wherein is the included reward function is the Sigmoid function represents the input dialogue state represents the action represents different actions indicates that the action is considered to be better or more preferred than the action Step 3.3.2: Supplement the preference ranking sequence to the policy preference dataset and update the parameters of the policy planning module through direct preference optimization. The optimized loss function is as follows: Among them , , represents an action that leads to a win or victory, represents an action that leads to a loss or defeat; denotes the learned policy parameters, denotes a parameterized policy, where is a set of parameters; represents a reference strategy; is the sigmoid function, which maps real values into the interval (0, 1); is a hyperparameter; is the probability of taking action under the state ; Step 3.3.3: Select the optimal new strategy by maximizing the expected return while considering the diversity of strategies , and the process of selecting the optimal strategy is as follows: , where denotes selecting the strategy that maximizes the expected return among all possible strategies π; Denote over all possible dialogue states Take the expectation; Denote over all possible actions given a policy and a dialogue state to take the expectation; is a hyperparameter used to balance policy selection and policy diversity; Representation strategy Relative to the initial strategy The KL divergence, used to measure the strategy and The difference between them.
7. A method for psychological companionship dialogue based on multi-agent collaboration according to claim 1, characterized in that Step 4 includes the following steps: Step 4.1: The scheduling agent calls the required vertical domain agents according to the response strategy. Step 4.2: The vertical domain agents selected in Step 4.1 generate response content for specific strategies based on the relevant information passed in. Step 4.3: The scheduling agent calls the vertical domain agents to continue generating response content for the next strategy and splices it after the result of Step 4.
2. Step 4.4: Loop through Steps 4.1 - 4.3 until all strategies in the strategy group are used to obtain the complete model response for this round.
8. A method for psychological companionship dialogue based on multi-agent collaboration according to claim 1, characterized in that Step 5 includes the following steps: Step 5.1: Store the user's question and the model response for this round in the dialogue history record of the short-term memory. Step 5.2: Update the dialogue state in the long-term memory according to the update result of Step 5.1, including dialogue history summary, event summary, emotion summary, and user profile. Step 5.3: Store the update result of Step 5.2 in the vector database, relational database, and graph database.
9. A storage medium, characterized in that, When the processor executes the program in the storage medium, it implements a multi-agent collaborative psychological companionship dialogue method as described in any one of claims 1 - 8.
Citation Information
Patent Citations
Psychological counseling methods, terminals, and storage media based on human-computer dialogue
CN111564202B
Emotion support dialogue generation method, system and device based on thinking chain reasoning
CN117932041A
Common-condition reply generation method and device, terminal and storage medium
CN115934909A
Intelligent dialogue system in psychological field
CN117609486A