Psychological counseling dialogue method and system based on psychological large model

By constructing a psychological counseling dialogue method for a large psychological model, using fine-grained psychological state recognition and empathetic text generation modules, combined with the improved LangChain framework and legal chain, the problem of reasonable suggestions and insufficient emotional expression of large language models in the field of mental health is solved, and a more efficient psychological counseling dialogue is achieved.

CN120376059APending Publication Date: 2025-07-25ZHEJIANG UNIV OF FINANCE & ECONOMICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510440901.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

There is still room for improvement in the ability to reasonably suggest and emotional expression in the field of mental health, and there is inaccuracy in medical summary prediction.

Method used

A psychological counseling dialogue method based on psychological large models is adopted, and a fine-grained psychological state recognition model and empathetic text generation module are used, combined with the improved LangChain framework and legal chain, a large language model for psychological counseling is constructed, including prompt word generation, empathetic text generation and large language model modules, and psychological counseling dialogue is conducted.

Benefits of technology

It improves the empathy and professional ability of large language models in psychological counseling, ensures that the generated text is ethical and legal, saves model development costs, and enhances the controllability and personalized answers of text emotions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120376059A_ABST
    Figure CN120376059A_ABST
Patent Text Reader

Abstract

The invention discloses a psychological counseling dialogue method and system based on a psychological large model. The method comprises the steps that psychological counseling description input by a user is acquired; the psychological counseling description is input into a pre-trained psychological counseling large language model, a corresponding estrus-sharing reply text is obtained, the psychological counseling large language model comprises a cue word generation module, an estrus-sharing text generation module and a large language model module, and the cue word generation module is used for generating the estrus-sharing reply text in combination with the historical description of the current dialogue; generating reply prompt words for the psychological consultation description; the large language model module is used for generating a corresponding reply text according to the reply prompt word; and the common-condition text generation module is used for converting the reply text into a common-condition reply text so as to realize psychological counseling dialogue with the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large models, and particularly relates to a psychological counseling dialogue method and system based on a psychological large model. Background Art

[0002] For large language models, they are proposed due to the performance mutation caused by expanding the scale of pre-trained models. In recent years, the field of large models has developed rapidly. OpenAI proposed GPT-1 as early as 2018. From GPT-1 to GPT-3, the expansion of the model scale has brought significant performance improvements. It shows very excellent performance in various NLP tasks and special designed tasks that require reasoning or domain adaptability.

[0003] Regarding the application of large language models and sentiment-based text generation models in the field of mental health. Chatbots are one of the important forms of NLP applications in the field of mental health. Before the widespread use of large models, the original technologies had deficiencies in terms of result diversity, content accuracy, application usability, etc. With the development of large language models and the need for the model-generated text sentiment in the field of mental health, large language models and empathy text generation models are often combined for application. In the academic field, "Hsu SL, Shah RS, Senthil P, Ashktorab Z, Dugan C, Geyer W, Yang D. Helping the Helper: Supporting Peer Counselors via AI-Empowered Practice and Feedback[J]. 2023. arXiv preprint arXiv:2305.08982." uses the BERT model to construct a classifier, determines the most suitable counseling strategy through contextualized language generation technology, and combines it with the large model "Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and William B Dolan. 2020. DIALOGPT: Large-Scale Generative Pre-training for Conversational Response Generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations. 270-278." to provide customized advice for the help seekers. "Chujie Zheng, Sahand Sabour, Jiaxin Wen, Zheng Zhang and Minlie Huang. AugESC: Dialogue Augmentation with Large Language Models for Emotional Support Conversation[J]. 2023. arXiv preprint arXiv:2202.13047." enhances the dialogue by fine-tuning the 6B GPT-J model and the empirical dialogue method, enabling the model to complete multi-topic conversations.In terms of applications, "Yan X, Xue D. MindChat: Psychological Large Language Model [DB / OL]. 2023." constructed MindChat with empathy output and value guidance, while "Chen Y, Xing X, Wang Z, Xu X. SoulChat: A Large Mental Health Model - Enhancing Empathetic Capabilities of Large Models through Mixed Pretraining with Long Text Clinical Prompts and Multi - turn Empathetic Dialogue Data [DB / OL]. 2023." established SoulChat with empathy and listening capabilities through mixed fine - tuning. Although large language models have made progress in the field of mental health, there are still many problems that need to be solved. At the academic level, the transformer in large language models will predict known medical summaries (such as "Javaid M, Haleem A, Singh R P. ChatGPT for healthcare services: An emerging stage for an innovative perspective [J]. BenchCouncil Transactions on Benchmarks, Standards and Evaluations, 2023, 3(1): 100105. ISSN 2772 - 4859."), and the prediction results are inaccurate to a certain extent. At the application level, applications such as MIndChat and SoulChat still have room for improvement in terms of empathy and emotional expression, and the model's suggestion - giving ability also needs to be improved.

[0004] In summary, large language models and emotional text generation models have made some progress in the field of mental health, but the reasonable suggestion - giving ability and emotional expression ability still need to be further improved. Summary of the Invention

[0005] Aiming at the problems existing in the prior art, the purpose of the embodiments of the present application is to provide a psychological counseling dialogue method and system based on a psychological large model.

[0006] According to the first aspect of the embodiments of the present application, a psychological counseling dialogue method based on a psychological large model is provided, including:

[0007] Obtain the psychological counseling description input by the user;

[0008] Input the psychological counseling description into a pre-trained large language model for psychological counseling to obtain a corresponding empathic response text. The large language model for psychological counseling includes a prompt word generation module, an empathic text generation module, and a large language model module. The prompt word generation module is used to generate a response prompt word for the psychological counseling description by combining the historical description of the current conversation. The large language model module is used to generate a corresponding response text based on the response prompt word. The empathic text generation module is used to convert the response text into an empathic response text, thereby realizing a psychological counseling conversation with the user.

[0009] Furthermore, the prompt word generation module includes a fine-grained psychological state recognition model and a prompt word corpus. The fine-grained psychological state recognition model is used to generate multi-dimensional labels for the current description and historical description of the user in the current conversation through a bidirectional long short-term memory network and an attention mechanism. The prompt word corpus is used to generate corresponding prompt words based on the multi-dimensional labels.

[0010] Furthermore, the empathic text generation module includes an encoder and a decoder. The encoder includes several layers of bidirectional long short-term memory network layers, and extracts the feature information of the sentence based on the relationship between each word in the input response text and its context. The decoder part includes several layers of long short-term memory network layers, and predicts the empathic response text based on the encoding result of the encoder and logical coherence.

[0011] Furthermore, in the empathic text generation module, the output layer after the decoder is an OverlappingSoftmax layer. In the OverlappingSoftmax layer, the probability of the output falling into each cluster is calculated through the first softmax layer, the specific position of falling into the cluster is calculated through the second softmax layer, and the third softmax layer provides a complementary item for the number of predicted categories. Among them, the number of prediction types to be constructed in the first softmax layer is The number of prediction types in the second softmax layer is The number of prediction types in the third softmax layer is L is the length of the vocabulary.

[0012] Furthermore, the large language model for psychological counseling also includes a knowledge base constructed based on an improved LangChain framework and multi-agents. The knowledge base is used to generate corresponding text prompt words for the psychological counseling description, and add the text prompt words to the historical description of the current conversation;

[0013] Among them, the construction process of the knowledge base includes:

[0014] Read the psychological counseling knowledge file, convert it into text data and split it into text blocks;

[0015] Extract the keyword and summary information of each text block, and store the keyword, summary information and the corresponding text block in a database, thereby completing the construction of the knowledge base.

[0016] Furthermore, the psychological counseling large language model further includes a legal chain, which is used to judge whether the empathic response text complies with ethics and law. If not, the response needs to be regenerated.

[0017] The process of judging whether the empathic response text complies with ethics and law is as follows:

[0018] Extract the response keyword vector and response sentence vector from the empathic response text.

[0019] Match the response keyword vector and response sentence vector with the law sentence vector and law text summary vector in the legal document tree to obtain several most relevant legal text blocks and convert them into legal prompt words. The legal document tree is formed by splitting and storing laws according to chapters and articles, and each node of the legal document tree is set with the law keyword vector and law text summary vector for matching.

[0020] Based on the legal prompt words, use the large language model module to analyze the empathic response text and judge whether it complies with ethics and law.

[0021] According to the second aspect of the embodiments of the present application, there is provided a psychological counseling dialogue system based on a psychological large model, including:

[0022] A user input acquisition module, which is used to acquire the psychological counseling description input by the user.

[0023] An empathic response generation module, which is used to input the psychological counseling description into a pre-trained psychological counseling large language model to obtain a corresponding empathic response text. The psychological counseling large language model includes a prompt word generation module, an empathic text generation module and a large language model module. The prompt word generation module is used to combine the historical description of the current conversation to generate a response prompt word for the psychological counseling description; the large language model module is used to generate a corresponding response text according to the response prompt word; the empathic text generation module is used to convert the response text into an empathic response text, so as to realize a psychological counseling dialogue with the user.

[0024] Furthermore, the system is provided with three storage modules, namely a dialogue module, a virtual memory module, and a memory module. The dialogue module has the functions of learning, querying and reading, writing, checking and correcting. The virtual memory module has the functions of truncating and compressing, querying and reading, writing, checking and correcting, and summary extraction. The long-term memory mechanism is realized through the interaction between the three storage modules.

[0025] When the user inputs a psychological counseling description to the large model, the dialogue module analyzes and triggers a query and a read request through a parser, searches the content in the virtual memory module, and the searched content is historical information related to the current description. If no relevant content is found, the parser sends a query and a read request from the virtual memory module to the memory module to search for historical content related to the current description in the memory module. After extracting the content into the virtual memory, relevant parts are imported into the dialogue module according to requirements to assist the large model in answering.

[0026] According to a third aspect of the embodiments of the present application, there is provided an electronic device, including:

[0027] One or more processors;

[0028] A memory for storing one or more programs;

[0029] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in the first aspect.

[0030] According to a fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in the first aspect are implemented.

[0031] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:

[0032] The present application uses the prompt word generation module of the CARE+Prompt Lib model to evaluate the current description and historical descriptions of the user in the current dialogue, generate appropriate reply prompt words, apply the large language model in the field of psychology and psychological counseling, generate psychological counseling reply texts, and then introduce a sentiment-based text generation model to make the text sentiment controllable. The present application is applicable to both general psychological large models and general large language models. For general psychological large models, the psychological large model framework can endow general psychological large models with stronger empathy and professional capabilities; for general large language models, the framework can enable them to have the capabilities of psychological large models without fine-tuning or with only a small amount of fine-tuning. While achieving good results in psychological counseling for the model, the model development cost is greatly saved.

[0033] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0035] Figure 1 It is a schematic diagram of a psychological counseling dialogue method based on a large psychological model shown according to an exemplary embodiment.

[0036] Figure 2 It is a schematic diagram of a prompt word generation module shown according to an exemplary embodiment.

[0037] Figure 3 It is a schematic diagram of the CARE model shown according to an exemplary embodiment.

[0038] Figure 4 It is a schematic diagram of the Teacher Forcing Seq2Seq model shown according to an exemplary embodiment.

[0039] Figure 5 It is a flowchart of the operation of the OverlappingSoftmax layer shown according to an exemplary embodiment.

[0040] Figure 6 It is a framework diagram of LangChain+MultiformChat shown according to an exemplary embodiment.

[0041] Figure 7 It is a framework diagram of Long-term Memory shown according to an exemplary embodiment.

[0042] Figure 8 It is a framework diagram of LawChain shown according to an exemplary embodiment.

[0043] Figure 9 It is a schematic diagram of a large language model for psychological counseling shown according to an exemplary embodiment.

[0044] Figure 10 It is a block diagram of a psychological counseling dialogue system based on a large psychological model shown according to an exemplary embodiment.

[0045] Figure 11 It is a schematic diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0046] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with the present application.

[0047] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0048] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".

[0049] A large language model (LLM), also known as a large language model, is an artificial intelligence model designed to understand and generate human language. They are trained on a large amount of text data and can perform a wide range of tasks. In this application, a psychological counseling large language model is constructed by using psychological counseling-related corpus for model fine-tuning on the basis of a general large model and building a local knowledge base.

[0050] Figure 1 is a flowchart of a psychological counseling dialogue method based on a psychological large model shown according to an exemplary embodiment, as Figure 1 shown, the method may include the following steps:

[0051] S1: Obtain the psychological counseling description input by the user;

[0052] In a specific implementation, the user only needs to describe the content to be consulted in natural language in a normal conversation manner.

[0053] For example:

[0054] I've been unhappy recently. I feel like I don't have many friends around me, but I don't want to meet new people either. I don't know how to communicate or what to say so as not to be disliked by others. I also don't know how to maintain relationships or start conversations. I don't make friends much and I feel very lonely. Please give me some advice.

[0055] It should be noted that a user's psychological counseling may result in a conversation containing multiple descriptions and corresponding large model responses. Therefore, the user's descriptions in a conversation are divided into the current description and the historical description.

[0056] S2: Input the psychological counseling description into a pre-trained large language model for psychological counseling to obtain a corresponding empathic response text. The large language model for psychological counseling includes a prompt word generation module, an empathic text generation module, and a large language model module. The prompt word generation module is used to generate corresponding response prompt words for the psychological counseling description in combination with historical data. The large language model module is used to generate a corresponding response text according to the response prompt words. The empathic text generation module is used to convert the response text into an empathic response text.

[0057] In a specific implementation, the prompt word generation module adopts the CARE+Prompt Lib model. The fine-grained psychological state recognition model (hereinafter referred to as the CARE model) constructed by this application is used to evaluate the current description and historical description of the user in the current conversation, generate multi-dimensional labels, and feedback the multi-dimensional labels to the prompt word corpus (PromptLib) to generate corresponding response prompt words. The historical description of the user in the current conversation and the current description combined with the response prompt words are input into the large language model module together, thereby completing the adjustment and optimization of the model input stage.

[0058] For the CARE+Prompt Lib model, its basic structure is as Figure 2 shown. Among them, the CARE model is a fine-grained psychological state recognition model based on the Attention-BiLSTM model. The model structure is as Figure 3 shown, including an input layer, a word encoding (Word2Vec) layer, a bidirectional long short-term memory network (Bidirectional Long Short-Term Memory, BiLSTM) layer, an attention layer, a fully connected layer, and an output layer. The sentence enters the input layer and is tokenized into S1, S2, S3…S t , and each word is encoded into a corresponding vector x1, x2, x3…x i by the word encoding layer. The vector is input into the BiLSTM layer and the attention layer to dynamically extract the sequence context information and capture the key features of the input sequence. h1, h2, h3…h i represents the final calculation result of the BiLSTM. is the parameter of the hidden layer state, which is used to reflect the features calculated by the BiLSTM in two directions of the input sequence. The output layer outputs multi-dimensional labels, and the labels respectively reflect the types of problems encountered, the types of mental illnesses, and the urgency of the situation. The labels specifically correspond to the "id" numbers of each prompt word unit in the Prompt Lib.

[0059] According to experimental findings, the fine-grained mental state recognition model based on the Attention-BiLSTM model has a higher accuracy rate and a faster convergence speed in mental state assessment tasks compared to other models, and its performance is manifested in three mental state assessment tasks from different perspectives (Task 1 to 3 are the recognition tasks of problem types encountered, mental illness recognition tasks, and situation urgency recognition tasks).

[0060] Table 1 Model Convergence Situation Table

[0061]

[0062] For Prompt Lib, it contains many professional and different prompt word units for various mental states. Combining with the guidance of the CARE model, that is, according to multiple classification labels obtained from the CARE model, select the corresponding prompt word units from Prompt Lib, and reconstruct and unify the corpus format within the selected prompt word units and combine the content into a whole. Finally, construct high-quality and targeted reply prompt words to prompt the large model to generate targeted and more user-experienced texts.

[0063] When the reply text generated by the CARE+Prompt Lib model is used in psychological counseling conversations, it sometimes appears rather rigid. Therefore, it is necessary to endow the reply text with emotions to provide users with a better experience. Empathetic text generation is a technology that endows the text generated by machines with emotions. The mainstream processing method in this field is to add some mechanisms, such as emotion vectors, emotion memories, emotion dictionaries, etc. on the basis of the traditional Encoder-Decoder to make the generated reply have emotional factors. Since the seq2seq model plays the role of transferring the meaning expression in one language style to another language style, which is similar to text style transfer, this application introduces the text style transfer task by using the seq2seq framework and improves the seq2seq to construct the Teacher ForcingSeq2Seq model for converting the reply text into an empathetic reply text, thereby realizing the rich emotionality of the output language of the mental large language model. As Figure 4As shown in the figure, the empathy text generation module includes an encoder and a decoder. The encoder part needs to learn text information as much as possible. Therefore, a bidirectional LSTM (BiLSTM) model is used to encode the response text. This encoding method has better effects than traditional models such as LSTM. For the decoder part, since the errors are prone to accumulate when using a sequence-dependent model, a model with stronger fault tolerance is required. Therefore, several groups of LSTM models are used to implement decoding, so as to output the empathy response text. The encoder extracts the feature information of the sentence based on the relationship between each word in the input text and its context. The decoder predicts a new sentence based on the encoding result and the logical coherence from front to back. In addition, when training, the text to be predicted is added to the decoder, which can assist in training the logical coherence of the prediction result. After the encoder obtains the encoding information using this model, training can be carried out with the sequence to be predicted input to the decoder at the same time, which helps to accelerate the convergence of the model. When using this module to process the response text, all or part of the sentences in the response text can be transformed according to the actual situation.

[0064] In order to enable the Teacher Forcing Seq2Seq model to have better empathy effects, in the early stage of this application, a text non-emotional information extraction model is used to collect training data texts: crawl psychological counseling texts and limit the text sources. For example, only select the remarks from professional psychological counselor users (in this embodiment, it is default that the answers of professional psychological counselor users have a certain degree of scientificity and empathy ability). In this way, approximate texts with empathy ability can be obtained. Then, using the abstract extraction technology, the core information of the text is extracted, and the modifying words expressing emotions are removed. Finally, a set of simple non-empathy emotion-rich empathy emotion parallel texts are obtained. Since it is a rough extraction, although a large amount of parallel texts can be obtained in a simple way, there are also certain limitations. The model can only learn about the general empathy emotion transfer method through the simple parallel texts, and there are certain deficiencies in performance. Therefore, high-quality parallel texts are manually screened and used to retrain the empathy text generation module trained with the roughly extracted parallel texts, and the model performance is improved with a small amount of high-quality data.

[0065] In one embodiment, the final training results of the Teacher Forcing Seq2Seq model are shown in Table 2 below.

[0066] Table 2 Comparison table of input and output of the empathy text generation module

[0067]

[0068]

[0069] It is not difficult to find from Table 2 that the Teacher Forcing Seq2Seq model can enhance the empathy ability of the text.

[0070] In natural language generation tasks, the output layer of the network is usually a softmax layer, and its number of parameters is proportional to the size of the vocabulary. The word-based vocabulary model requires a large number of parameters to establish the relationship between words, increasing the modeling difficulty and training complexity. In contrast, the word-based model only needs to consider the relationship between words, with fewer parameters and simplified modeling. However, to improve the effect, a large vocabulary is often required, resulting in a sharp increase in the number of parameters in the softmax layer and significant video memory consumption. In response to this situation, this application constructs a new type of softmax model, OverlappingSoftmax, to replace the output softmax layer after the decoder in the empathy text generation module.

[0071] As Figure 5 shown, the improved softmax needs to construct at most three softmax layers. The first softmax layer calculates the probability that the output falls into each cluster. The role of the second softmax layer is to calculate the specific position within the cluster. The third softmax layer is the complement term for the total number of categories to be predicted and the number of predicted categories provided by the OverlappingSoftmax model.

[0072] The calculation of the number of parameters of these three softmax layers is as follows. Define the vocabulary length of the word-based model using OverlappingSoftmax to reduce parameters as L. Let the size of the first softmax layer to be constructed be x + a (a is a correction parameter), the size of the third softmax layer be y, the number of units in the penultimate layer of the network using OverlappingSoftmax be n, and there are k parameters in each unit of the penultimate layer of the network. If we hope to minimize the number of parameters after construction, we get the following optimization model:

[0073]

[0074] The above optimization model is an approximate calculation method for the number of parameters in the last layer of the network of the model using softmax as the output layer. Solving the above optimization model, we can obtain that the number of predicted categories of the first softmax layer to be constructed is The number of predicted categories of the second softmax layer is The number of predicted categories of the third softmax layer is Adjust the model according to the above parameters, and successfully reduce the parameters to one-thousandth of the original, resulting in a huge improvement in computing efficiency.

[0075] In one embodiment, for the large language model module, a psychological large model can be adopted, or a general large language model can be used. For an ordinary psychological large model, the present application can endow it with stronger empathy ability and professional ability; for a general large language model, it can be made to have the ability of a psychological large model without fine-tuning or with a small amount of fine-tuning. While achieving good results in psychological counseling, the model development cost is greatly saved. Considering that the amount of psychological corpus data may be difficult to support the training of large models, the large model can be trained with a small amount of psychological corpus first, and then the large model can be applied to the conventional domain open corpus to extract the psychological-related corpus for further training. In one embodiment, the ChatGLM series of models can be used for model construction. In addition, the large language model module is also compatible with the Baichuan series, etc., which is convenient to select a more suitable large language model module according to the existing conditions.

[0076] In one embodiment, the answers of the psychological large model module are optimized by introducing LangChain+MultiformChat and the long-term memory mechanism.

[0077] Specifically, LangChain+MultiformChat queries and parses the knowledge base information based on the improved LangChain framework, and combines with the multi-agent with large language model (Muti Agent with LLM) to generate different forms of Prompts and text processing methods to display or implicitly answer the user's psychological counseling description. Among them, the display answer is to directly answer the question after sorting out the extracted knowledge base information, and the implicit answer is to use the extracted knowledge base information as the context to participate in the question answering. The framework structure is as Figure 6 shown.

[0078] As Figure 6 shown, LangChain+MultiformChat, that is, the improved LangChain framework + multi-agent, first needs to read files from local documents, load ocr-type data through paddleocr, read binary-encoded document data through pdfplumber, python-docx, etc., and read plain text data through pandas, etc. After reading, all the data is converted into plain text type data and split into text chunks. The framework extracts the key words (Key Words) and summary information (Summary) of each text chunk by means of word frequency statistics and encodes them into vectors, and stores the vectors corresponding to the text chunks in the database, thus completing the construction of the knowledge base. When in use, various types of books related to psychological counseling can be uploaded according to needs, and this module will convert the book files into a knowledge base in the above manner.

[0079] After the knowledge base is built, when using the LangChain+MultiformChat framework, the description of psychological counseling input by the user is first received. The multi-agent (Muti Agent with LLM) extracts the keywords of the psychological counseling question by filling in relevant prompt words during the input process of the large language model, and at the same time identifies the answering method for the question, and answers different questions in a form suitable for the question through implicit answers (by reconstructing the historical information of the large model, that is, reconstructing the set of the user's historical descriptions in the current conversation, setting the context for it to affect the answering effect. In each round of conversation of the large language model, generally the user's current input and historical information are simultaneously input into the large language model for it to output. The reconstruction in this embodiment is to directly add the extracted knowledge base information to the historical description set) or explicit answers (let the large model directly answer the question by organizing language according to the provided text in the knowledge base). Then, the question and keywords are encoded into vectors. After finding relevant text chunks in the database, knowledge prompt words are generated according to the text chunks, and the knowledge prompt words are added to the historical description of the current conversation, so as to guide the large language model module to make implicit or explicit answers combined with the knowledge prompt words.

[0080] Figure 7 It is the long-term memory mechanism in the model architecture. The long-term memory mechanism Long-term Memory introduces three storage modules, Dialogue (D), Virtual Memory (VM), and Memory (M), and realizes the long-term memory ability of the model through the interaction between these three storage systems.

[0081] The framework constructs four functions, Learn (learning), Query&Read (querying and reading), Write (writing), and Check&Modify (checking and modifying) for the Dialogue (conversation) part, and constructs five functions, Truncate&Compress (truncating and compressing), Query&Read (querying and reading), Write (writing), Check&Modify (checking and modifying), and Summary (abstract extraction) for the Vitrual Memory (virtual memory) part.

[0082] For the D module, when the user inputs a psychological counseling description to the large model, the D module analyzes and triggers a Query&Read request through the parser to search for the content in the VM module. The searched content is the historical information related to the current description. If nothing is found, the VM module sends a Query&Read request to M through the parser to search for the content in M that is related to the current question in history. After extracting the content into the VM, the relevant part is imported into the D module according to the requirements to assist the large model in answering the question.

[0083] If the current conversation reaches a certain number of rounds, or the total amount of conversations stored in the virtual memory VM exceeds a certain length limit, etc., a Write request will be triggered through the parser to write the content as a new conversation into the corresponding next-level module M and mark information such as the timestamp. In addition, when the length of the newly stored original conversations in the virtual memory VM that have not been processed by Tuncate&Compress (truncation and compression) reaches a certain length, the parser will perform a Tuncate&Compress (truncation and compression) operation on the original conversation at the very end.

[0084] In addition, a Check&Modify (check and correction) request is provided to replace the context information that has been negated by the user or does not conform to the actual situation during the work, and to delete incorrect remarks, remarks containing negative guidance, etc. during the work. When the VM initiates a write request to M, a Summary (abstract extraction) request is generally triggered to extract the main information from the new conversation to save storage space.

[0085] When the large language model module answers, according to the type of question and the extraction situation of the historical memory in the VM, a Learn request in the parser is triggered. This request mainly integrates the historical conversations similar to the current conversation retrieved as the context into the D module, guiding the large model to enhance the answer emotionally by referring to the user's language habits in history. After the answer is completed, the context is removed to ensure the rationality of the next round of conversation.

[0086] In one embodiment, the model determines whether the answer conforms to ethics and law through the LawChain. For generative AI, to make it widely used in the industrial field, it is essential to ensure the security of its output content. Especially when generative AI is applied in the field of psychological counseling, it is even more necessary to ensure that the output of the model conforms to ethics and law. This embodiment combines the traditional knowledge base framework to construct a LawChain framework that can be used to check whether the output content conforms to law and ethics. The framework is as Figure 8 shown.

[0087] The LawChain framework mainly analyzes specific laws and regulations and the input content through large models to determine whether the input content violates morality and laws. However, there are many difficulties in the judgment, such as the huge number of laws and regulations, and the limited understanding ability of large models and the length of input texts. Therefore, in response to the above difficulties, a hierarchical matching of laws is proposed. First, the laws are split according to different chapters and different articles, and a tree storage structure is used to reflect the subordinate relationships between its chapters, articles, etc. Thus, it is constructed into a Legal Document Tree. After the construction of the tree storage structure is completed, keyword vectors and text summary vectors of laws are extracted for each node for querying and retrieval.

[0088] When using the LawChain framework of the legal chain, first, the input text (here refers to the empathy response text) is converted into a response keyword vector and a response sentence vector through encoding (Embedding), and the vectors are matched with the corresponding law keyword vectors and law text summary vectors in the legal document tree. After matching the most relevant multiple legal text blocks, they are converted into legal prompt words. Taking the legal prompt words as the task requirements and context, a large language model service is called once and the legal prompt words are passed as parameters into the current description and historical description of the current conversation of the large language model (when used alone, a large language model service needs to be provided separately. In this application, for each module that needs to provide a large language service, its functions can be realized by sequentially accessing the large language model service started in this application). The large language model will analyze and give feedback on the input text according to the imported law information. For the feedback results, keyword matching such as "violation" and "non-compliance" is set, and it can accurately judge and obtain whether the input content violates laws or morality.

[0089] The LawChain framework of the legal chain can accurately detect the legality and morality of texts and give feedback, which greatly ensures the security of the model output.

[0090] In summary, the psychological counseling large language model constructed in this application is mainly composed of modules such as CARE+Prompt Lib, empathy text generation model, LawChain (legal chain), LangChain+MultiformChat, and Long-term Memor combined with the large model module LLM. LLM integrates the interface LLM Core for accessing different large models, and different large models can be accessed through this. In one embodiment, the ChatGLM2-6B model is adopted. For the framework itself, in one embodiment, a management module Manager is also constructed. By controlling and scheduling the computing units occupied when different models are loaded and used, and monitoring the data flow process, the operation efficiency and performance of the psychological large model framework are guaranteed.

[0091] For the data flow process of the framework, first, the large language model, database, and legal statute files are loaded. Under the monitoring of the Manager, the data first flows through CARE + Prompt Lib. CARE is applied to make a preliminary professional judgment on the psychological problems of the inquirer, and then different prompt words are generated according to Prompt Lib to enhance the input of the large model and improve the accuracy and usability of the model's psychological counseling responses. When the input question is too long, the large model will first decompose the question and then solve it step by step. After the large model completes the output, the output text will be passed through the emotional text generation model to enhance the text empathy ability. Thus, the basic psychological counseling ability and emotional expression ability of the ordinary large model are enhanced.

[0092] If it is necessary to further enhance the scientificity and rationality of the model, it can be considered to be achieved by introducing the optimized local knowledge base LangChain + MultiformChat and the long-term memory mechanism Long-term Memory. Before the final output of the large model, LangChain + MultiformChat is selectively introduced to implicitly and explicitly answer the question in combination with the content of the knowledge base. Psychological counseling cases, psychological counseling books, etc. can be stored in the knowledge base. In addition, the long-term memory mechanism Long-term Memory is introduced. The long-term memory mechanism realizes the backtracking and reference of the current model's answer to historical information through the interaction between constructing Dialogue, Virtual History, and Memory. The addition of the long-term memory mechanism helps to better conduct psychological counseling, and the large model can better enhance the emotion of the output text by referring to the user's language preferences in history.

[0093] After all the outputs of the model are completed, the output results are analyzed through the LawChain to determine whether the text complies with laws and ethics. When the security of the text is confirmed, it can be used as the final output. Otherwise, it will return to the large model module and regenerate the data flow.

[0094] For the model deployment control of the framework, the Manager respectively uses each computing unit (CPU, GPU) for the loading and use of different models according to actual requirements, so that the model can be flexibly deployed on the CPU and GPU, greatly reducing the dependence of the model on the GPU configuration.

[0095] The constructed large model is tested and found to have a good answering effect. The model shows good effects in terms of empathy ability and the ability to solve psychological counseling problems. ChatGLM2-6B is used as the large language model module in the psychological counseling large language model of this application, and it is found that its effect in psychological counseling has improved compared with ChatGLM2-6B. The comparison results are shown in Table 3.

[0096] Table 3 Comparison of the answering effects of ChatGLM2-6B before and after using the framework of this application

[0097]

[0098]

[0099] As can be seen from the above embodiments, this application has the following beneficial effects:

[0100] 1. Innovation in the application scenarios of technology

[0101] At present, few people study and develop large language models dedicated to the field of mental health. This application innovatively proposes to apply it in the fields of psychology and psychological counseling. And after considering the problem of the lack of psychological corpus in applying the technology in this field, it is proposed to train the model with a small amount of psychological corpus, and then apply the model to extract the psychological-related corpus from the open corpus in the conventional field.

[0102] 2. Introducing CARE+Prompt Lib to enhance the personalized psychological counseling ability of the large model

[0103] The multidimensional nature of psychological problems increases the complexity of psychological counseling. Psychological problems include a wide range of disorders, and each patient has their own unique individual characteristics and life background, making the problems complex and individualized. General language models have limitations in dealing with such problems and are prone to inaccuracies, which may instead have a wrong impact on the inquirer. Therefore, introducing the CARE+Prompt Lib technology aims to improve the professionalism of large language models in the field of psychological counseling. The model makes a multi-dimensional evaluation of the user's psychological condition, etc., and generates a prompt word framework for the user from multiple perspectives according to the evaluation, so as to enhance its ability at the model input stage. In addition, prompt word enhancement is performed on each question to be answered in each round, and the original question is stored in the model's historical record after the answer is completed, helping the large model to build a context or strengthen the large model's understanding and expression ability of the requirements. Experiments show that in this way, the personalization and professionalism of the model's generated results can be enhanced without affecting the model's memory ability of the chat history.

[0104] 3. Considering the impact of text emotion on users in text generation applications and proposing a text emotion controllable scheme

[0105] For conventional text generation applications, due to scenario issues, developers mostly pursue the accuracy of generated texts and rarely consider the emotions of generated texts and their impact on users. In the field of mental health, due to technical limitations, conventional applications generally provide direct answers through information in the knowledge base, and it is technically impossible to control the emotions of texts. Emotions can greatly affect the psychological state of users and help them in a subtle way. Making AI-generated text controllable can better help users out of difficulties. Therefore, technical innovations are made to introduce emotion-based text generation models to make text emotions controllable.

[0106] 4. Simplify the parameters

[0107] Usually, NLP tasks require the use of a softmax layer to output results. The size of the softmax layer is related to the size of the vocabulary. When the vocabulary is large, the layers before and after the softmax layer will have a huge amount of parameters, increasing the training burden. In conventional projects, the vocabulary length is generally long. For this reason, softmax is optimized and the OverlappingSoftmax layer is proposed. The OverlappingSoftmax layer reduces the parameters of the original softmax position to one thousandth, greatly reducing the training burden.

[0108] 5. Build a LawChain that can judge the legality of specific laws

[0109] By adding LawChain to the big model, LawChain can refer to the imported laws and make judgments on whether the output of the big model violates the law or ethics. LawChain helps to ensure the security of the big model.

[0110] 6. Propose a Long-term Memory framework to increase the model's long-term memory capability

[0111] Psychological consultation is a long-term task, and the information in the historical chat may have an impact on the information in the current chat. Therefore, the long memory mechanism is introduced to selectively connect the information in the historical chat with the current conversation, making the entire psychological consultation process more continuous and reasonable.

[0112] 7. Introducing the Big Psychological Model Framework

[0113] A psychological large model framework is introduced. The framework adjusts the user input and model output to enhance the model's capabilities in the field of psychological counseling. The framework is applicable to both general psychological large models and general large language models. For general psychological large models, the psychological large model framework can endow them with stronger empathy and professional capabilities; for general large language models, the framework can enable them to have the capabilities of psychological large models without fine-tuning or with only a small amount of fine-tuning. While achieving good results in psychological counseling, it greatly saves the model development cost. In addition, the psychological counseling large language model framework proposed in this application can schedule the use of components on different computing units.

[0114] Corresponding to the foregoing embodiments of the psychological counseling dialogue method based on a psychological large model, the present application also provides embodiments of a psychological counseling dialogue system based on a psychological large model.

[0115] Figure 10 It is a block diagram of a psychological counseling dialogue system based on a psychological large model shown according to an exemplary embodiment. Referring to Figure 10 , the system may include:

[0116] A user input acquisition module 21, configured to acquire a psychological counseling description input by a user;

[0117] An empathy reply generation module 22, configured to input the psychological counseling description into a pre-trained psychological counseling large language model to obtain a corresponding empathy reply text, where the psychological counseling large language model includes a prompt word generation module, an empathy text generation module, and a large language model module. The prompt word generation module is configured to generate a reply prompt word for the psychological counseling description by combining the historical description of the current conversation; the large language model module is configured to generate a corresponding reply text according to the reply prompt word; the empathy text generation module is configured to convert the reply text into an empathy reply text, so as to realize a psychological counseling dialogue with the user.

[0118] Regarding the system in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0119] For system embodiments, since they basically correspond to method embodiments, reference may be made to the partial descriptions of the method embodiments for relevant parts. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0120] Correspondingly, this application also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the above-mentioned psychological large model-based psychological counseling dialogue method.

[0121] Correspondingly, this application also provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned psychological large model-based psychological counseling dialogue method. As Figure 11 shown, it is a hardware structure diagram of any device with data processing capabilities where the psychological large model-based psychological counseling dialogue system provided by the embodiment of the present invention is located. In addition to Figure 11 the processors, memory, and network interfaces shown, any device with data processing capabilities where the device in the embodiment is located usually includes other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here.

[0122] Correspondingly, this application also provides a computer-readable storage medium, on which computer instructions are stored, which, when executed by a processor, implement the above-mentioned psychological large model-based psychological counseling dialogue method. The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store the data that has been output or will be output.

[0123] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the content disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present application.

Claims

1. A psychological counseling dialogue method based on a large psychological model, characterized in that, Including: Obtain the psychological counseling description input by the user; Input the psychological counseling description into a pre-trained large language model for psychological counseling to obtain a corresponding empathic response text. The large language model for psychological counseling includes a prompt word generation module, an empathic text generation module, and a large language model module. The prompt word generation module is used to generate a response prompt word for the psychological counseling description by combining the historical description of the current conversation. The large language model module is used to generate a corresponding response text according to the response prompt word. The empathic text generation module is used to convert the response text into an empathic response text, so as to realize a psychological counseling conversation with the user.

2. The method according to claim 1, wherein The prompt word generation module includes a fine-grained mental state recognition model and a prompt word corpus. The fine-grained mental state recognition model is used to generate multi-dimensional tags for the current description and historical description of the user in the current conversation through a bidirectional long short-term memory network and an attention mechanism. The prompt word corpus is used to generate corresponding prompt words based on the multi-dimensional tags.

3. The method according to claim 1, wherein The empathic text generation module includes an encoder and a decoder. The encoder includes several layers of bidirectional long short-term memory network layers, and extracts the feature information of the sentence based on the relationship between each word in the input response text and its context. The decoder part includes several layers of long short-term memory network layers, and predicts the empathic response text based on the encoding result of the encoder and logical coherence.

4. The method according to claim 3, wherein In the empathy text generation module, the output layer after the decoder is an OverlappingSoftmax layer. In the OverlappingSoftmax layer, the probability of the output falling into each cluster is calculated through the first softmax layer, the specific position of falling into the cluster is calculated through the second softmax layer, and the complement item of the predicted class number is provided through the third softmax layer, where the number of prediction types of the first softmax layer to be constructed is The number of prediction types of the second softmax layer is The number of prediction types of the third softmax layer is L is the length of the vocabulary.

5. The method according to claim 1, characterized in that The large language model for psychological counseling also includes a knowledge base constructed based on an improved LangChain framework and multi-agents. The knowledge base is used to generate corresponding text prompt words for the psychological counseling description, and add the text prompt words to the historical description of the current conversation; Among them, the construction process of the knowledge base includes: Read the psychological counseling knowledge file, convert it into text data and split it into text blocks; Extract the keywords and summary information of each text block, and store the keywords, summary information and corresponding text blocks in the database, thus completing the construction of the knowledge base.

6. The method according to claim 1, characterized in that The large language model for psychological counseling also includes a legal chain, which is used to judge whether the empathic response text complies with ethics and laws. If not, a new response needs to be generated; The process of judging whether the empathic response text complies with ethics and laws is: Extract the response keyword vector and response sentence vector from the empathic response text; Match the response keyword vector and response sentence vector with the legal article sentence vector and legal article text summary vector in the legal document tree to obtain the most relevant several legal text blocks and convert them into legal prompt words. The legal document tree is formed by splitting and storing the legal articles according to chapters and articles. Each node of the legal document tree is set with the legal article keyword vector and legal article text summary vector for matching; Based on the legal prompt words, use the large language model module to analyze the empathic response text and judge whether it complies with ethics and laws.

7. A psychological counseling dialogue system based on a large psychological model, characterized in that, Including: A user input acquisition module, which is used to obtain the psychological counseling description input by the user; An empathy response generation module, which is used to input the psychological counseling description into a pre-trained large language model for psychological counseling to obtain a corresponding empathy response text. The large language model for psychological counseling includes a prompt word generation module, an empathy text generation module, and a large language model module. The prompt word generation module is used to generate a response prompt word for the psychological counseling description by combining the historical description of the current conversation. The large language model module is used to generate a corresponding response text according to the response prompt word. The empathy text generation module is used to convert the response text into an empathy response text, so as to realize a psychological counseling conversation with the user.

8. The system according to claim 7, characterized in that, The system is provided with three storage modules, namely a conversation module, a virtual memory module, and a memory module. The conversation module has the functions of learning, querying and reading, writing, checking and correcting. The virtual memory module has the functions of truncating and compressing, querying and reading, writing, checking and correcting, and abstract extraction. The long-term memory mechanism is realized through the interaction between the three storage modules. When the user inputs a psychological counseling description into the large model, the conversation module analyzes and triggers a query and read request through a parser, searches the content in the virtual memory module, and the searched content is historical information related to the current description. If no relevant content is found, the parser is used to send a query and read request from the virtual memory module to the memory module, search for historical content related to the current description in the memory module, extract the content into the virtual memory, and then import the relevant part into the conversation module according to the requirements to assist the large model in answering.

9. An electronic device, characterized in that, Comprising: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.

10. A computer-readable storage medium having computer instructions stored thereon, characterized in that, When the instruction is executed by the processor, it implements the steps of the method according to any one of claims 1-6.