Methods, devices, equipment, storage media, and products for multi-turn dialogue in psychological counseling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]然而,当前模型在心理咨询过程中输出的内容常聚焦于单一维度,难以覆盖用户潜在的多元心理需求
本申请提供的心理咨询多轮对话方案,通过响应于多轮对话指令,从多轮对话指令中提取本轮问题,基于至少一轮历史问题对应的历史意图,对本轮问题进行意图识别,得到本轮问题对应的本轮意图。这种时序关联的意图识别机制不仅能捕捉用户当前表达的浅层诉求,更能通过历史对话的上下文推理挖掘用户潜在的深层心理需求,从而显著提升意图理解的准确性,使得本轮意图能够更精准地反映用户实际需求。因此,基于本轮意图,生成的多个心理学话题不仅能够紧密围绕用户核心诉求,还能覆盖相关联的多元心理维度。在此基础上,通过大语言模型,基于本轮意图和多个心理学话题,生成本轮问题对应的回答,既能确保模型回答始终贴合用户真实意图,又能避免回答陷入单一维度的局限,使得回答内容在不偏离用户意图的前提下,增加话题丰富度,从而有效改善了咨询效果。
Smart Images

Figure CN121543712B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing technology, and in particular to a method, apparatus, device, storage medium and product for multi-turn dialogue in psychological counseling. Background Technology
[0002] Psychological counseling is a service that uses psychological theories and methods to help individuals resolve psychological distress and improve their mental state through professional communication. It plays an important role in areas such as emotion regulation and interpersonal relationship management. With the development of natural language processing technology, large language models are increasingly being applied to psychological counseling scenarios due to their efficiency and convenience.
[0003] However, current models often focus on a single dimension in their output during psychological counseling, failing to cover the diverse potential psychological needs of users. For example, when addressing parent-child conflicts, the model might only focus on "communication skills," neglecting related topics such as "emotional guidance" and "the impact of the family ecosystem," resulting in insufficient topic richness and affecting counseling effectiveness. Therefore, there is an urgent need to provide a method to solve this problem.
[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of this application is to provide a method, device, equipment, storage medium and product for multi-turn dialogue in psychological counseling, which can effectively improve the topic richness of the model in multi-turn dialogue in psychological counseling.
[0006] To achieve the above objectives, this application proposes a multi-round dialogue method for psychological counseling, the method comprising: In response to a multi-turn dialogue instruction, extract the current round question from the multi-turn dialogue instruction; Based on the historical intent corresponding to at least one round of historical questions, the intent of the current round of questions is identified to obtain the current round intent corresponding to the current round of questions. Based on the stated intent of this round, several psychological topics were generated; Using a large language model, answers to the questions in this round are generated based on the current intent and the multiple psychological topics.
[0007] Optionally, the step of generating the answer to the question in the current round based on the current round's intent and the multiple psychological topics using a large language model includes: Using a large language model, multiple alternative answers are generated based on the stated intent and the multiple psychological topics. Determine the reward points for each alternative answer; The answer with the highest reward score among the multiple alternative answers will be selected as the answer to the question in this round.
[0008] Optionally, determining the reward score for each alternative answer includes: For any given answer, obtain at least one of the following scores: word count in a single round of dialogue, number of rounds in a multi-round dialogue, degree of deviation from intent, and richness of topic; The score is obtained by weighted summation of the at least one score, and the reward score is obtained for any candidate answer.
[0009] Optionally, for any candidate answer, a score is obtained to indicate the degree of deviation from intent, including: Based on the intent of the question corresponding to any of the alternative answers, search for target psychological knowledge from the psychological knowledge base; Based on the semantic similarity between any of the alternative answers and the target psychological knowledge, the intention deviation score is determined, and the semantic similarity is positively correlated with the intention deviation score.
[0010] Optionally, for any of the alternative answers, a topic richness score is obtained, including: A codebook is constructed based on psychological knowledge. The codebook includes multiple semantic vectors, and each semantic vector has a unique semantic identifier. Convert the intent of the question corresponding to any of the alternative answers into an intent vector; Based on the vector similarity between the intent vector and each semantic vector in the codebook, a set of semantic identifiers is searched. Based on the domain topics corresponding to each semantic identifier in the semantic identifier set, the topic richness score corresponding to any candidate answer is determined.
[0011] Optionally, the step of identifying the intent of the current round's question based on the historical intent corresponding to at least one round of historical questions, to obtain the current round's intent corresponding to the current round's question, includes: Obtain intent recognition prompts, wherein the intent recognition prompts include historical intents corresponding to at least one round of historical questions; Based on the intent recognition prompts, the intent of the current round question is obtained by using the large language model to identify the intent of the current round question.
[0012] Optionally, the step of identifying the intent of the current round's question based on the historical intent corresponding to at least one round of historical questions, to obtain the current round's intent corresponding to the current round's question, includes: From the at least one round of historical questions, identify any omitted questions that were not covered by the historical answers; Based on the historical intent corresponding to at least one round of historical questions, the intent of the omitted questions and the current round of questions is identified to obtain the intent of the current round.
[0013] Optionally, based on the intent of this round, multiple psychological topics are generated, including: Obtain rich topic prompts, and add the intent of this round to the rich topic prompts; Based on the rich prompts for the topic, multiple psychological topics related to the current intention are generated through the large language model.
[0014] Optionally, based on the intent of this round, multiple psychological topics are generated, including: Determine the content coverage of at least one round of historical answers to the at least one round of historical questions; Based on the content coverage rate, a multi-turn dialogue round score is generated, and the content coverage rate is positively correlated with the multi-turn dialogue round score. If the score of the number of rounds in the multi-round dialogue is less than the score threshold, multiple psychological topics are generated based on the intent of the current round.
[0015] Optionally, before generating the answer to the question in the current round based on the current round's intent and the multiple psychological topics using a large language model, the method further includes: Obtain multi-turn dialogue training data, wherein the multi-turn dialogue training data includes at least one round of questions; The large language model is used to answer the questions in at least one round, and the reward score corresponding to the generated answers in at least one round is determined. The large language model is trained based on the reward scores corresponding to the at least one round of responses, so that the large language model maximizes the reward scores.
[0016] Optionally, after determining the reward score corresponding to at least one round of responses, the method further includes: Multiple rounds of sampling are performed on the at least one round of responses, and each round of sampling is used to extract at least one round of responses from the at least one round of responses; Calculate the information entropy based on the reward score corresponding to at least one round of answers in each round of sampling; Determine the target round of sampling in which the information entropy is greater than the information entropy threshold among the multiple rounds of sampling; Based on the at least one round of responses corresponding to the target round number and the question corresponding to the at least one round of responses, a new multi-round dialogue training corpus is constructed.
[0017] Optionally, before generating the answer to the question in the current round based on the current round's intent and the multiple psychological topics using a large language model, the method further includes: Obtain seed question-and-answer pairs; Based on the seed question-answer pairs, expanded question-answer pairs are generated using the teacher big model; Select target question-answer pairs that meet the quality criteria from the expanded question-answer pairs; The seed question-answer pairs and the target question-answer pairs are used as the training corpus for the large language model.
[0018] Optionally, the step of generating expanded question-answer pairs based on the seed question-answer pairs using the teacher's large model includes: Clustering is performed on the seed question-answer pairs to obtain clustering results, which include multiple category semantic centers; Using a large teacher model, psychological counseling questions related to the semantic center of the category are generated, and corresponding answers to the psychological counseling questions are generated. The psychological counseling questions and answers generated by the aforementioned teacher big data model constitute the expanded question-and-answer pairs.
[0019] Furthermore, to achieve the above objectives, this application also proposes a multi-turn dialogue device for psychological counseling, the device comprising: The question extraction module is used to extract the current question from the multi-turn dialogue instructions in response to the multi-turn dialogue instructions; The intent recognition module is used to perform intent recognition on the current round question based on the historical intent corresponding to at least one round of historical questions, so as to obtain the current round intent corresponding to the current round question. The topic generation module is used to generate multiple psychology topics based on the stated intent of this round. The answer generation module is used to generate answers to the questions in the current round based on the current round's intent and the multiple psychological topics, using a large language model.
[0020] Optionally, the answer generation module includes: The answer generation unit is used to generate multiple alternative answers based on the current intent and the multiple psychological topics using a large language model; The score determination unit is used to determine the reward score for each candidate answer; The answer selection unit is used to select the answer with the highest reward score from the multiple alternative answers as the answer to the question in this round.
[0021] Optionally, the score determination unit is used to obtain at least one of the following scores for any candidate answer: single-round dialogue word count score, multi-round dialogue round count score, intention deviation score, and topic richness score; and to perform a weighted summation of the at least one score to obtain the reward score corresponding to any candidate answer.
[0022] Optionally, the score determination unit is used to search for target psychological knowledge from a psychological knowledge base based on the intent of the question corresponding to any of the candidate answers; and to determine the intent deviation score based on the semantic similarity between any of the candidate answers and the target psychological knowledge, wherein the semantic similarity is positively correlated with the intent deviation score.
[0023] Optionally, the score determination unit is used to construct a codebook based on psychological knowledge, the codebook including multiple semantic vectors, and each semantic vector having a unique semantic identifier; convert the intent of the question corresponding to any candidate answer into an intent vector; search a set of semantic identifiers based on the vector similarity between the intent vector and each semantic vector in the codebook; and determine the topic richness score corresponding to any candidate answer based on the domain topic corresponding to each semantic identifier in the set of semantic identifiers.
[0024] Optionally, the intent recognition module is used to obtain intent recognition prompts, the intent recognition prompts including historical intents corresponding to at least one round of historical questions; based on the intent recognition prompts, the intent of the current round of questions is recognized by the large language model to obtain the intent of the current round of questions.
[0025] Optionally, the intent recognition module is used to determine the missing questions not covered by historical answers from the at least one round of historical questions; and to perform intent recognition on the missing questions and the current round of questions based on the historical intent corresponding to the at least one round of historical questions, so as to obtain the intent of the current round.
[0026] Optionally, the topic generation module is used to obtain rich topic prompts, in which the current intent is added; based on the rich topic prompts, multiple psychological topics related to the current intent are generated through the large language model.
[0027] Optionally, the topic generation module is used to determine the content coverage of at least one round of historical answers to the at least one round of historical questions; based on the content coverage, generate a multi-round dialogue round score, wherein the content coverage is positively correlated with the multi-round dialogue round score; and if the multi-round dialogue round score is less than a round score threshold, generate multiple psychology topics based on the intent of the current round.
[0028] Optionally, the device further includes: The first training module is used to acquire multi-turn dialogue training corpus, which includes at least one round of questions; to answer the at least one round of questions sequentially using the large language model, and to determine the reward score corresponding to the generated at least one round of answers; and to train the large language model based on the reward score corresponding to the at least one round of answers, so that the large language model maximizes the reward score.
[0029] Optionally, the device further includes: The first corpus generation module is used to perform multiple rounds of sampling on the at least one round of responses, with each round of sampling used to extract at least one round of responses from the at least one round of responses; calculate information entropy based on the reward score corresponding to the at least one round of responses corresponding to each round of sampling; determine the target round of sampling in the multiple rounds of sampling whose corresponding information entropy is greater than the information entropy threshold; and construct a new multi-round dialogue training corpus based on the at least one round of responses corresponding to the target round of sampling and the questions corresponding to the at least one round of responses.
[0030] Optionally, the device further includes: The second corpus generation module is used to obtain seed question-answer pairs; generate expanded question-answer pairs based on the seed question-answer pairs using the teacher large model; select target question-answer pairs that meet the quality conditions from the expanded question-answer pairs; and use the seed question-answer pairs and the target question-answer pairs as the training corpus of the large language model.
[0031] Optionally, the second corpus generation module is used to cluster the seed question-answer pairs to obtain clustering results, the clustering results including multiple category semantic centers; generate psychological counseling questions related to the category semantic centers through a teacher big model, and generate answers corresponding to the psychological counseling questions; and construct the expanded question-answer pairs by combining the psychological counseling questions and answers generated by the teacher big model.
[0032] Optionally, the device further includes: The second training module is used to acquire a secure training corpus, in which the questions and answers do not contain sensitive content; and to fine-tune the large language model based on the secure training corpus.
[0033] In addition, to achieve the above objectives, this application also proposes a multi-turn dialogue device for psychological counseling, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the multi-turn dialogue method for psychological counseling as described above.
[0034] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the multi-turn dialogue method for psychological counseling as described above.
[0035] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the multi-turn dialogue method for psychological counseling as described above.
[0036] One or more technical solutions proposed in this application have at least the following technical effects: The multi-turn dialogue solution for psychological counseling provided in this application extracts the current round's question from the multi-turn dialogue instructions in response to those instructions. Based on the historical intent corresponding to at least one previous round's question, it identifies the intent of the current round's question, thus obtaining the intent corresponding to the current round's question. This temporally related intent identification mechanism not only captures the user's current superficial needs but also mines the user's potential deep psychological needs through contextual reasoning from historical dialogues, significantly improving the accuracy of intent understanding and enabling the current round's intent to more accurately reflect the user's actual needs. Therefore, the multiple psychological topics generated based on the current round's intent not only closely revolve around the user's core needs but also cover related multi-dimensional psychological aspects. Furthermore, through a large language model, based on the current round's intent and multiple psychological topics, it generates answers corresponding to the current round's question. This ensures that the model's answers always align with the user's true intent while avoiding the limitations of a single dimension. This allows the answer content to increase topic richness without deviating from the user's intent, thereby effectively improving the counseling outcome. Attached Figure Description
[0037] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0038] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a schematic diagram of an implementation environment for the multi-round dialogue method in psychological counseling proposed in this application; Figure 2 A flowchart illustrating the first embodiment of the multi-round dialogue method for psychological counseling in this application; Figure 3A flowchart illustrating the second embodiment of the multi-round dialogue method for psychological counseling in this application; Figure 4 A schematic diagram of a codebook construction process provided in an embodiment of this application; Figure 5 This is a flowchart illustrating the third embodiment of the multi-round dialogue method for psychological counseling in this application. Figure 6 This is a schematic diagram illustrating fine-tuning training of a large language model, as provided in an embodiment of this application. Figure 7 This is a schematic diagram of the neural network structure of a security adapter provided in an embodiment of this application; Figure 8 This is a schematic diagram illustrating reinforcement learning training of a large language model, provided as an embodiment of this application. Figure 9 This is a schematic diagram of the module structure of the multi-turn dialogue device for psychological counseling, as described in an embodiment of this application. Figure 10 This is a schematic diagram of the hardware operating environment involved in the multi-turn dialogue method for psychological counseling in the embodiments of this application.
[0040] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0041] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0042] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0043] Figure 1 This is a schematic diagram of an implementation environment provided by an embodiment of this disclosure. See also... Figure 1 The implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 are connected via a wireless or wired network. For example, the terminal 101 is installed with a target application provided by the server 102, and the terminal 101 can perform functions such as data transmission and message interaction through the target application.
[0044] For example, terminal 101 can be a computer, mobile phone, tablet computer, or other terminal. For example, the target application can be a target application within the operating system of terminal 101, or a target application provided by a third party. For instance, the target application could be a chat application, a search application, etc., which has multi-turn dialogue functionality and can provide psychological counseling services. For example, server 102 can be the backend server corresponding to the target application. Accordingly, server 102 could be a chat application server, a search application server, etc.
[0045] In this application, terminal 101 is used to respond to a multi-turn dialogue instruction, extract the current round question from the multi-turn dialogue instruction, and send the current round question to server 102. Server 102 is used to perform intent recognition on the current round question based on the historical intent corresponding to at least one previous round of historical questions, to obtain the current round intent corresponding to the current round question; based on the current round intent, generate multiple psychological topics; and through a large language model, generate the answer corresponding to the current round question based on the current round intent and the multiple psychological topics. Then, the answer corresponding to the current round question is sent to terminal 101. Terminal 101 is used to display the answer corresponding to the current round question in the multi-turn dialogue interface.
[0046] Alternatively, the aforementioned multi-round psychological counseling process can also be completed by terminal 101 alone. Alternatively, terminal 101 can complete it through the installed target application. This application embodiment does not impose any limitations on this.
[0047] The multi-round dialogue method for psychological counseling provided in this application is applicable to various scenarios. For example, in the scenario of child psychological counseling: children often face problems such as academic pressure and emotional fluctuations, and their deep-seated troubles may gradually be revealed in multiple rounds of dialogue. This application identifies the intent of the current round by combining historical intent and generates rich psychological topics based on the intent of the current round, so that the answers can revolve around the core needs, thereby assisting in adjusting the child's mindset from multiple levels.
[0048] Workplace psychological support scenario: Professionals may experience stress and interpersonal difficulties. This application identifies current intent by combining historical intent with current intent, and generates various related topics such as stress management, career planning, and communication advice based on current intent. By using current intent and related topics, the application helps a large language model generate responses, enabling the model to provide more comprehensive workplace psychological advice.
[0049] Figure 2 This is a flowchart illustrating the first embodiment of the multi-round dialogue method for psychological counseling in this application. (Refer to...) Figure 2 Taking the implementing entity as the terminal as an example, this multi-round dialogue method for psychological counseling includes the following steps S10~S40: Step S10: In response to a multi-turn dialogue instruction, extract the current question from the multi-turn dialogue instruction.
[0050] The multi-turn dialogue method in psychological counseling is a dialogue interaction method applied in psychological counseling scenarios. Its characteristic is that through multiple rounds of communication between the user and the terminal, combined with contextual information and user intent, it generates coherent and personalized dialogue responses with professional psychological support.
[0051] Multi-turn dialogue commands are instructions triggered by user input during continuous interaction with a terminal. For example, when a user enters dialogue content in a multi-turn dialogue interface using voice input or keyboard input, the terminal generates a multi-turn dialogue command. The dialogue content entered by the user constitutes the question carried in the multi-turn dialogue command.
[0052] The question in this round is the content of the latest statement entered by the user in the current round of the dialogue, which serves as the main basis for the terminal to generate a response.
[0053] Step S20: Based on the historical intent corresponding to at least one round of historical questions, perform intent recognition on the current round of questions to obtain the current round intent corresponding to the current round of questions.
[0054] The current intent refers to the core psychological need or consultation goal expressed by the user through the questions asked in the current dialogue round. The historical intent corresponding to historical questions refers to the intent reflected in the questions asked by the user in previous dialogue rounds with the terminal. Intent recognition refers to the process of identifying the core goals or psychological needs expressed by users from their input statements using natural language processing technology.
[0055] Optionally, based on the historical intent corresponding to at least one round of historical questions, intent recognition is performed on the current round of questions to obtain the intent corresponding to the current round of questions, including: obtaining intent recognition prompts, which include the historical intent corresponding to at least one round of historical questions; and based on the intent recognition prompts, intent recognition is performed on the current round of questions through a large language model to obtain the intent corresponding to the current round of questions.
[0056] Intent recognition prompts are a set of structured text information used to guide large language models in intent recognition. They contain key intent information from historical dialogues, helping the model to more accurately understand the background and semantics of the current question.
[0057] This application's embodiments introduce intent recognition prompts, integrating intent information from the user's historical dialogues into the current intent recognition process. Leveraging the semantic understanding capabilities of a large language model, more accurate intent recognition can be achieved. Compared to traditional methods that rely solely on the current question for isolated judgment, this approach effectively utilizes contextual information, improving the accuracy and robustness of intent recognition. Especially in scenarios like psychological counseling, which heavily rely on contextual understanding, this improvement helps the terminal better grasp the user's potential psychological state and long-term counseling focus, thus providing a solid foundation for generating multi-dimensional, personalized psychological topics and professional responses, effectively enhancing the effectiveness of psychological counseling.
[0058] Optionally, based on the historical intent corresponding to at least one round of historical questions, the intent of the current round of questions is identified to obtain the intent of the current round of questions, including: identifying the omitted questions not covered by historical answers from at least one round of historical questions; and identifying the intent of the omitted questions and the current round of questions based on the historical intent corresponding to at least one round of historical questions to obtain the intent of the current round.
[0059] Historical questions refer to at least one question asked by the user before the current round during a multi-round dialogue with the terminal. Historical answers are the responses generated by the terminal in response to the user's historical questions.
[0060] Missed issues refer to psychological topics or counseling requests that users raised in previous conversations, but which the terminal's historical responses failed to fully cover or address. These issues may be problems that users are still concerned about and have not yet resolved.
[0061] In this embodiment, the solution addresses the shortcomings of traditional intent recognition methods that rely solely on the current question while ignoring unanswered historical questions by introducing a mechanism for identifying missed questions. By re-examining user questions not covered by historical responses and combining them with the current question for intent recognition, a more comprehensive understanding of the user's persistent psychological distress and deep-seated needs can be achieved. This enhances the depth of understanding of the user's true intent, thereby strengthening the coherence and personalization of the dialogue and significantly improving the consultation effect.
[0062] Step S30: Based on the intent of this round, generate multiple psychological topics.
[0063] Psychology topics are discussion topics related to psychological theories, practices, or intervention methods, covering multiple dimensions such as emotion regulation, cognitive restructuring, interpersonal relationship management, and family systems, and are used to guide the terminal to generate response content with professional psychological support.
[0064] Optionally, based on the intent of this round, multiple psychological topics are generated, including: obtaining rich topic prompts, adding the intent of this round to the rich topic prompts; and generating multiple psychological topics related to the intent of this round through a large language model based on the rich topic prompts.
[0065] Topic-rich prompts are structured textual information containing the current intent and related guidance information. They are used to prompt the large language model to generate multiple topics related to the current intent and covering different psychological dimensions. For example, topic-rich prompts can also include historical intents corresponding to historical questions.
[0066] This application's embodiments introduce topic-enriching prompts, explicitly embedding the user's current intent into the topic generation process. This leverages the language understanding and creative capabilities of a large language model to generate multiple psychological topics highly relevant to the user's psychological needs but from different perspectives. Compared to traditional single-topic responses, this mechanism effectively enhances the diversity and professional depth of responses, avoiding overly limited or repetitive answers. Simultaneously, multi-faceted psychological topics help stimulate users' self-exploration and reflection, enhancing the inspiration and interactivity of the dialogue, thereby improving the overall comprehensiveness and personalization of the consultation service.
[0067] Optionally, based on the intent of the current round, multiple psychological topics are generated, including: determining the content coverage of at least one round of historical answers to at least one round of historical questions; generating a multi-round dialogue round score based on the content coverage, wherein the content coverage is positively correlated with the multi-round dialogue round score; and generating multiple psychological topics based on the intent of the current round if the multi-round dialogue round score is less than a round score threshold.
[0068] Content coverage refers to the degree to which historical responses respond to the core psychological content or potential needs of users contained in historical questions. It can be evaluated through semantic analysis or keyword matching, and reflects whether the terminal's response content fully covers the key points of the user's questions.
[0069] The multi-turn dialogue score is an indicator calculated based on content coverage, used to measure the depth and breadth of a terminal's responses to user questions during multi-turn dialogues. The higher the content coverage, the higher the score, indicating that the terminal has comprehensively addressed the user's needs.
[0070] The round-count score threshold is a preset scoring standard value used to determine whether the terminal has adequately responded to the user's previous questions. When the multi-round conversation score is lower than this threshold, it indicates that there are still psychological issues that have not been explored in depth, and further topic expansion is needed.
[0071] This application's embodiments achieve dynamic monitoring of the completeness and depth of the psychological counseling process by quantitatively evaluating the terminal's content coverage of user questions in historical dialogues and generating a multi-turn dialogue score accordingly. When the multi-turn dialogue score falls below a set threshold, the terminal proactively generates multiple psychological topics based on the intent of the current turn, helping to compensate for potential coverage deficiencies in previous dialogues and preventing the dialogue from becoming one-dimensional or repetitive. This mechanism enhances the intelligence and adaptability of the psychological counseling system, meeting users' diverse psychological needs and improving the overall counseling experience and service quality.
[0072] For example, when the round score in a multi-turn dialogue is greater than or equal to a round score threshold, the terminal directly generates an answer to the question in the current round based on the intent of that round using a large language model. When the round score is greater than or equal to the threshold, it indicates that the terminal has already comprehensively and deeply covered the user's psychological issues in previous conversations, possessing sufficient contextual understanding. In this case, the terminal directly generates an answer based on the intent of the current round using a large language model, which can improve dialogue efficiency and response speed while ensuring the professionalism of the response and avoiding redundant information from interfering with the user. Furthermore, this approach also avoids excessive elaboration on psychological topics, helping to end the dialogue earlier, reducing the number of rounds, and ensuring efficient resource utilization.
[0073] Step S40: Using a large language model, based on the intent of this round and multiple psychological topics, generate the answer to the question in this round.
[0074] Large language models are artificial intelligence models trained on large-scale corpora that possess language understanding and generation capabilities. They can generate structured, semantically coherent, and content-rich natural language outputs based on input text.
[0075] Optionally, using a large language model, based on the current intent and multiple psychological topics, the answer to the current question is generated. This includes: generating contextual information based on the current intent and multiple psychological topics; and generating the answer to the current question based on the contextual information using the large language model. Contextual information refers to the complete background information relied upon when generating the answer, including the current intent, multiple psychological topics, historical intent, and historical dialogue content, etc., to provide sufficient semantic support for the large language model to generate more accurate, coherent, and personalized responses.
[0076] The multi-turn dialogue solution for psychological counseling provided in this application extracts the current round's question from the multi-turn dialogue instructions in response to those instructions. Based on the historical intent corresponding to at least one previous round's question, it identifies the intent of the current round's question, thus obtaining the intent corresponding to the current round's question. This temporally related intent identification mechanism not only captures the user's current superficial needs but also mines the user's potential deep psychological needs through contextual reasoning from historical dialogues, significantly improving the accuracy of intent understanding and enabling the current round's intent to more accurately reflect the user's actual needs. Therefore, the multiple psychological topics generated based on the current round's intent not only closely revolve around the user's core needs but also cover related multi-dimensional psychological aspects. Furthermore, through a large language model, based on the current round's intent and multiple psychological topics, it generates answers corresponding to the current round's question. This ensures that the model's answers always align with the user's true intent while avoiding the limitations of a single dimension. This allows the answer content to increase topic richness without deviating from the user's intent, thereby effectively improving the counseling outcome.
[0077] Based on the first embodiment described above, a second embodiment of this application is proposed. Contents that are the same as or similar to the first embodiment can be referred to the above description and will not be repeated hereafter. (Refer to...) Figure 3 In the second embodiment, step S40 includes steps S401 to S403: Step S401: Using a large language model, generate multiple alternative answers based on the current round's intent and multiple psychological topics.
[0078] Alternative answers refer to multiple candidate answers generated by the large language model during the answer generation process based on the current round's intent and multiple psychological topics, for use in subsequent screening.
[0079] Step S402: Determine the reward score for each alternative answer.
[0080] The reward score is a quantitative indicator used to measure the quality of alternative responses. It can cover multiple dimensions of evaluation results. The higher the reward score, the more the response meets the standards of high-quality psychological counseling.
[0081] Optionally, the reward score for each candidate answer is determined, including: for any candidate answer, obtaining at least one of the following scores: single-round dialogue word count score, multi-round dialogue round count score, intention deviation score, and topic richness score; and weighted summing of the at least one score to obtain the reward score corresponding to any candidate answer.
[0082] The single-turn dialogue word count score is assigned based on the number of words in the current round's response, measuring the information content and completeness of the expression. A reasonable range can be set, and the single-turn dialogue word count score is given based on whether the current round's response falls within this range, avoiding insufficient information due to too little content or difficulty in understanding due to too much content. The multi-turn dialogue round score is an indicator calculated based on the content coverage of historical answers to historical questions, reflecting the comprehensiveness of the system's response to the user's psychological questions. The greater the content coverage, the higher the multi-turn dialogue round score. The intent deviation score measures the semantic matching degree between the current candidate answer and the user's intent in this round. A higher score indicates that the answer is more in line with the user's intent and the smaller the deviation; a lower score indicates that the answer deviates more from the user's intent. The topic richness score assesses the diversity and relevance of the psychological topics covered in the answer, determining whether the answer provides valuable content from multiple psychological perspectives.
[0083] Of course, in addition to the dimensions mentioned above, other dimensions can be combined to determine the reward score for the answer, such as a correctness score, which measures whether the answer is correct. This application does not limit this approach.
[0084] For example, the aforementioned reward score can be a PRM (Process Reward Model) score, which is a score given by the process reward model to evaluate the quality of the answer. That is, this embodiment of the application scores the candidate answers using a PRM model. This PRM model can be trained in advance using scoring training samples. The scoring training samples include questions, answers, and the corresponding scores for the answers. The scores for the answers in the scoring training samples can be manually labeled or calculated by combining scores given by other scoring models.
[0085] This application's embodiments introduce a multi-dimensional scoring mechanism, including scores for single-turn dialogue word count, multi-turn dialogue number of turns, degree of intent deviation, and topic richness, to construct a more scientific and objective evaluation system for alternative responses. By weighted summing of these scores, the terminal can accurately identify the high-quality responses that best meet user needs, have a reasonable structure, and rich content, thereby improving the response quality and personalization level of psychological counseling.
[0086] Optionally, for any candidate answer, an intention deviation score is obtained, including: searching for target psychological knowledge from a psychological knowledge base based on the intention of the question corresponding to any candidate answer; determining the intention deviation score based on the semantic similarity between any candidate answer and the target psychological knowledge, wherein the semantic similarity is positively correlated with the intention deviation score.
[0087] The psychology knowledge base is a dataset containing a wide range of psychological theories, concepts, treatment methods, and other information. Target psychological knowledge is relevant knowledge retrieved from the knowledge base based on the user's intent in asking the question. Semantic similarity is a metric that measures the semantic closeness between two texts, used to compare the relevance between alternative answers and target psychological knowledge.
[0088] For example, a Variational Autoencoder (VAE) can be used to transform textual knowledge into representation vectors. That is, a VAE encoder can be used to convert any candidate answer and the target psychological knowledge into representation vectors, and then the semantic similarity between any candidate answer and the target psychological knowledge can be determined by the distance between the two representation vectors. This VAE encoder has been specifically trained with psychological knowledge. Therefore, it can more accurately capture the unique semantic features in psychological texts, and the generated representation vectors will better fit the logical system of psychological knowledge in the latent space.
[0089] This application's embodiments retrieve target psychological knowledge matching the user's question intent from a psychological knowledge base. This target psychological knowledge represents the most professional and accurate theoretical expression of the current intent. If the candidate answer is highly semantically similar to this knowledge, it indicates that the answer closely aligns with the user's intent, and the content is targeted and professional. Conversely, if the semantic similarity is low, it indicates that the answer deviates from the user's core needs. Therefore, determining the degree of intent deviation score based on the semantic similarity between the candidate answer and the target psychological knowledge can accurately reflect the degree of fit between the answer and the intent, thereby ensuring the accuracy and relevance of the answer content.
[0090] Optionally, for any candidate answer, a topic richness score is obtained, including: constructing a codebook based on psychological knowledge, the codebook including multiple semantic vectors, and each semantic vector having a unique semantic identifier; converting the intent of the question corresponding to any candidate answer into an intent vector; searching a set of semantic identifiers based on the vector similarity between the intent vector and each semantic vector in the codebook; and determining the topic richness score corresponding to any candidate answer based on the domain topics corresponding to each semantic identifier in the set of semantic identifiers.
[0091] A codebook is a structured representation of psychological knowledge, composed of multiple semantic vectors. Each vector represents a specific psychological topic or concept and is equipped with a unique semantic ID, such as "root cause of behavior" or "intervention method." Semantic vectors are high-dimensional vectors generated from text in psychological knowledge using vector transformation techniques, such as word embedding, to represent the semantic information of the text. Semantic IDs are unique identifiers assigned to each semantic vector in the codebook, used to associate the vector with the corresponding domain topic. Intent vectors are high-dimensional vectors generated from the intent of a question using vector transformation techniques, used to quantify the semantic features representing the user's intent. Vector similarity is an indicator that measures the closeness between intent vectors and semantic vectors in high-dimensional space, used to determine their semantic relevance, such as cosine similarity. The semantic ID set is a collection of semantic IDs selected from the codebook through vector similarity matching, showing a high degree of semantic relevance to the intent vectors. Domain topics are the psychological professional topics corresponding to the semantic IDs.
[0092] For example, RQ-VAE (Residual-Quantized Variational Autoencoder) can be used to encode psychological knowledge to obtain a codebook. Specifically, this involves: organizing psychological knowledge and extracting multi-dimensional knowledge; converting multi-dimensional knowledge into representation vectors using a residual-quantized variational autoencoder; and then constructing the codebook using vector space quantization technology, which applies vector space quantization to all generated representation vectors to perform clustering and compression. Specifically, similar representation vectors are grouped into the same class using an algorithm, with each class corresponding to a "cluster center vector," and the set of these cluster center vectors constitutes the codebook. Simultaneously, a unique semantic identifier is assigned to each cluster center vector. The multi-dimensional knowledge can include topics, categories, titles, content, and other dimensions, but this embodiment does not limit this.
[0093] Figure 4 This is a schematic diagram of the codebook construction process provided in an embodiment of this application. (Reference) Figure 4 This method extracts subtopic content information from psychological knowledge, including subtopic identifiers, titles, descriptions, and categories. This subtopic content information is then input into a residual quantization variational autoencoder, which converts it into representation vectors. Cluster quantization is then performed, grouping similar representation vectors into the same class, with each class corresponding to a "cluster center vector." A unique semantic identifier is assigned to each cluster center vector, and the set of these cluster center vectors constitutes the codebook.
[0094] Table 1 below lists the compiled knowledge of child psychology, and is for illustrative purposes only.
[0095] Table 1
[0096] For example, the method for determining the topic richness score of any candidate answer based on the domain topics corresponding to each semantic identifier in the semantic identifier set is as follows: First, count the total number of all domain topics corresponding to the semantic identifier set; then, identify the number of domain topics actually covered by the candidate answer; determine the topic coverage rate by the proportion of the actual covered topics to the total number of topics; and determine the topic richness score based on the topic coverage rate. The topic coverage rate and the topic richness score are positively correlated.
[0097] This application's embodiments transform psychological knowledge into a structured codebook, which is then used to quickly filter out multiple domain topics that match the user's intent. This allows for an objective, quantitative assessment of the richness of candidate answer topics based on these multiple domain topics. The codebook, constructed based on psychological knowledge, ensures the professionalism and relevance of the identified topics. By calculating the similarity between the intent vector and the semantic vector in the codebook, domain topics that highly match the user's intent can be accurately selected. Therefore, the topic richness score determined based on these domain topics accurately reflects the breadth and diversity of the answers in the psychological dimension. Combining this topic richness score effectively identifies high-quality answers that are more comprehensive and better meet the user's diverse psychological needs.
[0098] Step S403: Select the answer with the highest reward score from the multiple alternative answers as the answer to the question in this round.
[0099] This application embodiment generates multiple candidate answers using a large language model and introduces a reward score mechanism to evaluate and select the best answer. This allows for the exploration of responses that better meet the user's psychological needs from multiple perspectives, thereby improving the professionalism of the answers. This multi-candidate generation and scoring optimization mechanism not only enhances the flexibility and robustness of the terminal in dealing with complex psychological problems but also helps to increase users' trust and satisfaction with psychological counseling.
[0100] Based on the first embodiment of this application described above, a third embodiment of this application is proposed. Contents that are the same as or similar to the first embodiment can be referred to the above description, and will not be repeated hereafter. See also... Figure 5 In the third embodiment, steps S01 to S03 are included before step S10.
[0101] Step S01: Obtain multi-turn dialogue training data, which includes at least one round of questions.
[0102] Multi-turn dialogue training corpus is a dataset consisting of multiple dialogue rounds, each round containing a question and its corresponding answer, used to train and optimize the performance of large language models in multi-turn dialogues.
[0103] Step S02: Use the large language model to answer at least one round of questions sequentially, and determine the reward score corresponding to the generated at least one round of answers.
[0104] At least one round of questions refers to a set of one or more sets of questions posed by users within the training corpus, constituting a complete dialogue flow used to simulate the interaction process in real-world psychological counseling scenarios. At least one round of responses refers to the response content generated by the large language model in response to at least one round of questions.
[0105] It should be noted that the method for determining the reward score for each round of answers is the same as the method for determining the reward score for each candidate answer in the above embodiment. That is, determining the reward score for at least one generated round of answers includes: for any round of answers, obtaining at least one of the following scores: single-round dialogue word count score, multi-round dialogue round count score, intention deviation score, and topic richness score; and performing a weighted sum of the at least one score to obtain the reward score for any round of answers.
[0106] Step S03: Based on the reward scores corresponding to at least one round of responses, train a large language model to maximize the reward scores of the large language model.
[0107] Training a large language model involves adjusting its parameters through a feedback mechanism, enabling it to generate high-quality, high-reward-score responses more effectively when faced with new psychological counseling questions.
[0108] Optionally, after determining the reward score corresponding to the generated at least one round of answers, the method further includes: performing multiple round sampling on the at least one round of answers generated by the model, each round sampling being used to extract at least one round of answers from the at least one round of answers; calculating information entropy based on the reward score corresponding to the at least one round of answers in each round sampling; determining a target round sampling in the multiple round sampling where the corresponding information entropy is greater than the information entropy threshold; and constructing a new multi-round dialogue training corpus based on the at least one round of answers corresponding to the target round sampling and the question corresponding to the at least one round of answers.
[0109] Round sampling is the process of randomly or systematically selecting a portion of dialogue rounds from at least one round of responses generated by the model. Information entropy is used to measure the uncertainty of a set of sampled responses in the reward distribution. The higher the information entropy, the greater the diversity and the more dispersed the distribution; the lower the information entropy, the higher the consistency and the less the variation.
[0110] The information entropy threshold is a preset value used to determine whether the information entropy of a particular round of sampling is at a high level. If it is higher than this threshold, the sample is considered to have high response diversity and potential learning value. Target round sampling refers to the sampling results in multiple rounds where the corresponding information entropy is greater than the set information entropy threshold, indicating that the responses it contains have high diversity and learning value.
[0111] Constructing new multi-turn dialogue training corpora involves combining high-quality and diverse responses selected from the target number of turns of sampling with their corresponding questions to form new dialogue samples, which are used to expand the original training dataset and improve the breadth and generalization ability of model training.
[0112] In this embodiment, by sampling the generated responses multiple times and combining this with information entropy analysis to select dialogues with high diversity and learning value, a new, more representative multi-turn dialogue training corpus is constructed. This information entropy-driven data augmentation mechanism can effectively identify dialogue segments that exhibit inconsistent reward distributions but possess potential learning significance, preventing training data from falling into a single pattern or local optima. Training the model with new, high-value multi-turn dialogue training corpus can further enhance the generalization, adaptability, and creative expression capabilities of the large language model in psychological counseling scenarios, enabling it to generate more accurate, richer, and more personalized content when facing complex and diverse psychological counseling issues.
[0113] The above description of this embodiment illustrates the process of reinforcement learning training for a large language model. Prior to this, training corpus needs to be acquired. Specifically, seed question-answer pairs are obtained; expanded question-answer pairs are generated based on the seed pairs using the teacher large model; target question-answer pairs that meet quality criteria are selected from the expanded pairs; and the seed and target question-answer pairs are used as training corpus for the large language model. The seed question-answer pairs are an initial, representative set of question-answer samples, which can be generated from professional psychological counseling data or expert annotations, serving as the basis for subsequent question-answer pair expansion. The teacher large model is a high-performance large language model with strong language understanding and generation capabilities, used to assist in generating more high-quality question-answer pairs. Expanded question-answer pairs are additional question-answer samples generated by the teacher large model based on the seed pairs, aiming to expand the original corpus and improve the diversity and coverage of the training data. Quality criteria are a set of standards used to select generated question-answer pairs, ensuring the professionalism and usability of the generated content. The training corpus is a dataset used to train the large language model, containing multiple question-answer pairs. The model learns from this data to improve its performance in psychological counseling dialogues.
[0114] For example, a large language model can be used to filter out target question-answer pairs that meet the quality criteria from the expanded question-answer pairs. For example, prior to this, textual lexical and semantic methods can be used to de-duplicate the expanded question-answer pairs and perform preliminary low-quality filtering.
[0115] Optionally, based on seed question-answer pairs, an expanded question-answer pair is generated using a large teacher model. This includes: clustering the seed question-answer pairs to obtain clustering results, which include multiple category semantic centers; generating psychological counseling questions related to the category semantic centers using the large teacher model, and generating corresponding answers to the psychological counseling questions; and combining the psychological counseling questions and answers generated by the large teacher model to form expanded question-answer pairs. Clustering involves grouping similar question-answer pairs into one category, forming multiple clusters. The category semantic center is the "semantic center point" of all question-answer pairs in each cluster, which can be understood as the core semantic expression of that type of question, representing the main characteristics of a certain type of psychological counseling topic.
[0116] It should be noted that, to improve the quality of the training corpus, manual assistance can be used to filter the data during the acquisition of the training corpus through the teacher-led model. For example, after generating psychological counseling questions related to the semantic center of the category using the teacher-led model, high-quality psychological counseling questions can be selected from the generated questions through manual labeling for use in the subsequent response generation by the teacher-led model. Similarly, after generating the responses to each psychological counseling question, high-quality responses can be selected from the generated responses through manual labeling for use in constructing the training data.
[0117] This application's embodiments extract multiple category semantic centers with representative semantic features through cluster analysis of seed question-answer pairs. Then, using a large teacher model, new psychological counseling questions and their professional answers are generated around these category semantic centers, thereby constructing high-quality expanded question-answer pairs. This approach not only mines latent semantic structures from existing data but also ensures the professionalism and diversity of the generated content. Compared to simply rewriting seed data, the question-answer pairs generated by this method are closer to real-world psychological counseling scenarios and can cover a wider range of psychological issues.
[0118] This application's embodiments utilize a large-scale teacher model to generate diverse extended question-answer pairs based on existing seed question-answer pairs. Then, by filtering out target question-answer pairs that meet set quality criteria, a rich and professional training corpus is constructed. This effectively solves the problem of insufficient high-quality training data in the field of psychological counseling. This approach not only reduces the cost of manual annotation but also improves the generalization ability and response quality of the large language model in psychological counseling scenarios.
[0119] It should be noted that the training corpus obtained using the large teacher model can be used for both fine-tuning and reinforcement learning training of the large language model, and this application does not limit this.
[0120] Optionally, to ensure the compliance and security of the content generated by the model, the large language model needs to undergo secure alignment training before providing psychological counseling services. Specifically, the terminal obtains secure training corpus, in which neither the questions nor the answers contain sensitive content; based on the secure training corpus, the large language model is fine-tuned. The secure training corpus is a set of filtered and processed question-and-answer data, in which neither the questions nor the answers contain any illegal or sensitive content, ensuring the security and compliance of the content during model training.
[0121] For example, LoRA (Low-Rank Adaptation) technology can be used to fine-tune the training of large language models, enhancing their safety alignment capabilities while minimizing the forgetting of psychological knowledge. LoRA is an efficient model fine-tuning technique that freezes most of the model's parameters and only adds low-rank matrix parameters to specific layers for training. This preserves the model's original knowledge (i.e., psychological knowledge) during fine-tuning while allowing the model to adapt to new task requirements with minimal computational cost, thus enhancing safety in psychological counseling scenarios.
[0122] For example, a sparse mask can be selected based on the logit weight threshold of each layer of the model. That is, a threshold is set, and the logit weights of each layer are compared to this threshold. The mask portion corresponding to weights less than the threshold is set to 0, forming a sparse mask. When training the model using a secure training corpus for LoRA, the parameters in the low-rank matrix whose corresponding logit weights are not less than the threshold are updated. These parameters are related to security knowledge. For parameters storing core knowledge in the field of psychology that do not need to be changed, their update magnitude is limited or updates are prevented through the sparse mask. This allows the model to learn secure alignment knowledge while preserving its knowledge of child psychology to the maximum extent, achieving targeted improvement in model capabilities.
[0123] Figure 6 This is a schematic diagram illustrating fine-tuning training of a large language model according to an embodiment of this application. (Reference) Figure 6 During the fine-tuning training of the large language model using a secure training corpus, the backbone weights and parameter A in the right-hand dense connection layer are frozen, and only parameter B in the right-hand sparse connection layer can be adjusted.
[0124] Figure 7 This is a schematic diagram of the neural network structure of a safety adapter provided in an embodiment of this application. It corresponds to the above... Figure 6 The sparse and dense connection layers in the reference. Figure 7During the fine-tuning training of a large language model using a safety training corpus, parameters of the densely connected layers labeled "Safety in Child Psychology" and "Safety in Developmental Psychology" within the model's safety adapter are frozen; only the corresponding parameters of the sparsely connected layers within the safety adapter can be adjusted. In this way, the model's core knowledge in areas such as child psychology and developmental psychology is preserved, while simultaneously enhancing its safety capabilities in a targeted manner.
[0125] This application's embodiments effectively ensure the content compliance and ethical safety of the large language model when generating responses by introducing a safe training corpus and fine-tuning the training based on it. This helps to provide psychological counseling services that conform to social norms and improve user experience and trust.
[0126] This application's embodiments train a large language model by introducing multi-turn dialogue training corpora and combining a reward score mechanism to evaluate and provide feedback on the model's generated responses, thereby achieving continuous model optimization. During this process, the model continuously learns how to generate responses that are both professional and relevant to user needs, while also covering a wide range of topics. This reward-driven training method not only improves the model's comprehension and expressive abilities in psychological counseling scenarios but also enhances its adaptability and stability in complex contexts.
[0127] This application addresses the challenges in psychology, including the high cost of dataset expansion and the inability to collect rich, high-quality dialogue data in real-time to adequately train large models with strong generalization capabilities. Furthermore, in multi-turn dialogue scenarios, as the conversation progresses, issues such as intent deviation, content violations, and topic limitations may arise. Therefore, this application utilizes a combination of multiple technical solutions described in the above embodiments to effectively solve these technical problems.
[0128] Figure 8 This is a schematic diagram illustrating reinforcement learning training of a large language model, as provided in an embodiment of this application. (Reference) Figure 8 For multi-turn dialogue training corpora, the large language model infers based on the current weights to generate answers to questions in each turn. Then, a reward score is determined for each answer, covering factors such as word count in a single turn, number of turns in a multi-turn dialogue, degree of intent deviation, topic richness, and correctness. Based on the reward scores for each answer, the model loss is calculated, and the model weights are adjusted to maximize the reward.
[0129] Another point to note is that the above examples are only for understanding this application and do not constitute a limitation on the multi-round dialogue method of psychological counseling in this application. Any simple modifications based on this technical concept are within the scope of protection of this application.
[0130] This application also provides a multi-turn dialogue device for psychological counseling; please refer to [reference needed]. Figure 9 The multi-round dialogue device for psychological counseling includes: Question extraction module 10 is used to extract the current question from the multi-turn dialogue command in response to the multi-turn dialogue command; The intent recognition module 20 is used to perform intent recognition on the current round of questions based on the historical intent corresponding to at least one round of historical questions, so as to obtain the current round intent corresponding to the current round of questions. Topic generation module 30 is used to generate multiple psychology topics based on the intent of this round; The answer generation module 40 is used to generate answers to the questions in this round based on the intent of this round and multiple psychological topics through a large language model.
[0131] Optionally, the answer generation module includes: The answer generation unit is used to generate multiple alternative answers based on the current intent and multiple psychological topics using a large language model; The score determination unit is used to determine the reward score for each candidate answer; The answer selection unit is used to select the answer with the highest reward score from multiple alternative answers as the answer to the question in this round.
[0132] Optionally, the score determination unit is used to obtain at least one of the following scores for any candidate answer: single-turn dialogue word count score, multi-turn dialogue round count score, intention deviation score, and topic richness score; and to perform a weighted summation of the at least one score to obtain the reward score corresponding to any candidate answer.
[0133] Optionally, the score determination unit is used to search for target psychological knowledge from a psychological knowledge base based on the intent of the question corresponding to any candidate answer; and to determine the degree of intent deviation score based on the semantic similarity between any candidate answer and the target psychological knowledge, wherein the semantic similarity is positively correlated with the degree of intent deviation score.
[0134] Optionally, the score determination unit is used to construct a codebook based on psychological knowledge. The codebook includes multiple semantic vectors, and each semantic vector has a unique semantic identifier. The intent of the question corresponding to any candidate answer is converted into an intent vector. Based on the vector similarity between the intent vector and each semantic vector in the codebook, a set of semantic identifiers is searched. Based on the domain topics corresponding to each semantic identifier in the set of semantic identifiers, the topic richness score corresponding to any candidate answer is determined.
[0135] Optionally, the intent recognition module 20 is used to obtain intent recognition prompts, which include historical intents corresponding to at least one round of historical questions; based on the intent recognition prompts, the intent of the current round of questions is recognized through a large language model to obtain the intent of the current round of questions.
[0136] Optionally, the intent recognition module 20 is used to identify omitted questions not covered by historical answers from at least one round of historical questions; and to perform intent recognition on the omitted questions and the current round of questions based on the historical intent corresponding to at least one round of historical questions to obtain the intent of the current round.
[0137] Optionally, the topic generation module 30 is used to obtain rich topic prompts, which contain the intent of the current round; based on the rich topic prompts, multiple psychological topics related to the intent of the current round are generated through a large language model.
[0138] Optionally, the topic generation module 30 is used to determine the content coverage of at least one round of historical answers to at least one round of historical questions; based on the content coverage, a multi-round dialogue round score is generated, and the content coverage is positively correlated with the multi-round dialogue round score; if the multi-round dialogue round score is less than the round score threshold, multiple psychology topics are generated based on the intent of the current round.
[0139] Optionally, the device further includes: The first training module is used to acquire multi-turn dialogue training corpus, which includes at least one round of questions; to use a large language model to answer the at least one round of questions sequentially and determine the reward score corresponding to the generated at least one round of answers; and to train the large language model based on the reward score corresponding to the at least one round of answers so that the large language model maximizes the reward score.
[0140] Optionally, the device further includes: The first corpus generation module is used to perform multiple rounds of sampling on at least one round of responses. Each round of sampling is used to extract at least one round of responses from at least one round of responses. The information entropy is calculated based on the reward score corresponding to the at least one round of responses for each round of sampling. The target round of sampling is determined when the information entropy is greater than the information entropy threshold in the multiple rounds of sampling. Based on the at least one round of responses corresponding to the target round of sampling and the questions corresponding to the at least one round of responses, a new multi-round dialogue training corpus is constructed.
[0141] Optionally, the device further includes: The second corpus generation module is used to obtain seed question-answer pairs; based on the seed question-answer pairs, expanded question-answer pairs are generated using the teacher's large model; target question-answer pairs that meet the quality conditions are selected from the expanded question-answer pairs; and the seed question-answer pairs and target question-answer pairs are used as training corpus for the large language model.
[0142] Optionally, the second corpus generation module is used to cluster the seed question-answer pairs to obtain clustering results, which include multiple category semantic centers; through the teacher big model, psychological counseling questions related to the category semantic centers are generated, and answers to the psychological counseling questions are generated; the psychological counseling questions and answers generated by the teacher big model are combined to form an expanded question-answer pair.
[0143] Optionally, the device further includes: The second training module is used to acquire a secure training corpus, in which the questions and answers do not contain sensitive content; and to fine-tune the large language model based on the secure training corpus.
[0144] The multi-turn dialogue device for psychological counseling provided in this application, employing the multi-turn dialogue method for psychological counseling in the above embodiments, can solve the technical problem in related technologies where the output content of the model during psychological counseling focuses on a single dimension and has low topic richness. Compared with the prior art, the beneficial effects of the multi-turn dialogue device for psychological counseling provided in this application are the same as those of the multi-turn dialogue method for psychological counseling provided in the above embodiments, and other technical features in the multi-turn dialogue device for psychological counseling are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0145] This application provides a multi-turn dialogue device for psychological counseling, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the multi-turn dialogue method for psychological counseling in the above embodiments.
[0146] The following is for reference. Figure 10 The diagram illustrates a structural schematic suitable for implementing a multi-turn dialogue device for psychological counseling in the embodiments of this application. The multi-turn dialogue device for psychological counseling in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 10 The illustrated multi-turn dialogue device for psychological counseling is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0147] like Figure 10As shown, the multi-turn dialogue device for psychological counseling may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the multi-turn dialogue device for psychological counseling. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the multi-turn counseling device to communicate wirelessly or wiredly with other devices to exchange data. While the figure shows multi-turn counseling devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0148] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0149] The multi-turn dialogue device for psychological counseling provided in this application, employing the multi-turn dialogue method for psychological counseling in the above embodiments, can solve the technical problem in related technologies where the output content of the model during psychological counseling focuses on a single dimension and has low topic richness. Compared with the prior art, the beneficial effects of the multi-turn dialogue device for psychological counseling provided in this application are the same as those of the multi-turn dialogue method for psychological counseling provided in the above embodiments, and other technical features in this multi-turn dialogue device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0150] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0151] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0152] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the multi-turn dialogue method for psychological counseling in the above embodiments.
[0153] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0154] The aforementioned computer-readable storage medium may be included in a multi-turn dialogue device for psychological counseling; or it may exist independently and not be assembled into a multi-turn dialogue device for psychological counseling.
[0155] The aforementioned computer-readable storage medium carries one or more programs that, when executed by a multi-turn dialogue device for psychological counseling, cause the multi-turn dialogue device to: respond to a multi-turn dialogue instruction, extract the current round question from the multi-turn dialogue instruction; perform intent recognition on the current round question based on the historical intent corresponding to at least one previous round question, and obtain the current round intent corresponding to the current round question; generate multiple psychological topics based on the current round intent; and generate an answer corresponding to the current round question through a large language model, based on the current round intent and the multiple psychological topics.
[0156] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0157] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0158] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0159] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the aforementioned multi-turn dialogue method in psychological counseling. This addresses the technical problem in related technologies where the output content of the model during psychological counseling focuses on a single dimension and has low topic richness. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the multi-turn dialogue method in psychological counseling provided in the above embodiments, and will not be elaborated upon here.
[0160] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the multi-turn dialogue method in psychological counseling as described above.
[0161] The computer program product provided in this application can solve the technical problem in related technologies that the output content of the model in the psychological counseling process focuses on a single dimension and has low topic richness. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the multi-turn dialogue method in psychological counseling provided in the above embodiments, and will not be repeated here.
[0162] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A multi-round dialogue method in psychological counseling, characterized in that, The method includes: In response to a multi-turn dialogue instruction, extract the current round question from the multi-turn dialogue instruction; Based on the historical intent corresponding to at least one round of historical questions, the intent of the current round of questions is identified to obtain the current round intent corresponding to the current round of questions. Based on the stated intent of this round, several psychological topics were generated; Using a large language model, multiple alternative answers are generated based on the stated intent and the multiple psychological topics. For any given answer, obtain at least one of the following scores: word count in a single round of dialogue, number of rounds in a multi-round dialogue, degree of deviation from intent, and richness of topic; The score is obtained by weighted summation of the at least one score, and the corresponding reward score for any candidate answer is obtained. The answer with the highest reward score among the multiple alternative answers will be selected as the answer to the question in this round. For each candidate answer, a topic richness score is awarded, including: A codebook is constructed based on psychological knowledge. The codebook includes multiple semantic vectors, and each semantic vector has a unique semantic identifier. Convert the intent of the question corresponding to any of the alternative answers into an intent vector; Based on the vector similarity between the intent vector and each semantic vector in the codebook, a set of semantic identifiers is searched. Based on the domain topics corresponding to each semantic identifier in the semantic identifier set, the topic richness score corresponding to any candidate answer is determined.
2. The method as described in claim 1, characterized in that, For any given answer, obtain a score indicating the degree of deviation from intent, including: Based on the intent of the question corresponding to any of the alternative answers, search for target psychological knowledge from the psychological knowledge base; Based on the semantic similarity between any of the alternative answers and the target psychological knowledge, the intention deviation score is determined, and the semantic similarity is positively correlated with the intention deviation score.
3. The method as described in claim 1, characterized in that, The step of identifying the intent of the current round of questions based on the historical intent corresponding to at least one round of historical questions, to obtain the current round's intent corresponding to the current round of questions, includes: Obtain intent recognition prompts, wherein the intent recognition prompts include historical intents corresponding to at least one round of historical questions; Based on the intent recognition prompts, the intent of the current round question is obtained by using the large language model to identify the intent of the current round question.
4. A multi-turn dialogue device for psychological counseling, characterized in that, The device includes: The question extraction module is used to extract the current question from the multi-turn dialogue instructions in response to the multi-turn dialogue instructions; The intent recognition module is used to perform intent recognition on the current round question based on the historical intent corresponding to at least one round of historical questions, so as to obtain the current round intent corresponding to the current round question. The topic generation module is used to generate multiple psychology topics based on the stated intent of this round. The answer generation module is used to generate answers to the questions in this round based on the current intent and the multiple psychological topics using a large language model. The answer generation module includes: The answer generation unit is used to generate multiple alternative answers based on the current intent and the multiple psychological topics using a large language model; The scoring unit is used to obtain at least one of the following scores for any candidate answer: single-round dialogue word count score, multi-round dialogue round count score, intention deviation score, and topic richness score; and to perform a weighted summation of the at least one score to obtain the reward score corresponding to any candidate answer. The answer selection unit is used to select the answer with the highest reward score from the multiple alternative answers as the answer to the question in this round. The scoring unit is configured to construct a codebook based on psychological knowledge, the codebook including multiple semantic vectors, each semantic vector having a unique semantic identifier; convert the intent of the question corresponding to any candidate answer into an intent vector; search a set of semantic identifiers based on the vector similarity between the intent vector and each semantic vector in the codebook; and determine the topic richness score corresponding to any candidate answer based on the domain topic corresponding to each semantic identifier in the set of semantic identifiers.
5. A multi-turn dialogue device for psychological counseling, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the multi-turn dialogue method for psychological counseling as described in any one of claims 1 to 3.
6. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the multi-round dialogue method for psychological counseling as described in any one of claims 1 to 3.
7. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the multi-turn dialogue method for psychological counseling as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Response statement generation method and device, storage medium and electronic equipment
CN117874197A
Preference alignment training method of large language model LLM, electronic equipment and storage medium
CN120069082A