Two-Stage Retrieval-Enhanced LLM-Based Psychological Support Q&A Generation Method and System

By conducting two-stage searches in the external resource library to obtain professional and targeted answers, the problem of difficult to control the generation of psychological support answers by LLM is solved, and more efficient professionalism and personalized answer generation is achieved, improving the quality of psychological support answers of LLM.

CN119760102BActive Publication Date: 2025-06-17豫章师范学院
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510261082.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-17
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The prior art is difficult to accurately control the professionalism and pertinence of psychological support answers generated by large language models (LLM), and the high consumption of hardware resources by fine-tuning LLMs limits its application promotion.

Method used

A psychological support question-and-answer generation method based on two-stage search enhancement LLM is proposed. By conducting two-stage search in an external resource library, professional resources are obtained and LLM is guided to generate more professional and personalized psychological support answers.

Benefits of technology

The professionalism, relevance and empathy ability of generating answers has been significantly improved, and the overall performance of LLM's automatic psychological support answers has been improved. The evaluation results show that the BLEU value has been significantly improved, and the help and empathy have also been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119760102B_ABST
    Figure CN119760102B_ABST
Patent Text Reader

Abstract

The present invention proposes a method and system for generating psychological support Q&A based on two-stage retrieval enhanced LLM. In the first stage, the original question is used as a retrieval query condition to perform a retrieval operation from an external resource library. Then, the original question and the retrieval results are used as context cues to construct a prompt and input it into the LLM to generate several generated answers. The original question and the generated answers are fused as a retrieval query condition to perform a second-stage retrieval operation from the external resource library. Then, the original question and the retrieval results are used as context cues to construct a prompt and input it into the LLM to generate the final answer, obtaining a psychological support answer. Aiming at the problem that simply using the LLM to generate psychological support answers often makes it difficult to accurately control the professionalism and pertinence of the generated answers, through the two-stage retrieval enhancement method, professional resources are obtained from the external resource library to guide the LLM to generate more professional and personalized psychological support answers, so as to improve the overall performance of the LLM for automatically generating psychological support answers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and particularly to a method and system for generating psychological support Q&A enhanced by two-stage retrieval based on a large language model (LLM). Background Art

[0002] Due to the fact that traditional community interaction models often fail to meet users' needs for immediacy and personalization, the anonymity and convenience of online psychological support platforms have led more and more people with psychological or emotional distress to be inclined to confide their feelings on these platforms. However, the rapid development of the platforms has also increased users' requirements for service quality and efficiency. Facing a vast amount of user questions and diverse mental health problems, professional psychological counselors and community managers often feel overwhelmed and find it difficult to respond quickly and comprehensively. This not only affects the user experience but may also result in the failure to meet users' needs in a timely manner in some emergency situations. Against this background, the psychological support automatic Q&A system has emerged. With its characteristics of instant response and round-the-clock service, it has become a new way to solve mental health problems. Compared with manual consultation services, the psychological support automatic Q&A system can quickly respond to user questions, provide personalized psychological support and suggestions, greatly improving the convenience and accessibility of the service. In addition, through artificial intelligence technology, the automatic Q&A system can process a large amount of user data, analyze mental health trends, and provide more accurate services for users.

[0003] A large language model (LLM) is a natural language processing technology based on deep learning. By training a large-scale corpus with model parameters in the billions to hundreds of billions, it can understand and generate natural language text by learning the structure, grammar, and semantics of language. Currently, commonly used LLMs include foreign ones such as the GPT series (GPT-3, GPT-3.5, GPT-4, GTP-4o, etc.), the LLaMA series, and Claude, as well as domestic ones such as Wenxin Yiyan, Zhipu Qingyan, Tongyi Qianwen, Doubao, and Baichuan. With the rapid development and wide application of LLMs, LLMs for psychological support have also attracted more and more attention from scholars. Some scholars have applied a theory based on five elements in human emotions: self-awareness, self-regulation, motivation, empathy, and social skills, through Chain-of-Thought (CoT), to LLM prompt engineering to guide the LLM to generate more empathetic psychological support answers. Some scholars have combined cognitive behavioral therapy (CBT) in psychology with LLM prompt engineering to construct a psychological support Q&A large language model specifically designed for cognitive behavioral therapy. There are also scholars who have fine-tuned LLMs using a psychological support dialogue dataset to construct an empathy dialogue LLM containing multiple emotional support scenarios.

[0004] However, it is often difficult to accurately control the professionalism and pertinence of the generated answers by prompting the LLM to generate psychologically supportive answers. To prompt the LLM to generate more professional and personalized answers, users usually need to add some content related to psychology majors in the prompt. This method has caused new problems; the more prompt information is provided, the more likely the generated supportive answers will only focus on describing the given content, restricting the personalization and pertinence of the answers. At the same time, the method of fine-tuning the LLM to generate psychologically supportive answers is limited by the high requirements and high consumption of hardware resources, restricting its popularization and application in actual use. Summary of the Invention

[0005] In view of the above situation, the main purpose of the present invention is to propose a method and system for generating psychologically supportive Q&A enhanced by two-stage retrieval for the LLM to solve the above technical problems.

[0006] The present invention proposes a method for generating psychologically supportive Q&A enhanced by two-stage retrieval for the LLM, and the method includes the following steps:

[0007] Step 1, providing an external resource library and an original question;

[0008] Step 2, performing a first-stage retrieval operation from the external resource library using the original question as the retrieval query condition;

[0009] During the first-stage retrieval process, perform a preliminary screening from the external resource library according to the original question to obtain candidate answers, and then perform another screening according to the local word order information and global semantic information of each word in the original question in the candidate answers to obtain the first candidate answer;

[0010] Step 3, using the original question and the first candidate answer as context prompts to construct a prompt, and inputting it into the LLM to generate several generated answers;

[0011] Step 4, fusing the original question and the generated answers as the retrieval query condition to perform a second-stage retrieval operation from the external resource library to obtain a second candidate answer;

[0012] Step 5, using the original question and the second candidate answer as context prompts to construct a prompt, and inputting it into the LLM to generate the final answer to obtain a psychologically supportive answer.

[0013] The present invention also proposes a system for generating psychologically supportive Q&A enhanced by two-stage retrieval for the LLM. Among them, the system applies the method for generating psychologically supportive Q&A enhanced by two-stage retrieval for the LLM as described above, and the system includes:

[0014] A data acquisition module, used for:

[0015] providing an external resource library and an original question;

[0016] The first-stage retrieval module is used for:

[0017] Performing a first-stage retrieval operation from an external resource library using the original question as the retrieval query condition;

[0018] During the first-stage retrieval process, perform a preliminary screening from the external resource library according to the original question to obtain candidate answers, and then perform another screening according to the local word order information and global semantic information of each word in the original question in the candidate answers to obtain the first candidate answer;

[0019] The first-stage generation module is used for:

[0020] Using the original question and the first candidate answer as context cues to construct a prompt, and inputting it into the LLM to generate several generated answers;

[0021] The second-stage retrieval module is used for:

[0022] Fusing the original question and the generated answer as the retrieval query condition to perform a second-stage retrieval operation from the external resource library to obtain the second candidate answer;

[0023] The second-stage generation module is used for:

[0024] Using the original question and the second candidate answer as context cues to construct a prompt, and inputting it into the LLM to generate the final answer to obtain the psychological support answer.

[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0026] In view of the problem that simply using the LLM to generate psychological support answers often makes it difficult to accurately control the professionalism and pertinence of the generated answers, the present invention proposes a two-stage retrieval enhancement method. Professional resources are obtained through retrieval from the external resource library, and the professional resources are effectively utilized to guide the LLM to generate more professional and personalized psychological support answers, thereby improving the overall performance of the LLM to automatically generate psychological support answers. The professionalism, relevance, and empathy ability of the generated answers are effectively improved. In terms of automatic evaluation, BLEU, ROUGE, D-1, and D-2 metrics are used for comprehensive evaluation. The evaluation results show that the BLEU value of the two-stage retrieval enhancement method proposed by the present invention has been significantly improved, indicating that the generated psychological support answers are closer to the reference text at the lexical and phrase levels, and the accuracy and standardization of the answers have been improved. In terms of manual evaluation, metrics such as pertinence, relevance, helpfulness, and empathy are used for evaluation. The evaluation results find that the present invention has achieved good improvement in terms of helpfulness and empathy, indicating that the present invention can more effectively identify and solve the psychological problems of the seekers, the suggestions and strategies provided are more in line with the needs of the seekers, and at the same time, there is also good improvement in understanding and caring for human emotions.

[0027] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a flowchart of a method for generating psychological support question and answer based on two-stage retrieval enhanced LLM proposed by the present invention;

[0029] Figure 2 It is an architecture diagram of a method for generating psychological support question and answer based on two-stage retrieval enhanced LLM proposed by the present invention;

[0030] Figure 3 It is a schematic structural diagram of a system for generating psychological support question and answer based on two-stage retrieval enhanced LLM proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.

[0032] These and other aspects of the embodiments of the present invention will be clear with reference to the following description and drawings. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0033] Please refer to Figure 1 and Figure 2 , this embodiment provides a method for generating psychological support question and answer based on two-stage retrieval enhanced LLM, and the method includes the following steps:

[0034] Step 1: Given an external resource library and an original question;

[0035] Step 2: Perform a first-stage retrieval operation from the external resource library using the original question as a retrieval query condition to obtain a first candidate answer;

[0036] As a preferred embodiment of the present invention, the following relational expression exists in the process corresponding to performing the first-stage retrieval operation from the external resource library using the original question as a retrieval query condition:

[0037]

[0038] Wherein, represents the first-stage retrieval operation, Represents the original question, Represents according to the original question q Retrieved from the external resource library, the first k candidate answer formed by the first several candidate answers, Represents the set of questions in the external resource library, Represents the length of.

[0039] As a preferred embodiment of the present invention, the first-stage retrieval operation specifically includes the following steps:

[0040] Taking the original question as the retrieval query condition, using the BM25 retrieval method to calculate the similarity between the original question and each question in the external resource library, and obtaining the BM25 score. The corresponding process has the following relational formula:

[0041] ;

[0042] Among them, Represents a word in the original question q in, Represents the word t in the question the word frequency, Represents the average document length of the entire set, Represents two different hyperparameters, Represents the question the length of, Represents the word t the inverse document frequency, the word t the calculation process of the inverse document frequency of has the following relational formula:

[0043] ;

[0044] Among them, Represents the logarithmic function, ; Represents the number of questions in the external resource library that contain the word t in.

[0045] According to the BM25 score, obtain the first several candidate answers ;

[0046] Traverse the candidate answers to obtain the occurrence positions of the words in the candidate answers;

[0047] According to the occurrence positions of the words in the candidate answers, calculate the minimum distance between each word pair in the original question in the candidate answer and calculate the context window score according to the minimum distance and the word occurrence frequency in the candidate answer. The corresponding process has the following relational formula:

[0048] ;

[0049] wherein, represents a word pair, represents a number of candidate answers obtained according to the BM25 score, represents the context window score, represents the word in the candidate answer the frequency of occurrence, represents the word in the candidate answer the frequency of occurrence, represents the word pair in the candidate answer the minimum distance of occurrence;

[0050] According to the consistency of the word order of the words in the candidate answer with the word order of the original question words, the corresponding word order weights are assigned, and the context window score is adjusted using the word order weights to obtain the adjusted context window score. The corresponding process has the following relational expression:

[0051] ;

[0052] wherein, represents the adjusted context window score, represents the order factor, represents order consistency, represents order inconsistency;

[0053] Accumulate the context window scores of all word pairs to obtain the final context window score. The corresponding process has the following relational expression:

[0054] ;

[0055] wherein, represents the final context window score;

[0056] Use the pre-trained DPR model to encode the original question and the candidate answer respectively to generate the semantic embedding vectors of the original question and the candidate answer. The corresponding process has the following relational expression:

[0057] ;

[0058] wherein, represents the query encoder, represents the passage encoder, represents the semantic embedding vector of the original question, represents the semantic embedding vector of the candidate answer;

[0059] The context window score of the candidate answer is spliced as a scalar feature into the semantic embedding vector of the candidate answer to obtain an enhanced semantic embedding vector of the candidate answer. The corresponding process has the following relational expression:

[0060] ;

[0061] Among them, represents the enhanced semantic embedding vector of the candidate answer, represents the splicing operation;

[0062] According to the semantic embedding vector of the original question and the enhanced semantic embedding vector of the candidate answer, the cosine similarity is used to calculate the semantic relevance between the original question and the candidate answer, and a semantic embedding score is obtained. The corresponding process has the following relational expression:

[0063] ;

[0064] Among them, represents the semantic embedding score, represents the norm;

[0065] The BM25 score, the final context window score, and the semantic embedding score are combined to obtain the final candidate answer score. The corresponding process has the following relational expression:

[0066] ;

[0067] Among them, represents the final candidate answer score, represents the BM25 score, represents the final context window score, represents the semantic embedding score, respectively represent the weight coefficients corresponding to the BM25 score, the final context window score, and the semantic embedding score, represents the original question, represents each question in the external resource library;

[0068] According to the final candidate answer score, the top k candidate answers are selected as the first candidate answers .

[0069] In the above solution, in order to find a balance between efficiency and effectiveness, the present invention adopts a phased retrieval strategy, that is, candidate answers are quickly screened out by an efficient algorithm first, and then these candidate answers are refined by applying an algorithm. This method can significantly reduce the computational cost while retaining high retrieval performance.

[0070] Among them, BM25 is a very efficient retrieval algorithm. Based on the inverted index calculation, it can quickly process large-scale corpora. Through the method of staged retrieval, by using the efficient initial screening ability of BM25 and the deep fine-ranking ability of the context window and semantic embedding, efficient and accurate candidate answer retrieval can be achieved under large-scale corpora.

[0071] To ensure better screening effects, in the process of calculating the context window score in the present invention, word order and word frequency are considered to take into account both global semantic relevance and local word arrangement relationships, which is more applicable to phrase-based scenarios and further enhances its effects and expressive capabilities.

[0072] Moreover, the context window score is a scalar feature, and the concatenation operation is simple, having a relatively small impact on the model's computational overhead. To ensure better performance in phrase-based scenarios, the context window score is concatenated with the embedding vectors of words and candidate answers. The context window score provides the positional relationship of the query words in the candidate answers, which can help the model make a more accurate evaluation of the matching of phrases and word combinations. The concatenated vector contains both the global semantic information of the DPR model and the local word order information of the context window, and can more comprehensively measure the relevance of candidate answers. It has a significant effect on phrase-based queries and multi-word queries and can better match relevant candidate answers. And in the ranking task, the k relevance of the top candidate answers is improved, and the user experience is enhanced.

[0073] Step 3: Use the original question and the first candidate answer as context cues to construct a prompt, and input it into the LLM to generate several generated answers;

[0074] As a preferred embodiment of the present invention, in the process of using the original question and the first candidate answer as context cues to construct a prompt and inputting it into the LLM to generate several generated answers, the following relational expressions exist:

[0075] ;

[0076] Among them, represents the generated answer, represents the large language model, represents the task instruction, represents using the task instruction s , the original question q and the first candidate answer retrieved to construct the prompt, represents the n th answer in the generated answers.

[0077] An example of the LLM prompt of the present invention is shown in the following table.

[0078]

[0079] Step 4: Integrate the original question and the generated answer as a retrieval query condition to perform a second-stage retrieval operation from an external resource library to obtain a second candidate answer;

[0080] As a preferred embodiment of the present invention, integrating the original question and the generated answer as a retrieval query condition to perform a second-stage retrieval operation from an external resource library to obtain a second candidate answer specifically includes the following steps:

[0081] Perform a concatenation operation on each answer in the original question and the generated answer to obtain an enhanced query. The corresponding process has the following relational expression:

[0082] ;

[0083] where, represents the enhanced query, represents the concatenation operation;

[0084] Based on the enhanced query, perform a second-stage retrieval operation in the external resource library to obtain a second candidate answer. The corresponding process has the following relational expression:

[0085] ;

[0086] Here, represents the second candidate answer formed by the top m candidate answers retrieved from the external resource library according to the enhanced query, represents the length of, represents the second-stage retrieval operation.

[0087] In this embodiment, the second-stage retrieval operation and the first-stage retrieval operation adopt the same method, and there are only differences in the input and the number of candidate answers finally selected to form the second candidate answer. Therefore, no further elaboration is provided.

[0088] Step 5: Use the original question and the second candidate answer as context cues to construct a prompt, and input it into the LLM to generate a final answer to obtain a psychological support answer.

[0089] As a preferred embodiment of the present invention, using the original question and the second candidate answer as context cues to construct a prompt and inputting it into the LLM to generate a final answer to obtain a psychological support answer has the following relational expression:

[0090] ;

[0091] where, represents the final answer, indicating the m candidate answers retrieved from the external resource library by the enhanced query.

[0092] Please refer to Figure 3 , this embodiment also provides a psychological support Q&A generation system based on two-stage retrieval enhanced LLM. Among them, the system applies the psychological support Q&A generation method based on two-stage retrieval enhanced LLM as described above. The system includes:

[0093] A data acquisition module for:

[0094] Given an external resource library and an original question;

[0095] A first-stage retrieval module for:

[0096] Performing a first-stage retrieval operation from the external resource library with the original question as the retrieval query condition;

[0097] During the first-stage retrieval process, perform a preliminary screening from the external resource library according to the original question to obtain candidate answers, and then perform another screening according to the local word order information and global semantic information of each word in the original question in the candidate answers to obtain the first candidate answer;

[0098] A first-stage generation module for:

[0099] Using the original question and the first candidate answer as context prompts to construct a prompt, and inputting it into the LLM to generate several generated answers;

[0100] A second-stage retrieval module for:

[0101] Fusing the original question and the generated answers as the retrieval query condition to perform a second-stage retrieval operation from the external resource library to obtain the second candidate answer;

[0102] A second-stage generation module for:

[0103] Using the original question and the second candidate answer as context prompts to construct a prompt, and inputting it into the LLM to generate the final answer to obtain the psychological support answer.

[0104] In order to more detailedly illustrate the advantages of the present invention over the prior art, in this embodiment, ChatGPT3.5 is used as the entire two-stage retrieval enhanced LLM, and ChatGPT3.5 is operated by calling the API interface provided by OpenAI.

[0105] There are three widely used psychological support Q&A datasets integrated in the external resource library, namely the PsyQA dataset constructed by Tsinghua University, the EfaQA jointly constructed by Stanford University, UCLA, and Fu Jen Catholic University's clinical psychology, and the SoulChatCorpus dataset constructed by South China University of Technology. After data preprocessing, a comprehensive psychological support Q&A resource library covering a wide range of psychological topics and diverse Q&A categories is constructed and named PesQA. The resource library exists in the form of Q&A pairs, including the question content of the seeker and the corresponding answer content of the psychological counselor for the question, serving as an external professional psychological support Q&A resource library for retrieval enhancement.

[0106] The present invention uses two methods, automatic evaluation and manual evaluation, for evaluation. Among them, the automatic evaluation metrics selected are BLEU, ROUGE, Distinct-1 (D1), and Distinct-2 (D2) metrics.

[0107] (1)BLEU (Bilingual Evaluation Understudy): A bilingual evaluation auxiliary tool. The core idea is to compare the overlapping degree of n-gram (consecutive n words or character sequences) between the candidate translation and the reference translation. The higher the overlapping degree, the higher the quality of the translation is considered. Unigram is used to measure the accuracy of word translation, and higher-order n-gram is used to measure the fluency of sentence translation. In the present invention, n = 1 to 4, and then the weighted average of each result is taken. The calculation formula of BLEU is as follows:

[0108] ;

[0109] Among them, P represents the probability that n consecutive words generated are the same. The penalty coefficient ranges from 0 to 1. When the generated result is the same length as the target result, the value is 1. When the generated result is shorter than the target result, the value is less than 1. Weight is a weight assigned to each gram. The value of BLEU ranges from 0 to 1, and the larger the value, the better the generated result.

[0110] (2)ROUGE (Recall-Oriented Understudy for Gisting Evaluation): It can be considered an improved version of BLEU, focusing on recall rather than precision. It mainly calculates how many n-gram phrases in the original reference sentences appear in the output. ROUGE is roughly divided into four types: ROUGE-N (optimizing the precision of BLEU to recall), ROUGE-L (optimizing the n-gramOptimized as the common subsequence), ROUGE-W (giving higher rewards to consecutive matches in ROUGE-L), ROUGE-S (allowing n-gram word skipping). The present invention uses ROUGE-N as the evaluation metric. "N" refers to n- gram , and its calculation method is similar to BLEU, except that BLEU is based on precision, while ROUGE is based on recall.

[0111] (3) Distinct: During the text generation process, the diversity of the text also needs to be pursued. The Distinct metric is used to evaluate the diversity of this text. The definition of Distinct is as follows:

[0112] ;

[0113] where represents the number of non-repeating n-gram in the generated text, represents the total number of n-gram words in the generated text. Distinct-1 represents 1-gram , and Distinct-2 represents 2-gram . The larger the Distinct- n , the higher the generated diversity. The present invention takes Distinct-1 (D-1) and Distinct-2 (D-2) as evaluation metrics.

[0114] To better evaluate the quality of the generated answers, the present invention additionally conducts a human evaluation. In the human evaluation task, the present invention mainly evaluates the generation effect of the psychological support answers generated by the LLM. The human evaluation metrics used in the present invention include:

[0115] (1) Relevance: Whether the answer is relevant and personalized to the user's answer;

[0116] (2) Relevance: Whether the answer is relevant to the user's question;

[0117] (3) Helpfulness: Whether the answer can provide helpful suggestions;

[0118] (4) Empathy: Whether the answer has appropriate emotional or empathetic responses, such as warmth, sympathy, and concern. A ten-point scoring method is adopted, where 1-3 points indicate that the generated answer is "very poor", and 8-10 points indicate that the generated answer is "very good".

[0119] 1. Experimental Results

[0120] The present invention conducts a comparative experiment on the proposed two-stage retrieval enhanced LLM method (abbreviated as index-2), and the comparative methods used are as follows:

[0121] (1)zero-shot: The zero-shot prompting method directly inputs the original question of the applicant and the corresponding task description into the LLM without adding any other contextual attachment conditions. This method relies solely on the LLM's own text understanding and processing capabilities to generate an answer corresponding to the original question. q and the corresponding task description into the LLM without adding any other contextual attachment conditions. This method relies solely on the LLM's own text understanding and processing capabilities to generate an answer corresponding to the original question. q

[0122] (2)few-shot: The few-shot prompting method provides the LLM with several input-output pairs related to the task. In this task, several example posts of similar questions from applicants and corresponding answer posts are input, enabling the LLM to learn the rules of the task through the prompts of the examples and the corresponding task description, and thus generate an answer corresponding to the original question. q

[0123] (3)CoT (Chain of Thought): The Chain of Thought prompting method uses the prompt "Please answer step by step", requiring the LLM to provide a train of thought or reasoning process for answering the question before generating the final answer, so as to better understand and generate the corresponding answer.

[0124] (4)index-1: The one-stage retrieval enhancement method, which is a simplified version of the two-stage retrieval enhancement LLM method (index-2) of the present invention, i.e., only the first-stage retrieval is performed without the second-stage enhancement query. The purpose of conducting the index-1 comparison is to verify the effectiveness of the second-stage enhancement query.

[0125] The experimental results of the automatic evaluation are shown in the following table.

[0126]

[0127] The experimental results of the manual evaluation are shown in the following table.

[0128]

[0129] ​​In view of the problem that generating psychological support answers only by using LLM prompts is difficult to accurately control the professionalism and personalization of the answers, the present invention proposes a method for generating psychological support question and answer based on two-stage retrieval enhanced LLM. First, a professional external psychological support question and answer resource library PesQA is constructed, and then the original question is used as a query condition to perform two-stage retrieval enhanced query in the external resource library PesQA. By adding professional resources in the external resource library, it guides and stimulates LLM to generate more professional and personalized psychological support answers, thereby improving the overall performance of LLM in automatically generating psychological support answers. The present invention has important practical significance for improving the generation effect and quality of the online psychological support automatic question and answer model, as well as the in-depth development of the online psychological support automatic question and answer technology, and at the same time provides useful insights and references for the automation and high-quality development of the online psychological support platform.

[0130] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0131] It should be understood that each part of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits with logic gate circuits for implementing logic functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.

[0132] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0133] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.

Claims

1. A psychological support question and answer generation method based on two-stage retrieval enhanced LLM, characterized in that: The method comprises the following steps: Step 1: Given an external resource library and an original question, the external resource library is in the form of a question-answer pair, which includes the content of the question of the help seeker and the corresponding answer of the psychological counselor to the question; Step 2: Perform the first-stage search operation from the external resource library using the original question as the search query condition to obtain the first candidate answer; Step 3: Use the original question and the first candidate answer as contextual prompts to construct prompts, and input them into the LLM to generate several generated answers; Step 4: The original question and the generated answer are combined as search query conditions to perform a second-stage search operation from an external resource library to obtain a second candidate answer; Step 5: Use the original question and the second candidate answer as contextual prompts to construct prompts, and input them into LLM to generate the final answer to obtain the psychological support answer; The first stage of retrieval operation specifically includes the following steps: The original question is used as the search query condition, and the BM25 search method is used to calculate the similarity between the original question and each question in the external resource library to obtain the BM25 score. The corresponding process has the following relationship: ; in, Represents the original problem q A word in Expressive words t In question The word frequency in represents the average document length of the entire collection, represents two different hyperparameters, Representation problem Length, Expressive words t The inverse document frequency of the word t The calculation process of the inverse document frequency has the following relationship: ; in, represents the logarithmic function, ; Indicates that the external resource library contains the word t The number of questions; Get the first several candidate answers based on BM25 score ; Traverse the candidate answers and obtain the occurrence position of the word in the candidate answers; According to the position of the word in the candidate answer, calculate the position of each word pair in the original question in the candidate answer The minimum distance in the answer is calculated, and the context window score is calculated based on the minimum distance and the frequency of occurrence of the word in the candidate answer. The corresponding process has the following relationship: ; in, Represents a word pair, Represents several candidate answers obtained based on BM25 scores. represents the context window score, Expressive words In the candidate answer The frequency of occurrence in Expressive words In the candidate answer The frequency of occurrence in Representing word pairs In the candidate answer The minimum distance that appears in ; According to the consistency between the word order in the candidate answer and the original question word, the corresponding word order weight is assigned, and the context window score is adjusted using the word order weight to obtain the adjusted context window score. The corresponding process has the following relationship: ; in, represents the adjusted context window score, represents the order factor, Indicates that the order is consistent. Indicates that the order is inconsistent; The context window scores of all word pairs are accumulated to obtain the final context window score. The corresponding process has the following relationship: ; in, represents the final context window score; The pre-trained DPR model is used to encode the original question and candidate answers respectively to generate semantic embedding vectors of the original question and candidate answers. The corresponding process has the following relationship: ; in, represents the query encoder, represents a paragraph encoder, The semantic embedding vector representing the original question, The semantic embedding vector representing the candidate answer; The context window score of the candidate answer is concatenated as a scalar feature into the semantic embedding vector of the candidate answer to obtain the enhanced semantic embedding vector of the candidate answer. The corresponding process has the following relationship: ; in, The semantic embedding vector representing the enhanced candidate answer, Represents a splicing operation; According to the semantic embedding vector of the original question and the semantic embedding vector of the enhanced candidate answer, the cosine similarity is used to calculate the semantic relevance of the original question and the candidate answer to obtain the semantic embedding score. The corresponding process has the following relationship: ; in, represents the semantic embedding score, represents the norm; The BM25 score, the final context window score, and the semantic embedding score are combined to obtain the final candidate answer score. The corresponding process has the following relationship: ; in, represents the final candidate answer score, represents the BM25 score, represents the final context window score, represents the semantic embedding score, They represent the weight coefficients corresponding to the BM25 score, the final context window score, and the semantic embedding score, respectively. Indicates the original problem, Represents each issue in the external repository; According to the final candidate answer score, select the top- k The candidate answer is the first candidate answer .

2. The psychological support question and answer generation method based on two-stage retrieval enhanced LLM according to claim 1 is characterized in that: In step 2, the first-stage search operation is performed from the external resource library using the original question as the search query condition, and the process of obtaining the first candidate answer corresponds to the following relationship: ; in, Represents the first stage of retrieval operation, Indicates the original problem, According to the original question q Top- k The first candidate answer is formed by candidate answers. represents a collection of issues in an external repository, express Length.

3. The psychological support question and answer generation method based on two-stage retrieval enhanced LLM according to claim 2 is characterized in that: In step 3, using the original question and the first candidate answer The process of constructing a prompt as a contextual prompt and inputting it into the LLM to generate several generated answers has the following relationship: ; in, Generates an answer. represents a large language model, Indicates task instructions. Indicates the use of task instructions s , original question q and the first candidate answer retrieved The built prompt, Indicates the first n Answers.

4. The psychological support question and answer generation method based on two-stage retrieval enhanced LLM according to claim 3 is characterized in that: In step 4, the original question and the generated answer are merged as search query conditions to perform a second-stage search operation from an external resource library, and obtaining the second candidate answer specifically includes the following steps: The original question and each of the generated answers are connected to obtain an enhanced query. The corresponding process has the following relationship: ; in, Indicates enhanced query, Indicates a connection operation; Based on the enhanced query, the second stage of retrieval operation is performed in the external resource library to obtain the second candidate answer. The corresponding process has the following relationship: ; here, Indicates the top-level search results retrieved from the external resource library based on the enhanced query. m The second candidate answer is formed by candidate answers. express Length, Represents a second-stage retrieval operation.

5. A psychological support question and answer generation system based on two-stage retrieval enhanced LLM, characterized in that: The system applies the psychological support question and answer generation method based on two-stage retrieval enhanced LLM as described in any one of claims 1 to 4, and the system comprises: Data acquisition module, used to: Given an external repository and the original problem; The first stage retrieval module is used to: Perform the first-stage search operation from the external resource library using the original question as the search query condition to obtain the first candidate answer; The first stage generates modules for: Use the original question and the first candidate answer as contextual prompts to construct prompts, and input them into the LLM to generate several generated answers; The second stage retrieval module is used to: The original question and the generated answer are combined as search query conditions to perform a second-stage search operation from an external resource library to obtain a second candidate answer; The second stage generates modules for: The original question and the second candidate answer were used as contextual prompts to construct prompts, which were input into LLM to generate the final answer and obtain the psychological support answer.

Citation Information

Patent Citations

  • Method for improving question answering accuracy based on context and LLM

    CN119271776A