Man-machine conversation method and device based on RAG, storage medium and program product

Through the analysis and rewriting of dialogue context information, and the construction of large model prompt words combined with relevant background knowledge, the problem information loss caused by unknown reference information in multiple rounds of RAG dialogue is solved, which improves the accuracy of retrieval and reply, and improves the application performance and user experience of multi-wheel human-computer interaction.

CN120104753APending Publication Date: 2025-06-06AISPEECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510238090.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In multiple rounds of RAG conversations, the reference information often occurs in user questions, resulting in the lack of problem information, which in turn affects the accuracy of the search and reply of the big model, resulting in poor hallucination reply and application performance.

Method used

By obtaining and analyzing the context information of the conversation, the questions entered by the user are rewrite and optimized, the rewrite and optimized questions are determined, and the big model prompt words are constructed based on relevant background knowledge to ensure the accuracy of RAG retrieval and the correctness of big model reply.

Benefits of technology

It effectively solves the problematic information loss caused by unknown reference information in multiple rounds of dialogue, improves the clarity and completeness of questions, ensures the accuracy of RAG retrieval and the intelligence of large model reply, and thus improves the application performance and user experience of multiple rounds of human-computer interaction scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104753A_ABST
    Figure CN120104753A_ABST
Patent Text Reader

Abstract

The invention provides an RAG-based man-machine conversation method and device, a storage medium and a program product, and the method comprises the steps: obtaining conversation context information associated with an input question, and completing the input question according to the conversation context information to determine a rewriting optimization question; performing RAG retrieval according to the rewriting optimization questions to determine at least one corresponding related background knowledge, and constructing large model cue words according to the related background knowledge and the rewriting optimization questions; and inputting the big model cue word into the man-machine conversation big language model, and determining a reply answer for the input question according to the output content of the man-machine conversation big language model. Therefore, through continuous tracking and analysis of the dialogue context information and rewriting of user questions, the retrieval accuracy, the generation intelligence and the dialogue consistency of the system are effectively improved, and then the application performance and the user experience in a multi-round human-computer interaction scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a human-computer dialogue method, device, storage medium and program product based on RAG. Background Art

[0002] With the rapid development of artificial intelligence technology, intelligent dialogue systems have been widely used in many fields, such as customer service, intelligent assistants, education and training, etc. At present, large language models have made remarkable achievements in the field of natural language processing. Some human-computer interaction products have greatly improved the intelligence level of human-computer interaction products by integrating large language models.

[0003] RAG (Retrieval-Augmented Generation) is a model that retrieves external knowledge content based on user questions and inputs it into the big model as prompt information, allowing the big model to read and understand it before replying. It effectively improves the performance of the big model and expands the application scope of the big model.

[0004] In the multi-round dialogue of the RAG mode, based on the chat characteristics of multi-round interactions, users may often use referential information, resulting in the lack of question information. At this time, directly performing RAG retrieval on user questions with missing information will not be able to recall effective background knowledge for prompt information, causing the large model to produce hallucinations and reply with wrong content, and also causing the RAG mode to have poor application performance in multi-round continuous human-computer interaction scenarios.

[0005] The industry has not yet proposed a better solution to the above problems. Summary of the invention

[0006] The present application provides a human-computer dialogue method, device, storage medium and program product based on RAG, which are used to at least solve the problem of large model output hallucination that often occurs in current RAG multi-round dialogues.

[0007] In a first aspect, an embodiment of the present application provides a human-computer dialogue method based on RAG, comprising: obtaining dialogue context information associated with an input question, and improving the input question according to the dialogue context information to determine a rewritten optimized question; performing a RAG search according to the rewritten optimized question to determine at least one corresponding relevant background knowledge, and constructing a large model prompt word according to each of the relevant background knowledge and the rewritten optimized question; inputting the large model prompt word into a large language model for human-computer dialogue, and determining a reply answer to the input question according to the output content of the large language model for human-computer dialogue.

[0008] In a second aspect, an embodiment of the present application provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the steps of the human-computer dialogue method of any embodiment of the present application.

[0009] In a third aspect, an embodiment of the present application provides a storage medium on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the human-computer dialogue method of any embodiment of the present application are implemented.

[0010] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the human-computer dialogue method of any embodiment of the present application.

[0011] The beneficial effects of the embodiments of the present application are: By acquiring and analyzing the conversation context information and rewriting and optimizing the user input questions, it can effectively solve the problem of missing question information caused by unclear reference information in user questions in multi-round conversation interaction scenarios, and improve the clarity and completeness of input questions, thereby ensuring that sufficient and accurate background knowledge can be obtained during RAG retrieval, and combining it with the optimized questions to build a more targeted and context-aware large model prompt word. Therefore, through the continuous tracking and analysis of conversation context information and rewriting user questions, the system's retrieval accuracy, generation intelligence and conversation consistency are effectively improved, thereby improving application performance and user experience in multi-round human-computer interaction scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0013] Figure 1 An operation flow chart of an example of a human-computer dialogue method based on RAG according to an embodiment of the present application is shown; Figure 2 An operation flow chart of an example of optimizing rewriting condition judgment according to an embodiment of the present application is shown; Figure 3 An operation flow chart showing an example of improving an input question according to conversation context information to determine a rewritten optimized question according to an embodiment of the present application; Figure 4 An operation flow chart showing another example of improving an input question according to conversation context information to determine a rewritten optimized question according to an embodiment of the present application; Figure 5 A flowchart showing another example of improving an input question according to conversation context information to determine a rewritten optimized question according to an embodiment of the present application; Figure 6 An operation flow chart of an example of starting multiple rounds of interaction in an optimized RAG scenario according to an embodiment of the present application is shown; Figure 7 It is a schematic structural diagram of an embodiment of an electronic device of the present application. DETAILED DESCRIPTION

[0014] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0015] Figure 1 An operation flow chart of an example of a RAG-based human-computer dialogue method according to an embodiment of the present application is shown.

[0016] like Figure 1 As shown, in step S110, the dialogue context information associated with the input question is obtained, and the input question is improved according to the dialogue context information to determine a rewritten optimized question.

[0017] It should be noted that in a multi-round conversation, each round of user questions may contain information such as reference and contextual relevance. Therefore, the user conversation context can be constructed through previous records (such as the user's previous round of questions, the system's answers, and other contextual information).

[0018] In some implementations, natural language processing (NLP) techniques, such as dependency parsing, coreference resolution, and entity recognition, can be used to extract core elements from contextual information. By comparing the current question with the previous conversation history, it is possible to determine missing key information or existing unclear references, infer missing parts of the user's intention based on the understanding of the previous conversation, and optimize the question sentence structure to ensure the clarity and specificity of the question.

[0019] For example, if a user asks "What is the impact on the workflow?", the system will determine the specific object that "it" refers to (such as "a newly launched function" or "a specific device") through the context, and convert the question into a more specific and clear rewritten optimized question, such as "What is the impact of the newly launched function on the workflow?". In this way, through the supplementation and optimization of context information, it ensures that the user's questions are clearer and more complete, and reduces incorrect answers due to missing information.

[0020] In step S120, a RAG search is performed based on the rewritten optimization question to determine at least one corresponding relevant background knowledge, and a large model prompt word is constructed based on each relevant background knowledge and the rewritten optimization question.

[0021] It should be understood that the rewritten optimized questions are supplemented and optimized with context information, and using them for RAG retrieval can be more targeted and can more accurately find background knowledge related to the question from the external knowledge base. In some examples, the RAG retrieval module can also combine semantic retrieval technology (such as semantic matching based on pre-trained models such as BERT and RoBERTa) to retrieve background knowledge related to the question, ensuring that the recalled content includes not only the knowledge itself, but also can be associated with the specific context of the question.

[0022] In some embodiments, the most relevant information is selected from multiple candidate background knowledge returned by RAG retrieval. Such background knowledge can come from multiple sources, such as the enterprise's internal knowledge base, external open APIs, Internet documents, etc. Then, the retrieved background knowledge is integrated and prioritized in combination with the context of the question (for example, whether the questioner has discussed similar issues before), and the knowledge that best matches the current conversation is selected. Finally, the prompt words of the large language model are constructed by combining the rewritten optimized question with the relevant background knowledge, which not only contains the optimized question content, but also embeds the relevant background information into it, ensuring the clarity and comprehensibility of the information when it is input into the large language model.

[0023] In some examples, a RAG search is performed on the background knowledge document based on the rewritten optimized questions to determine at least one corresponding relevant background knowledge, and the background knowledge document is constructed or updated based on the document uploaded by the user. Specifically, users are allowed to upload documents (such as PDF, Word, TXT and other formats), thereby supporting the human-computer dialogue function of document RAG, and a RAG search is performed on the uploaded custom document based on the rewritten optimized questions to obtain at least one relevant background knowledge fragment. As a result, the document upload and update mechanism ensures that the background knowledge base can reflect changes in user needs in a timely manner, especially when the document content or industry knowledge is updated, it can automatically synchronize and ensure that the latest knowledge is included in the retrieval and generation process, thereby improving the quality of answers.

[0024] In step S130, the large model prompt words are input into the large language model of human-computer dialogue, and the reply answer to the input question is determined according to the output content of the large language model of human-computer dialogue.

[0025] In some embodiments, the constructed prompt words (including the rewritten questions and related background knowledge) will be input into the large language model, so that the large language model can generate reasonable answers to user questions based on the information in the prompt words and the context. In one example, the output content of the large language model of human-computer dialogue can be directly used as a reply answer, thereby combining the RAG retrieval background knowledge of the optimized question with the powerful reasoning ability of the large model, so that the generated answer can more accurately answer the user's question and reduce errors caused by unclear understanding of the context. In another example, the output content of the large language model can be further optimized, for example, removing redundant information, correcting language expression, enhancing the professionalism of the reply, etc., to ensure that the quality of the reply reaches a high standard.

[0026] Figure 2 An operation flow chart of an example of optimizing rewriting condition judgment according to an embodiment of the present application is shown.

[0027] like Figure 2 As shown, in step S210, it is detected whether the input question involves context-related intention.

[0028] Here, the intent recognition of user input questions can be diverse, such as a text classification model based on BERT or RoBERTa, whose task is to determine whether the user's question depends on the context in the previous conversation. Through the judgment of the intent recognition model, it is possible to accurately distinguish between questions that require contextual association and independent questions, improving the system's ability to understand user needs in multiple rounds of conversations.

[0029] In step S220, if the input question is not related to the context, a RAG search is performed based on the input question.

[0030] Specifically, if it is detected that the input question has no obvious contextual association, for example, the input question is an independent question or an open-ended question, the original input question will be directly passed to the RAG retrieval module. For example, when the input question is "What is the weather like in Suzhou today?", it is an independent question and is generally unrelated to the previous conversation, so a RAG search can be performed based on the input question. In this way, by directly passing the input question to the RAG module, answers can be quickly obtained when the question has no complex contextual dependencies, reducing unnecessary context processing steps and improving response efficiency.

[0031] In step S230, if the input question is associated with the context, the dialog context information associated with the input question is obtained, and the input question is improved according to the dialog context information to determine a rewritten optimized question.

[0032] Specifically, if the system determines that the input question is related to the context (for example, the question contains pronouns or depends on previous information), it is necessary to first extract relevant contextual information from the conversation history, identify the key information on which the question depends from the conversation history, and "complete" and "optimize" the user's input question, including resolving pronouns, supplementing missing information, reconstructing unclear or ambiguous expressions, etc., to reduce misunderstandings caused by incomplete question information and ensure that subsequent answers are more in line with user expectations.

[0033] Figure 3 An operational flowchart of an example of improving an input question according to conversation context information to determine a rewritten optimized question according to an embodiment of the present application is shown.

[0034] like Figure 3 As shown, in step S310, it is detected whether the input question is a reference question.

[0035] In some implementations, it is identified whether the user's question contains referential content (for example, using pronouns such as "he", "she", "it", "that", "these", etc.), and the referential words in the question are marked and classified as referential questions. If the question does not use referential words (for example, directly asking "What is the definition of machine learning?"), it is judged as a non-referential question, and the subsequent referential processing steps are skipped to avoid unnecessary computing overhead and improve processing efficiency.

[0036] In step S320, when it is detected that the input question is a reference question, the reference content corresponding to the input question is determined according to the dialogue context information, and the input question is supplemented and improved according to the reference content to determine a rewritten optimized question.

[0037] In some embodiments, when it is detected that the question is a reference question, the antecedent of the reference word is searched through the context window (usually including the conversation content of the previous rounds before the question is asked), for example, using the context sliding window method to gradually check the previous conversation until the content matching the reference word is found. Then, the reference resolution algorithm is used to determine the specific content of the reference word. For example, if the user asks "What are his representative works?", the male characters or entities in the conversation history will be searched, and the specific object "he" refers to will be determined based on the context content.

[0038] Specifically, once the specific content of the pronoun is determined, the original question can be "supplemented" and "rewritten" according to the context. For example, if a user asks "What are his representative works?" and the antecedent mentioned in the context is "Zhang San", the system rewrites the original question to "What are Zhang San's representative works?" In this way, the antecedent information is obtained from the dialogue context, and a rewritten optimized question is generated to eliminate the problem of unclear reference in the original question, ensuring that the subsequent RAG retrieval or the answer generated by the large language model is more in line with the user's intention.

[0039] Figure 4 The operational flowchart of another example of improving an input question according to conversation context information to determine a rewritten optimized question according to an embodiment of the present application is shown.

[0040] like Figure 4 As shown, in step S410, it is detected whether the input question belongs to a multiple question.

[0041] It should be noted that multiple questions refer to users asking multiple questions in the same question, usually using words such as "and", "or", "in addition", "at the same time" to connect multiple sub-questions. Common multiple questions include two-part or multi-part questions, such as "Can you tell me about the weather today? And what about the weather tomorrow?".

[0042] Here, we can use techniques such as syntactic analysis, grammatical structure recognition or conjunction detection to determine whether the input question contains multiple questions. For example, we can analyze the structure of the question sentence and identify whether there are multiple verbs or typical conjunctions in parallel (such as "and", "or", "at the same time") to identify the structure of multiple questions.

[0043] In step S420, when it is detected that the input question is a multiple question, the input question is split into multiple question sentences to determine a rewritten optimized question.

[0044] Here, a multiple question is split into multiple independent sub-questions. For example, by identifying the conjunctions in the question, such as "and", "or", "at the same time", etc., the question is split into corresponding sub-questions. For example, in the questions "Do you like dogs or cats?" and "Which one do you think is smarter?", "or" and "which one" indicate the boundary between the two questions, and the system splits it into two independent sub-questions: "Do you like dogs or cats?" and "Which one is smarter?"

[0045] In this way, by processing the multiple sub-questions in parallel, performing independent RAG retrieval or large language model generation for each sub-question, and then integrating the answers to each sub-question into a complete response to provide to the user. In this way, complex multiple questions are split into multiple simple, independent questions, avoiding confusion and misunderstanding, making the question expression clearer, and reducing ambiguity caused by the complex structure of the original multiple questions.

[0046] Figure 5 The flowchart shows another example of improving an input question according to conversation context information to determine a rewritten optimized question according to an embodiment of the present application.

[0047] like Figure 5 As shown, in step S510, the input question and the conversation context information are respectively filled into the corresponding reserved question placeholder slots and context placeholder slots in the preset rewriting optimization prompt word template to obtain the rewriting optimization prompt word.

[0048] Regarding the description of the rewriting optimization prompt word template, it can contain multiple placeholders, which will be filled according to different needs and dialogue scenarios. The placeholder types mainly include question placeholders and context placeholders. Question placeholders are used to fill in the original user questions, and context placeholders are used to fill in dialogue context information related to the questions, such as the topics, events, and characters mentioned above. They can be distinguished by corresponding tags, such as "{{question}}" and "{{context}}".

[0049] In step S520, the rewriting optimization prompt words are input into the question rewriting language model to determine the corresponding rewriting optimization question.

[0050] Here, the content contained in the prompt word will be used by the question rewriting large language model to understand the user's intention and context. Combined with the previous conversation information, the question rewriting large language model generates an optimized question. Specifically, the question rewriting large language model can combine contextual information to understand the actual intention of the question, eliminate ambiguous or unclear expressions in the original question, make the question clearer and more precise, and adjust the current input question according to the conversation history to make it more consistent with the context of the current conversation, avoiding ignoring previous information when answering. Therefore, with the help of the powerful natural language understanding and text generation capabilities of the large language model, the system can automatically optimize the user's questions to make them clearer and more precise in semantics and context, and can avoid misunderstandings or wrong answers caused by vague or unclear expressions, thereby improving the interactive experience with the user.

[0051] In some examples of the embodiments of the present application, the rewriting optimization prompt word template includes at least one prompt word segment from the following: a context-related intention identification segment, a reference question optimization segment, or a multiple question optimization segment.

[0052] Therefore, the large language model performs corresponding operations on the input questions according to different prompt word fragments (such as context-related intent, reference question optimization, and multiple question optimization), understands the background of the input question, the entity referred to, and multiple questions, automatically removes ambiguity, clearly expresses the user's intention, and makes the question more accurate, thereby generating clear, accurate, and consistent questions in the conversation context. More details about the rewriting optimization prompt word template will be expanded in conjunction with other examples below.

[0053] It should be noted that the reason why the large language model can conduct multiple rounds of conversations is that the historical conversations between the user and the assistant are inserted into the input message group as prompt information and are role-identified.

[0054] The examples are as follows (Example 1: Input message group of a large model multi-round dialogue): "messages": [ { "role": "system", "content": "You are a geography expert and are only allowed to answer users' geography-related questions." }, { "role": "user", "content": "What is the land area of ​​China?" }, { "role": "assistant", "content": "China's total land area is approximately 9.6 million square kilometers." }, { "role": "user", "content": "What about Japan?" } ] In the above example, system is the system role, which is always the premise of inputting prompt words, user is the user role, and assistant is the assistant role of the big model. Therefore, when the message group is input into the big model, the big model will follow the premise of the system prompt words, obtain and understand the previous conversation history information between the user and the big model, and answer the user's latest question. The latest question in the example, "What about Japan?" actually contains multiple rounds of references. The big model will automatically reply to the user's question about the land area of ​​Japan based on the historical conversation information.

[0055] It should be noted that this similar technology can meet the needs of questions within the pre-trained knowledge of the big model itself. However, when the user asks multiple rounds of referential questions for the RAG scenario, this message structure will not be able to provide effective prompt information, causing the big model to hallucinate and reply with wrong content. The specific reason is that the RAG scenario is a mode in which external knowledge content is retrieved for user questions as prompt information and input into the big model for reading comprehension and reply. Directly performing RAG retrieval on user questions with missing information will not be able to recall effective background knowledge for prompt information.

[0056] In particular, RAG recall is often based on text vector models. The principle of vector retrieval is to find vectors similar to the target and calculate the score ranking. Therefore, when a question with referential information but no core semantics is directly searched for vectors, it will be difficult to retrieve some effective information. Therefore, in current human-computer dialogue systems, the multi-round interaction capabilities in RAG scenarios are usually limited, and multi-round interactions are not supported in document question-and-answer RAG scenarios.

[0057] Example 2: a conversation message group containing a reference question: "messages": [ { "role": "system", "content": "You are a document Q&A assistant and are only allowed to answer user questions based on the document content." }, { "role": "user", "content": "How many days of annual leave does the company have?" }, { "role": "assistant", "content": "According to the employee handbook, our employees are legally entitled to 10 days of annual leave." }, { "role": "user", "content": "How many days in advance should I apply for annual leave?" }, { "role": "assistant", "content": "According to the employee handbook, employees are generally required to submit a written application three days in advance for annual leave." }, { "role": "role", "content": "Do you need any materials?" } ] In Example 2 above, the first two questions asked by the user are about the number of days of annual leave and the method of applying for annual leave. When performing a RAG search in the employee handbook, relevant knowledge can be obtained to answer the user. However, for the third question, the user asks "What materials are needed?" When performing a RAG search directly in the employee handbook, it is possible to retrieve: materials required for employee onboarding, materials required for sick leave, etc. At this time, if this part of background knowledge is inserted into the prompt word information, it will cause the large model to have a deviation in understanding, and thus fail to accurately answer the question that the user really wants to ask, "What materials are needed to apply for annual leave?"

[0058] Taking into account the defects of the human-computer dialogue system of the above-mentioned large language model, the following concepts are proposed in the embodiments of the present application: First, the document question and answer RAG scenario needs to support multiple rounds of interaction, which is more in line with conversation habits; second, it can handle problems where the user lacks necessary semantics, such as incomplete information, reference disambiguation, etc., so as to correctly retrieve relevant background knowledge; third, questions that the user expresses normally can be left unprocessed.

[0059] In practice, for document question-answering RAG scenarios, a new round of "question optimization" can be added selectively. Using prompt word templates and placeholders, after filling in key information, the natural language processing capabilities of the large language model are used to rewrite and optimize user questions. Rewriting optimization supports reference to the conversation context to complete important information.

[0060] In some examples, the prompt word example template is as follows (Example 3: Question optimization prompt word template): # Role You are a Chinese language assistant who is good at understanding the meaning of the current user's question based on the context, and rewriting and optimizing the user's question for output.

[0061] ## Skill ### Skill 1: Disambiguate user questions - Based on the context, complete the content referred to in the user's question to make it a complete and easy-to-understand sentence.

[0062] ### Skill 2: User Problem Analysis - If the user's question contains multiple questions, please separate them into multiple question sentences.

[0063] ## limit - If the user question is not related to the context, there is no need to optimize it and the original question can be directly output.

[0064] - If the user question needs to be optimized, directly output the optimized user question. Do not answer the question or give some explanation or elaboration. Just return the conclusion directly.

[0065] # Example Example 1: Context: [Q: What is the weather like today? A: It is expected to be cloudy during the day, with a maximum temperature of 35℃ and light wind, and cloudy tonight, with a minimum temperature of 2℃ and light wind.] Current input: What about tomorrow? Output: What will the weather be like tomorrow? Example 2: Context: [Q: What is the weather like today? A: It is expected to be cloudy during the day, with a maximum temperature of 35℃ and light wind, and cloudy tonight, with a minimum temperature of 2℃ and light wind.] Current input: Recommend me some local food in Suzhou Output: Recommend me some local food in Suzhou # Now, let's start with the following context: {history_msg} # Current user questions: {last_query} In the above example 3, history_msg and last_query in parentheses are placeholders for the conversation context and user question, respectively.

[0066] Then based on the question in Example 2, the prompt word template can be filled in as follows: # Role You are a Chinese language assistant who is good at understanding the meaning of the current user's question based on the context, and rewriting and optimizing the user's question for output.

[0067] ## Skill ### Skill 1: Disambiguate user questions - Based on the context, complete the content referred to in the user's question to make it a complete and easy-to-understand sentence.

[0068] ### Skill 2: User Problem Analysis - If the user's question contains multiple questions, please separate them into multiple question sentences.

[0069] ## limit - If the user question is not related to the context, there is no need to optimize it and the original question can be directly output.

[0070] - If the user question needs to be optimized, directly output the optimized user question. Do not answer the question or give some explanation or elaboration. Just return the conclusion directly.

[0071] # Example Example 1: Context: [Q: What is the weather like today? A: It is expected to be cloudy during the day, with a maximum temperature of 35℃ and light wind, and cloudy tonight, with a minimum temperature of 2℃ and light wind.] Current input: What about tomorrow? Output: What will the weather be like tomorrow? Example 2: Context: [Q: What is the weather like today? A: It is expected to be cloudy during the day, with a maximum temperature of 35℃ and light wind, and cloudy tonight, with a minimum temperature of 2℃ and light wind.] Current input: Recommend me some local food in Suzhou Output: Recommend me some local food in Suzhou # Now, let's start with the following context: [Q: How many days of annual leave does the company have? A: According to the employee handbook, our employees are legally entitled to 10 days of annual leave. Q: How many days in advance do I need to apply for annual leave? A: According to the employee handbook, under normal circumstances, employees need to submit a written application 3 days in advance to apply for annual leave.] # Current user questions: Do you need any materials? At this point, the filled prompt words are input into the large model, and the user question with completed information can be output: What materials are needed to apply for annual leave? Then, the document RAG recall of "What materials are needed to apply for annual leave?" can be used to obtain relevant background knowledge as prompt information for the large model to reply. Assume that the background knowledge recalled by RAG is as follows: ## Background knowledge 1: Before taking a leave, you should make arrangements and coordination for matters related to your work and not affect your current work. If you violate this rule, the department head, supervisor or human resources department has the right to reject your leave application. If you take a leave without approval, the company will treat it as absenteeism and have the right not to pay your salary.

[0072] Annual leave is calculated based on the natural year. In principle, annual leave cannot be accumulated. However, unused annual leave in the previous year can be included in the first quarter of the following year, and will be automatically cleared if it is overdue.

[0073] If an employee has taken sick leave for 2 months cumulatively per year, he / she will no longer be entitled to annual leave benefits for that year.

[0074] Employees will continue to receive salary during their annual leave.

[0075] #WeddingLeave If an employee registers his / her marriage while employed, the procedures shall be based on local laws and regulations and the place of work where the contract is signed.

[0076] ## Background knowledge 2: #Leave Procedure Employees who apply for leave must obtain approval from the department head or supervisor. Those who take leave without approval will be treated as absenteeism.

[0077] Before taking leave, the person applying for leave should apply to the department head or supervisor in advance through the E-HR system and provide necessary supporting documents.

[0078] The department head or supervisor shall approve it and upload relevant supporting documents in the E-HR system (web or mobile) for filing with the human resources department.

[0079] I can only take leave after receiving approval from the department head and making arrangements for related work during the holiday.

[0080] In case of emergency (such as sudden illness, fire, theft, etc.) and there is no time to apply for leave in writing or by email, you should try to apply for leave by phone to the department head or supervisor in time, and complete the leave application procedures within 3 days and obtain the approval of the department head. Those who fail to apply for leave will still be considered absent from work.

[0081] Then the complete information finally spliced ​​into the message group will be as follows: "messages": [ { "role": "system", "content": "You are a document Q&A assistant and are only allowed to answer user questions based on the document content." }, { "role": "user", "content": "How many days of annual leave does the company have?" }, { "role": "assistant", "content": "According to the employee handbook, our employees are legally entitled to 10 days of annual leave." }, { "role": "user", "content": "How many days in advance should I apply for annual leave?" }, { "role": "assistant", "content": "According to the employee handbook, employees are generally required to submit a written application three years in advance for annual leave." }, { "role": "role", "content": "Now, the following document background is provided to you: ## Background knowledge 1: Before taking a leave, you should make arrangements and coordination for matters related to your work and not affect your current work. If you violate this rule, the department head, supervisor or human resources department has the right to reject your leave application. If you take a leave without approval, the company will treat it as absenteeism and have the right not to pay your salary.

[0082] Annual leave is calculated based on the natural year. In principle, annual leave cannot be accumulated. However, unused annual leave in the previous year can be included in the first quarter of the next year, and will be automatically cleared if it is overdue.

[0083] If an employee has taken sick leave for 2 months cumulatively in a year, he / she will no longer be entitled to annual leave benefits for that year.

[0084] Employees will continue to receive salary during their annual leave.

[0085] # Marriage Leave If an employee registers his / her marriage during employment, the marriage registration procedure shall be carried out according to local laws and regulations and the place of work where the contract is signed. ## Background knowledge 2: # Leave Procedure Employees who apply for leave must obtain approval from the department head or supervisor. Those who take leave without approval will be treated as absenteeism.

[0086] Before taking leave, the person applying for leave should apply to the department head or supervisor in advance through the E-HR system and provide necessary supporting documents.

[0087] The department head or supervisor must approve the application and upload relevant supporting documents to the E-HR system (web or mobile) for filing with the human resources department. After receiving the approval from the department head, the applicant must make arrangements for related work during the holiday before taking leave.

[0088] In case of emergency (such as sudden illness, fire, theft, etc.) and there is no time to apply for leave in writing or by email, you should try to apply for leave by phone to the department head or supervisor in time, and complete the leave application procedures within 3 days and obtain the approval of the department head. Those who fail to apply for leave will still be considered absent from work.

[0089] Please reply to the user's question based on the above content: What materials are needed to apply for annual leave? " } Finally, this message group is input into the big model to obtain the expected response effect of the user's original question "Do you need any materials?" The sample response effect is as follows: Employees who apply for leave must obtain approval from their department heads or supervisors. Before taking leave, the employee must apply to their department heads or supervisors in advance through the E-HR system and provide necessary supporting documents.

[0090] Figure 6 An operational flowchart of an example of starting multiple rounds of interaction in an optimized RAG scenario according to an embodiment of the present application is shown.

[0091] 1) User questions.

[0092] 2) Determine whether it is a document RAG scenario. If not, directly splice it into the message group according to the conventional multi-round dialogue method to execute the next process. If it is, proceed to step 3).

[0093] 3) Determine whether question optimization is enabled. If not, follow the normal process to directly perform RAG retrieval of background knowledge for the user question and then proceed to the next process. If enabled, proceed to step 4).

[0094] Here, question optimization is set as an optional configuration item, which can be set or adjusted according to needs. For example, when you need to experience the effect of multiple rounds of dialogue, you can choose to turn it on, and it supports flexible configuration.

[0095] 4) Optimize user questions based on context. Specifically, the big model is used to optimize the information of the current user question in combination with previous conversation records, and the optimized user question is output. When the big model determines that no optimization is needed, the original question is directly output.

[0096] Therefore, the reading comprehension ability of the big model is used to optimize and rewrite the original user question to obtain the complete user question, which improves the accuracy of RAG retrieval and affects the final response effect of the big model.

[0097] At the same time, because this step calls a large model, the response time may increase. In some business scenarios, other small models that have been trained specifically can be used to replace the large model in this step to optimize the problem and shorten the time of the entire link.

[0098] 5) Splice into message group. This step is to splice the relevant background knowledge retrieved from the user question RAG into the message group that is finally sent to the big model. If no relevant background knowledge is obtained, no splicing is required.

[0099] 6) Input the big model to return the answer. Finally, the big model responds to the user's question based on the system prompt words, conversation context, background knowledge and other prompt information in the message group.

[0100] Through the embodiments of the present application, user questions are associated with historical conversations, which can maintain the contextual consistency of the conversation. By optimizing the questions, clearer and more specific questions are provided, so that RAG retrieval can more efficiently obtain more accurate relevant background information from the background knowledge base, ensuring that the message group finally input to the big model contains the most relevant background knowledge, contextual information and optimized questions, avoiding interference from irrelevant or erroneous background information, increasing the accuracy of the big model's understanding and generation, and making the final answer more in line with the user's real needs.

[0101] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of actions combined, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application. In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0102] In some embodiments, an embodiment of the present application provides a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored. The execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any of the above-mentioned human-computer interaction methods of the present application.

[0103] In some embodiments, the embodiments of the present application also provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes any one of the above-mentioned human-computer interaction methods.

[0104] In some embodiments, an embodiment of the present application also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a human-computer dialogue method.

[0105] Figure 7 is a schematic diagram of the hardware structure of an electronic device for executing a human-computer dialogue method provided by another embodiment of the present application, such as Figure 7 As shown, the device includes: One or more processors 710 and memory 720, Figure 7 A processor 710 is taken as an example.

[0106] The device for executing the human-machine dialogue method may further include: an input device 730 and an output device 740 .

[0107] The processor 710, the memory 720, the input device 730 and the output device 740 may be connected via a bus or other means. Figure 7 The example of connecting through bus is taken in the following.

[0108] The memory 720 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the human-computer dialogue method in the embodiment of the present application. The processor 710 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory 720, that is, realizing the human-computer dialogue method in the above method embodiment.

[0109] The memory 720 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 720 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 720 may optionally include a memory remotely arranged relative to the processor 710, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0110] The input device 730 may receive input digital or character information and generate signals related to user settings and function control of the electronic device. The output device 740 may include a display device such as a display screen.

[0111] The one or more modules are stored in the memory 720, and when executed by the one or more processors 710, the human-computer dialogue method in any of the above method embodiments is executed.

[0112] The above-mentioned product can execute the method provided in the embodiment of the present application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of the present application.

[0113] The electronic devices of the embodiments of the present application exist in various forms, including but not limited to: (1) Mobile communication equipment: This type of equipment is characterized by having mobile communication functions and its main purpose is to provide voice and data communications. This type of terminal includes: smart phones, multimedia phones, functional phones, and low-end phones.

[0114] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have mobile Internet access features. These terminals include: PDA, MID and UMPC devices, etc.

[0115] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players, handheld game consoles, e-books, smart toys, and portable car navigation devices.

[0116] (4) Other onboard electronic devices with data interaction functions, such as on-board devices installed in vehicles.

[0117] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0118] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.

[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A human-computer dialogue method based on RAG, comprising: Acquire conversation context information associated with the input question, and improve the input question according to the conversation context information to determine a rewritten optimized question; Perform RAG search according to the rewritten optimization question to determine at least one corresponding relevant background knowledge, and construct a large model prompt word according to each of the relevant background knowledge and the rewritten optimization question; The large model prompt words are input into the large language model of human-computer dialogue, and a reply answer to the input question is determined according to the output content of the large language model of human-computer dialogue.

2. The method according to claim 1, wherein: The obtaining of the dialog context information associated with the input question, and improving the input question according to the dialog context information to determine a rewritten optimized question, includes: Detecting whether the input question involves contextual intent; If the input question is not related to the context, performing a RAG search based on the input question; and If the input question is associated with a context, dialog context information associated with the input question is obtained, and the input question is improved according to the dialog context information to determine a rewritten optimized question.

3. The method according to claim 1, wherein: The step of improving the input question according to the conversation context information to determine a rewritten optimized question includes: Detecting whether the input question is a referential question; In the case where it is detected that the input question is a reference question, the reference content corresponding to the input question is determined according to the conversation context information, and the input question is supplemented and improved according to the reference content to determine a rewritten optimized question.

4. The method according to claim 1, wherein: The step of improving the input question according to the conversation context information to determine a rewritten optimized question includes: Detecting whether the input question is a multiple question; When it is detected that the input question is a multiple question, the input question is split into multiple question sentences to determine a rewritten optimized question.

5. The method according to any one of claims 1 to 4, wherein: The step of improving the input question according to the conversation context information to determine a rewritten optimized question includes: Filling the input question and the conversation context information into the corresponding reserved question placeholder slot and context placeholder slot in the preset rewriting optimization prompt word template respectively to obtain the rewriting optimization prompt word; The rewriting optimization prompt words are input into the question rewriting language model to determine the corresponding rewriting optimization question.

6. The method according to claim 5, wherein: The rewriting optimization prompt word template includes at least one prompt word segment from the following: a context-related intention identification segment, a reference question optimization segment, or a multiple question optimization segment.

7. The method according to claim 1, wherein: The performing RAG search according to the rewritten optimized question to determine at least one corresponding relevant background knowledge comprises: A RAG search is performed on the background knowledge document according to the rewritten optimization question to determine at least one corresponding relevant background knowledge; the background knowledge document is constructed or updated based on the document uploaded by the user.

8. A storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.

9. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method described in any one of claims 1 to 7.

10. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Input system, device and method for artificial intelligence generation content of intelligent cabin and storage medium

    CN121075321A

  • Replay content confirmation method and device, storage medium and program product

    CN122240818A

  • Apparatus, method, and program for a generative AI model to reflect unlearned information.

    JP7847707B1