Human-computer interaction method, server, storage medium and program product
By generating query plans and response frameworks, including methods with multiple subqueries, the problem of monotonous and poorly organized response information in human-computer interaction systems is solved, enabling the generation of real-time, rich, and well-organized response information.
Patent Information
- Application Number
- CN202410496197.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-23
- Publication Date
- 2025-10-28
AI Technical Summary
In existing human-computer interaction systems, the generated response information is monotonous, poorly organized, and lacks real-time performance. Furthermore, the inconsistent quality of retrieval sources and interfaces across different demand scenarios leads to low-quality response information.
By generating query plans and response frameworks, including multiple subqueries, query results are searched and retrieved in real time. Response information is generated using auxiliary generation models and human-computer interaction models, ensuring the real-time nature, richness, and organization of the response information.
It achieves real-time and rich responses, improves response quality, and ensures that responses conform to the pre-planned framework and have good organization.
Smart Images

Figure CN120849570A_ABST
Abstract
Description
Technical Field
[0001] This application relates to computer technology, and more particularly to a human-computer interaction method, server, storage medium, and program product. Background Technology
[0002] With the development of artificial intelligence technology, natural language processing has been widely applied in various fields, such as medical consultation, online education, online shopping, and financial customer service. In human-computer interaction systems such as virtual digital humans, intelligent dialogue robots, digital customer service, and chatbots, human-computer interaction models can automatically generate responses to user queries. However, the generation of responses depends on the timeliness of the training set used in the training phase of the human-computer interaction model, resulting in poor real-time performance and the problem of knowledge illusion.
[0003] To improve the real-time nature of responses and address the knowledge illusion problem, augmented retrieval-based human-computer interaction solutions have emerged. These solutions first retrieve candidate knowledge based on the query information, and then generate responses using a human-computer interaction model. Since the responses are derived from real-time retrieved candidate knowledge, their real-time performance is relatively good. However, using different knowledge bases or search engines to search for real-time candidate knowledge to generate responses can lead to issues. Because different scenarios use different retrieval sources and interfaces, some search engines or knowledge bases are of poor quality, resulting in low-quality candidate knowledge and responses that are often simplistic and lack coherence. Summary of the Invention
[0004] This application provides a human-computer interaction method, server, storage medium, and program product to solve the problems of monotonous content and poor organization of generated response information in human-computer interaction systems.
[0005] Firstly, this application provides a human-computer interaction method, including:
[0006] Based on the input query information, a query plan and a response framework are generated. The query plan contains multiple subqueries for searching the information needed to respond to the query information. The queries are executed based on the multiple subqueries to obtain the query results. Based on the query results of the multiple subqueries and the response framework, response information for the query information is generated.
[0007] Secondly, a human-computer interaction method includes:
[0008] The system receives user questions sent by the receiving end-side device; generates a query plan and a response framework based on the user questions, wherein the query plan includes multiple sub-queries for searching for information required to answer the user questions; executes queries based on the multiple sub-queries to obtain query results for the multiple sub-queries; generates response information for the user questions based on the query results of the multiple sub-queries and the response framework; and returns the response information for the user questions to the receiving end-side device.
[0009] Thirdly, this application provides a human-computer interaction method, including:
[0010] Obtain the input query information; input the query information into a retrieval decision model, and determine whether a retrieval is needed through the retrieval decision model; based on the determination result, if a retrieval is needed, retrieve candidate knowledge matching the query information based on the query information, input the query information and the candidate knowledge into a human-computer interaction model, and generate response information for the query information through the human-computer interaction model based on the candidate knowledge; if a retrieval is not needed, input the query information into the human-computer interaction model, and generate response information for the query information through the human-computer interaction model.
[0011] Fourthly, this application provides a human-computer interaction method, comprising:
[0012] The system identifies the category of the input query information. If the query information is classified as a complex query, a query plan and a response framework are generated based on the input query information. The system executes queries based on multiple sub-queries included in the query plan, obtains the query results of the multiple sub-queries, and generates a response to the query information based on the query results of the multiple sub-queries and the response framework. If the query information is classified as a knowledge-based query, the query information is input into a human-computer interaction model, and a response to the query information is generated through the human-computer interaction model. If the query information is classified as a real-time fact-based query, candidate knowledge matching the query information is retrieved based on the query information. The query information and the candidate knowledge are input into the human-computer interaction model, and a response to the query information is generated through the human-computer interaction model based on the candidate knowledge.
[0013] Fifthly, this application provides a server, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the server to perform the methods provided in any of the foregoing aspects.
[0014] Sixthly, this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the method provided in any of the foregoing aspects.
[0015] In a seventh aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the methods provided in any of the foregoing aspects.
[0016] The human-computer interaction method, server, storage medium, and program product provided in this application generate a query plan and a response framework based on the input query information. The query plan includes multiple subqueries used to search for the information needed to respond to the query information. This allows for pre-planning of the query plan and response framework to be executed. Furthermore, by executing the query based on the multiple subqueries and obtaining the query results of the multiple subqueries, the rich information needed to respond to the query information can be searched in real time and fully. Furthermore, by generating response information based on the query results of the multiple subqueries and the response framework, the real-time and rich content of the generated response information can be guaranteed, and it conforms to the pre-planned response framework and has good organization. Thus, the response quality is comprehensively improved in terms of real-time performance, richness, and organization. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] Figure 1 This is a schematic diagram of an example system architecture to which this application applies;
[0019] Figure 2 A flowchart illustrating a human-computer interaction method provided in an exemplary embodiment of this application;
[0020] Figure 3 A flowchart illustrating a framework for generating a query plan and response, provided as an exemplary embodiment of this application;
[0021] Figure 4 A flowchart for generating response information based on query results and a response framework, provided as an exemplary embodiment of this application;
[0022] Figure 5 A general framework diagram of a human-computer interaction method provided for an exemplary embodiment of this application;
[0023] Figure 6 A flowchart illustrating the training process of an auxiliary generative model provided in an exemplary embodiment of this application;
[0024] Figure 7A flowchart illustrating a human-computer interaction method provided in another exemplary embodiment of this application;
[0025] Figure 8 An interaction flowchart of a human-computer interaction method provided as an exemplary embodiment of this application;
[0026] Figure 9 An interaction flowchart of a human-computer interaction method provided as another exemplary embodiment of this application;
[0027] Figure 10 This is a schematic diagram of the structure of a server provided in an embodiment of this application.
[0028] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0029] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0030] It should be noted that the user information (including but not limited to user device information, user attribute information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0031] First, let me explain the terms used in this application:
[0032] Prompt: This refers to the prompt words input to the large model. Conditions can be set to allow the model to output corresponding results according to instructions or requirements. For example, "Please write a poem, with the format XX, style YY, and theme ZZ." XX, YY, and ZZ can be filled in according to the actual application requirements.
[0033] Fine-tuning: In deep learning, it involves continuously training and updating the parameters (weights) of a pre-trained model in its deep network to fit a model that achieves the expected results.
[0034] Search: refers to finding specific information within a large amount of data, typically using keywords or algorithms / tools such as search engines. The goal of a search is to find information relevant to a given query and to rank and rate it.
[0035] Retrieval refers to obtaining specific information from stored data. It involves indexing the data to quickly retrieve relevant information, such as document retrieval in a library system.
[0036] Data distillation: Extracting specific data from a large model is called "data distillation". The extracted data usually has some noise.
[0037] Model fine-tuning, also known as supervised fine-tuning (SFT), involves fine-tuning a pre-trained model for a specific task. During fine-tuning, the model's parameters and structure are adjusted based on the task's characteristics to improve its performance. Various techniques can be used, such as data augmentation, regularization, and optimization algorithms. The advantage of SFT is its ability to quickly fine-tune for different tasks without retraining the entire model. Furthermore, large pre-trained models can be trained using massive amounts of text data, resulting in better performance. Commonly used SFT methods include P-Tuning v2, Low-Rank Adaptation (LoRA), Quantized LoRA (QLoRA), Freeze, and full-parameter fine-tuning.
[0038] Organization: In terms of form, the answer is presented in bullet points, which can be highlighted with numbers and punctuation marks; in terms of content, each bullet point explains one detail.
[0039] Logicality: It emphasizes whether the order of the key points in the answer is appropriate or whether the answer is organized with logical thinking. For example, answering in order from primary to secondary points reflects logicality; doing things from easy to difficult also reflects logicality.
[0040] Hierarchical structure: It reflects a hierarchical order, such as a progressive structure, where each level is a level; for example, if the answer follows a progressive logic, from shallow to deep or from deep to shallow, it reflects the hierarchical structure of the answer.
[0041] Comprehensiveness: When answering questions, the candidate is able to provide a comprehensive response from all aspects. The responses to each aspect are also organized, logical, and hierarchical.
[0042] Self-QA: A method that uses a large language model, prompts, and context to generate questions (Q) and answers (A). The characteristic of this method is that it asks and answers its own questions, and the questioning and answering can be done simultaneously or in separate steps.
[0043] Visual question answering task: Based on the input image and the question, determine the answer to the question from the visual information of the input image.
[0044] Image description task: Generate descriptive text for the input image.
[0045] Visual entailment task: Predict the semantic relevance between input images and text, i.e., entailment, neutrality, or contradiction.
[0046] The task of expression and comprehension involves locating the image region in the input image that corresponds to the input text.
[0047] Image generation task: Generate an image based on the input descriptive text.
[0048] Text-based sentiment classification task: Predict the sentiment classification information of input text.
[0049] Text summarization task: Generate a summary of the input text.
[0050] Multimodal tasks refer to downstream tasks that involve multiple modalities of data, such as images and text, in their input and output. Examples include visual question answering, image description, visual entailment, representation and understanding, and image generation.
[0051] Multimodal pre-trained models refer to pre-trained models whose input and output data involve multiple modalities such as images and text. After fine-tuning and training, they can be applied to multimodal task processing.
[0052] Pre-trained language model: A pre-trained model obtained by pre-training a large language model (LLM).
[0053] Large models refer to deep learning models with a massive number of parameters, typically containing hundreds of millions, tens of billions, or even trillions of parameters. Large models are also known as foundation models (FM), which are pre-trained on large-scale unlabeled corpora to produce pre-trained models with hundreds of millions of parameters. These models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and Multi-modal Pre-training Models.
[0054] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios of large models include digital assistants, intelligent chatbots, digital customer service, chatbots, search, online education, office software, e-commerce, and intelligent design.
[0055] This application provides a human-computer interaction method to address the problem that in search-enhanced human-computer interaction, different search sources and interfaces are used in different demand scenarios. Some search engines or knowledge bases are poorly designed, resulting in low-quality candidate knowledge and thus the generated response information is relatively simple and lacks organization.
[0056] This application provides a human-computer interaction method that generates a query plan and a response framework based on input query information. The query plan includes multiple subqueries used to search for the information needed to respond to the query information. This allows for pre-planning of the query plan and response framework to be executed. Furthermore, by executing the query based on the multiple subqueries and obtaining the query results of the multiple subqueries, the rich information needed to respond to the query information can be searched in real time and fully. Furthermore, by generating response information based on the query results of the multiple subqueries and the response framework, the real-time and rich content of the generated response information can be guaranteed, and it conforms to the pre-planned response framework and has good organization. Thus, the response quality is comprehensively improved in terms of real-time performance, richness, and organization.
[0057] Figure 1 This is a schematic diagram of an example system architecture to which this application applies. Figure 1 As shown, the system architecture includes a server and endpoint devices. The server and endpoint devices have a communication link, enabling communication between them.
[0058] The server is a computing device deployed in the cloud or locally, such as a cloud cluster. The server stores a human-computer interaction model and an auxiliary generation model. The server is responsible for using the auxiliary generation model to automatically plan and obtain a query plan and response framework based on given query information. The query plan contains multiple subqueries used to search for the information needed to respond to the query. The server executes the queries based on these subqueries to obtain their results. Furthermore, using the human-computer interaction model, the server generates a response to the query based on the results of the subqueries and the response framework.
[0059] The auxiliary generative model is used to generate query plans and response frameworks corresponding to query information. It can be obtained by fine-tuning a pre-trained model. The pre-trained model can be various pre-trained language models, such as pre-trained Large Language Models (LLM) or BERT (Bidirectional Encoder Representations from Transformers) models. In some example scenarios, the auxiliary generative model can be obtained by fine-tuning a smaller-scale pre-trained model, which can improve the efficiency of human-computer interaction.
[0060] The human-computer interaction model is a model used to generate response information for a query. In this embodiment, the human-computer interaction model uses multiple subqueries and query results in the query plan as candidate knowledge, and generates response information for the query based on the query results of multiple subqueries and the response framework. The human-computer interaction model can use various pre-trained language models, such as pre-trained large language models (LLM), multimodal pre-trained language models, etc., or any existing large model used for human-computer interaction models. This embodiment does not make any specific limitations.
[0061] Edge devices can be electronic devices that run downstream human-computer interaction systems. Specifically, they can be hardware devices with network communication, computing, and information display functions, including but not limited to smartphones, tablets, desktop computers, local servers, and cloud servers. For example, edge devices can be electronic devices that run various human-computer interaction systems such as virtual digital humans, intelligent chatbots, digital customer service, and chatbots.
[0062] When human-computer interaction is required, the user submits query information through a client device, which then sends the query information to the server. The server receives the query information from the client device, and based on the input query information, invokes an auxiliary generation model to generate a query plan and a response framework. The query plan contains multiple subqueries used to search for the information needed to respond to the query. The server executes the queries based on the multiple subqueries to obtain the query results. Further, the server invokes the human-computer interaction model to generate a response to the query based on the query results and the response framework.
[0063] Furthermore, the server returns the generated response information to the endpoint device. The endpoint device can then output the response information to the user.
[0064] The method in this embodiment can be applied to various human-computer interaction scenarios such as intelligent customer service, intelligent chatbots, and digital customer service in fields such as e-commerce, online education, finance, agriculture, and intelligent transportation. This embodiment does not impose specific limitations on these scenarios.
[0065] It should be noted that in another human-computer interaction system architecture, there are end-device devices, a server for the human-computer interaction system (referred to as the first server), a second server running the human-computer interaction model, and a third server running an auxiliary generation model. The third server provides the API of the auxiliary generation model to the first server of the human-computer interaction system. The first server calls the auxiliary generation model through the API to plan and generate a query plan and response framework based on the query information. The second server provides the application programming interface (API) of the human-computer interaction model to the first server of the human-computer interaction system. The first server calls the human-computer interaction model through the API to generate response information for the query information. The second and third servers can be the same or different cloud servers, and the second server and the first server are different servers.
[0066] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0067] Figure 2 This is a flowchart illustrating a human-computer interaction method provided in an exemplary embodiment of this application. The execution entity in this embodiment is the server in the aforementioned system architecture. Figure 2 As shown, the specific steps of this method are as follows:
[0068] Step S201: Based on the input query information, generate a query plan and a response framework. The query plan contains multiple subqueries used to search for the information needed to respond to the query information.
[0069] Among them, query information refers to the input information (query) when users use the human-computer interaction system. It can be questions, instructions, etc., submitted by users, which describe the user's interaction needs.
[0070] For example, in the context of intelligent customer service in e-commerce, information retrieval could involve users asking questions about products, such as how to use a product, its target audience, or how to store it. In the context of intelligent dialogue in finance, information retrieval could involve users providing instructions or asking questions about a specific financial service, such as how to activate SMS notifications or the specific details of a project.
[0071] After receiving the input query information, the server plans a query plan based on the query information. This query plan contains multiple subqueries. By executing the query based on multiple subqueries, the server can obtain the rich information needed to answer the query information. The query results of multiple subqueries are sufficient to comprehensively and accurately answer the query information.
[0072] The server can also plan a response framework based on the query information. This framework provides an overall structure for the generated response information, which can be understood as an outline of the response. The response framework includes logical and hierarchical key points. Responses that conform to this framework can provide a relatively complete and organized answer to the query information, meeting the user's needs.
[0073] In one optional implementation of this step, the query information is populated into a pre-configured second prompt template to generate second prompt information. This second prompt information is used to prompt the auxiliary generation model to generate a query plan and response framework based on the query information. Further, the second prompt information is input into the auxiliary generation model, which then generates the query plan and response framework based on the second prompt information.
[0074] The second prompt template is used to generate second prompt information based on the input query information, which is then used by the prompt-assisted generation model to generate a query plan and response framework. The second prompt template contains the locations where query information needs to be filled in; by filling the corresponding locations with the input query information, the corresponding second prompt information can be obtained. The second prompt information refers to a type of prompt information generated based on the second prompt template; different query information will produce different second prompt information.
[0075] For example, an example of a second prompt template is as follows:
[0076] "When answering a question, there is usually a complete answer framework, which allows for a more complete and organized response. A query plan (including multiple subqueries) is also needed to retrieve relevant information (ensuring the answer is complete and error-free). The current question is: \"{query}\". Based on the question above, please provide an answer framework and the corresponding complete query plan. The query plan should contain no more than 5 subqueries. Output format: Answer framework:\n\nQuery plan:". Here, {query} refers to the location where the query information needs to be filled in.
[0077] In addition, the second suggestion template can define the content of the query plan, such as the maximum number of subqueries that can be included in the query plan, and each subquery includes the query objective and query conditions. The query objective refers to the type of target data to be found, and the query conditions include the keywords that the subqueries must use. Furthermore, the second suggestion template can provide at least one example of a query plan and response framework for a query, allowing the larger model to better understand the format and content requirements of the query plan and response framework based on the example. The specific content of the second suggestion template can be configured and adjusted according to the needs of the actual application scenario; no specific limitations are set here.
[0078] For example, here is an example of a response framework for the query "How can high school students make money?":
[0079] Response framework:
[0080] 1. Introduction: Briefly introduce the necessity and feasibility of high school students making money.
[0081] 2. Explore ways to earn money in accordance with relevant regulations: Introduce some legal and compliant ways to earn money that are suitable for high school students, such as part-time jobs, online surveys, and writing submissions.
[0082] 3. Provide specific advice and resources: Offer high school students practical advice and resources to help them find suitable money-making opportunities, such as paying attention to job websites and joining clubs or organizations.
[0083] 4. Warn of the dangers of certain methods: Emphasize the harm and consequences of certain money-making methods and warn high school students not to try them easily.
[0084] 5. Conclusion: In summary, this paper encourages high school students to earn pocket money through legal and compliant means, while reminding them to pay attention to their studies and health while earning money.
[0085] Based on this response framework, a structured and comprehensive answer can be provided for the query "How can high school students make money?".
[0086] For example, an example of a query plan for the information "How can high school students make money?" is as follows:
[0087] Query plan:
[0088] 1. Labor laws and regulations concerning minors: age restrictions, working hours, and job type restrictions.
[0089] 2. The tutoring market: subject demand, tutoring fees, and tutoring platforms.
[0090] 3. Online job platform: user reviews, payment methods, job requirements
[0091] 4. Handicraft Market Analysis: Popular Products, Price Range, and Marketing Channels
[0092] 5. Career guidance resources: school career counseling, online career planning services
[0093] 6. Entrepreneurship Resources and Guidance: Youth Entrepreneurship Forums, Blogs, Online Courses
[0094] The query plan contains 6 subqueries. Each subquery provides the query target (such as "labor laws for minors", "home tutoring market", "online job platform", "handicraft market analysis", "career guidance resources", "entrepreneurship resources and guidance"), as well as query conditions (including keywords corresponding to each query target, such as "age restriction, working hours, job restriction", "tutoring fee, home tutoring platform"...).
[0095] In another alternative implementation of this step, a fifth prompt template for generating a query plan and a sixth prompt template for generating a response framework can be configured separately.
[0096] The fifth suggestion template must at least include the locations where query information needs to be filled in. The fifth suggestion template can also define the content of the query plan, examples of query plans, etc. For example, the maximum number of subqueries that a query plan can contain, and each subquery including the query objective, query conditions, etc.
[0097] Fill the corresponding positions in the fifth prompt template with the input query information to obtain the corresponding fifth prompt information. The fifth prompt information is used to prompt the auxiliary generation model to generate a query plan based on the query information. Input the fifth prompt information into the auxiliary generation model, and the auxiliary generation model will generate a query plan based on the fifth prompt information.
[0098] For example, a sample of the fifth prompt template is as follows: "When answering the question, a query plan (including multiple subqueries) is required to retrieve relevant information (ensuring the completeness and accuracy of the answer). The current question is: \"{query}\". Please provide a complete query plan based on the question above, with no more than 5 subqueries in the query plan. Output format: Query plan: ". Here, {query} refers to the location where the query information needs to be filled in.
[0099] The sixth prompt template must at least include the locations where query information needs to be filled in. The sixth prompt template can also define the format, content, and examples of the response framework. Filling the input query information into the corresponding locations in the sixth prompt template will generate the corresponding sixth prompt information. The sixth prompt information is used to prompt the auxiliary generation model to generate a response framework based on the query information. Inputting the sixth prompt information into the auxiliary generation model allows the model to generate a response framework based on the sixth prompt information.
[0100] For example, an example of the fifth prompt template is as follows: "When answering a question, there is usually a complete answer framework, which allows for a more complete and organized response. The current question is: \"{query}\". Please provide an answer framework based on the above question. Output format: Answer framework:". Here, {query} refers to the location where query information needs to be filled in.
[0101] Step S202: Execute the query based on multiple subqueries to obtain the query results of multiple subqueries.
[0102] After obtaining the query plan, this step involves executing queries for each subquery based on the multiple subqueries contained in the query plan, and performing searches and / or retrievals based on each subquery to obtain the query results for each subquery.
[0103] In one optional implementation of this step, for multiple subqueries in the query plan, the query results of each subquery can be obtained sequentially based on the search of each subquery in the order of the subqueries in the query plan.
[0104] In another optional implementation of this step, for multiple subqueries in the query plan, the subqueries without dependencies and the subqueries with dependencies are determined according to the dependencies between the multiple subqueries in the query plan; the subqueries without dependencies are executed in parallel, and the subqueries with dependencies are executed in the order of dependencies to obtain the query results of each subquery.
[0105] In this context, a subquery with a dependency relationship refers to a subquery that requires the query result of another subquery. For example, if subquery A depends on subquery B, then the query is first executed based on subquery B to obtain the query result of subquery B, and then the query is executed based on the query result of subquery B to obtain the query result of subquery A.
[0106] For example, subquery A is "the director's schedule on the 21st of this month", subquery B is "the name of our director", and subquery C is "the month of which year". It is clear that subquery A depends on subqueries B and C. Only after subqueries B and C are executed first to obtain "the name of the director" and "the month of which year", can the query result of subquery A be executed.
[0107] It should be noted that different search interfaces / tools can be invoked to retrieve the query results for different subqueries. Specifically, the search interface / tool to be invoked for the subquery can be determined by using a pre-trained model, or by determining the search interface / tool to be invoked based on pre-configured search rules. This embodiment does not impose any specific limitations here.
[0108] Step S203: Generate response information for the query information based on the query results and response framework of multiple subqueries.
[0109] After obtaining the query results of multiple subqueries and the generated response framework, response information is generated based on the query results of multiple subqueries and the response framework. This results in response information that conforms to the response framework, making the generated response information more organized. Furthermore, generating response information based on the query results of multiple subqueries can improve the real-time nature and richness of the response information.
[0110] In one optional implementation of this step, the query information, response framework, multiple subqueries, and query results are input into the human-computer interaction model. The human-computer interaction model then generates response information based on the response framework according to the query information, response framework, multiple subqueries, and query results.
[0111] Specifically, the query information, response framework, multiple subqueries, and query results are filled into the third prompt template to generate the third prompt information. The third prompt information is used to prompt the human-computer interaction model to generate response information based on the response framework according to the query information, response framework, multiple subqueries, and query results. The third prompt information is then input into the human-computer interaction model, which generates response information based on the response framework according to the third prompt information.
[0112] The third prompt template is used to generate prompts for the human-computer interaction model. Based on given query information, the results of multiple subqueries, and a response framework, it generates third prompt information based on the response framework. The third prompt template includes locations for filling in query information, multiple subqueries and their results, and the response framework. By filling in the corresponding locations in the third prompt template with the query information, response framework, multiple subqueries, and their results, the third prompt information is obtained. The third prompt information refers to a type of prompt information generated based on the third prompt template; filling different information into the third prompt template will produce different third prompt information.
[0113] For example, an example of a third-party prompt template is as follows:
[0114] "The question you need to answer is: \"{query}\", and you already know the answer framework as follows: {answer_plan}\n\nThe query results you have obtained are as follows:\n\n{searchlog}\n\nPlease output the answer information in a structured and organized manner based on the answer framework and query results."
[0115] Here, {query} refers to the location where query information needs to be populated, {answer_plan} refers to the location where the answer framework needs to be populated, and {searchlog} refers to the location where multiple subqueries and query results need to be populated.
[0116] In another optional implementation of this step, multiple subqueries and query results are divided into multiple groups. Based on each group of subqueries and query results, as well as the response framework, the response framework, the group of subqueries and query results are input into the human-computer interaction model. The human-computer interaction model generates response fragments based on the response framework, the group of subqueries and query results. Furthermore, the multiple response fragments generated are integrated based on the response framework to obtain the response information for the query information.
[0117] Specifically, multiple subqueries and their results are divided into multiple groups, with each group containing at least one subquery and its result. For each group's subqueries and results, these are populated into a third-prompt template to generate the corresponding third-prompt information. This third-prompt information is then input into a human-computer interaction model, which generates a corresponding response based on the third-prompt information. This response serves as a fragment of the final response, allowing us to obtain response fragments for each group. Furthermore, based on the response framework, the response fragments from each group are concatenated and duplicate content is removed to obtain the response information for the query.
[0118] Optionally, the corresponding response segments of each group can be sorted and concatenated in order according to the time sequence of the content in each group, and duplicate content can be removed to obtain the response information for the query information.
[0119] Furthermore, after generating the response information for the query, the server can output the response information. For example, the server can return the response information to the endpoint device that provided the query information. The endpoint device then outputs the response information to the user.
[0120] In addition, the server can output response information to a designated device according to a pre-configured output method; or, it can output response information through a front-end visual interactive interface. There are no specific restrictions on the output method of the response information here.
[0121] In an optional embodiment, the generated response information may further include an information source marker, which indicates which query result of each subquery the response content originates from. After (or before, or simultaneously with) the output of the response information, all information sources of the response information (including the query results of each subquery) and corresponding markers (used to mark and distinguish different query results) may also be output.
[0122] For example, given the query "How can high school students make money?" and the previously given response framework, an example of a response that conforms to the framework is as follows:
[0123] 1. Introduction: For high school students, earning money can not only develop their social skills but also improve their economic independence. However, while earning money, they also need to abide by relevant regulations and ensure that it does not affect their studies and health.
[0124] 2. Exploring methods that comply with relevant regulations: High school students can earn pocket money in several ways. First, they can try tutoring, using their strengths in certain subjects to help younger students improve their academic performance (Source 1). Second, they can sell learning materials or goods at school, such as books from bookstores or mobile phones from merchants (Source 1). In addition, they can also earn income by writing articles or participating in surveys (Source 2).
[0125] 3. Provide specific advice and resources: To find suitable earning opportunities, high school students can pay attention to recruitment websites, social media platforms, etc., to find suitable part-time jobs or internships (Source 2). They can also join school clubs and organizations to improve their abilities and skills through participation in various activities (Source not mentioned).
[0126] 4. Warning about the dangers of certain methods: Some ways of making money not only violate relevant regulations but may also lead to serious consequences, such as being punished according to regulations or suffering damage to one's own interests (Source 3). Therefore, high school students must choose methods that comply with relevant regulations when making money to avoid falling into traps that violate these regulations (Source 4).
[0127] 5. Conclusion: In general, high school students can earn pocket money in various ways that comply with relevant regulations, such as tutoring, selling goods, and writing articles. While earning money, they also need to pay attention to complying with relevant regulations, ensuring their own safety, and balancing their studies and health.
[0128] The method in this embodiment generates a query plan and a response framework based on the input query information. The query plan includes multiple subqueries used to search for the information needed to respond to the query information. A detailed query plan to be executed can be planned in advance, and a logical and hierarchical response framework can be planned. Furthermore, by executing the query based on the multiple subqueries and obtaining the query results of the multiple subqueries, the information needed to respond to the query information can be searched in real time and fully, providing rich and real-time reference information for generating response information. Furthermore, by generating response information based on the query results of the multiple subqueries and the response framework, the real-time and rich content of the generated response information can be guaranteed, and it conforms to the pre-planned response framework, with good organization. Thus, the response quality is comprehensively improved in terms of real-time performance, richness, and organization.
[0129] Figure 3 This is a flowchart illustrating the generation of a query plan and response framework based on input query information, provided as an exemplary embodiment of this application. Building upon the foregoing embodiments, in an optional embodiment, a search and / or retrieval can be performed first based on the query information to obtain matching knowledge. Then, a query plan and response framework can be generated based on the query information and the matching knowledge. This can improve the quality of the generated query plan and response framework, making the generated query plan more comprehensive and richer, and the generated response framework more reasonable.
[0130] like Figure 3 As shown, the generation of a query plan and response framework based on the input query information in the aforementioned step S201 can be specifically achieved through the following steps S2011-S2012:
[0131] Step S2011: Search and / or retrieve based on the input query information to obtain matching knowledge for the query information.
[0132] In this embodiment, based on the pre-built internal knowledge base of the human-computer interaction system, enhanced retrieval is performed on the input query information in the knowledge base to obtain enhanced retrieval results. The specific retrieval method can be any enhanced retrieval method in a human-computer interaction scheme based on enhanced retrieval; no specific limitation is made here. Based on a third-party search engine, a search is performed on the input query information to obtain search results. The matching knowledge for the query information in this step includes enhanced retrieval results and / or search results, which can be configured and adjusted according to the needs of the actual application scenario; no specific limitation is made here.
[0133] Step S2012: Generate a query plan and response framework based on the query information and the matching knowledge of the query information.
[0134] After obtaining the matching knowledge of the query information, this step inputs the query information and the matching knowledge of the query information into the auxiliary generation model. The auxiliary generation model generates a query plan and a response framework based on the query information and the matching knowledge of the query information, so as to improve the quality of the generated query plan and response framework, making the generated query plan more comprehensive and richer, and the generated response framework more reasonable.
[0135] In one optional implementation of this step, the query information and its matching knowledge are populated into the pre-configured first prompt template to generate first prompt information. This first prompt information is used to prompt the auxiliary generation model to generate a query plan and response framework based on the query information and its matching knowledge. Further, the first prompt information is input into the auxiliary generation model, which then generates the query plan and response framework based on the first prompt information.
[0136] The first prompt template is used by the prompt-assisted generation model to generate the first prompt information for the query plan and response framework based on the input query information and the matching knowledge of the query information. The first prompt template contains the locations where query information needs to be filled and the locations where matching knowledge needs to be filled. By filling the query information and its matching knowledge into the corresponding locations in the first prompt template, the corresponding first prompt information can be obtained. The first prompt information refers to a type of prompt information generated based on the first prompt template; filling different information into the first prompt template will produce different first prompt information.
[0137] For example, an example of a first prompt template is as follows:
[0138] "When answering a question, there is usually a complete answer framework, which allows for a more complete and organized response. A query plan (including multiple subqueries) is also needed to retrieve relevant information (ensuring the completeness and accuracy of the answer). The current question is: \"{query}\", and the matching knowledge found is: \"{content}\". Based on the question and the matching knowledge found, please provide an answer framework and the corresponding complete query plan. The query plan should contain no more than 5 subqueries. Output format: Answer framework:\n\nQuery plan:". Where {query} refers to the location where the query information needs to be filled, and {content} refers to the location where the matching knowledge needs to be filled.
[0139] In addition, the first suggestion template can define the content of the query plan, such as the maximum number of subqueries that can be included in the query plan, and each subquery includes the query objective and query conditions. The query objective refers to the type of target data to be found, and the query conditions include the keywords used by the subqueries. Furthermore, the first suggestion template can provide at least one example of a query plan and response framework for a query, allowing the larger model to better understand the format and content requirements of the query plan and response framework based on the example. The specific content of the first suggestion template can be configured and adjusted according to the needs of the actual application scenario; no specific limitations are set here.
[0140] In another optional implementation of this step, a seventh prompt template for generating a query plan based on query information and matching knowledge, and an eighth prompt template for generating a response framework based on query information and matching knowledge can also be configured.
[0141] The seventh suggestion template should at least include the locations where query information needs to be filled and the locations where matching knowledge needs to be filled. The seventh suggestion template can also define the content of the query plan, examples of query plans, etc. For example, the maximum number of subqueries that a query plan can contain, and each subquery including the query objective and query conditions.
[0142] Fill the corresponding positions in the seventh prompt template with the input query information and the matching knowledge of the retrieved query information to obtain the corresponding seventh prompt information. The seventh prompt information is used to prompt the auxiliary generation model to generate a query plan based on the query information and matching knowledge. Input the seventh prompt information into the auxiliary generation model, and the auxiliary generation model will generate a query plan based on the seventh prompt information.
[0143] For example, an example of the seventh prompt template is as follows: "When answering a question, a query plan (including multiple subqueries) is needed to retrieve relevant information (ensuring the completeness and accuracy of the answer). The current question is: \"{query}\", and the matching knowledge found is: \"{content}\"". Based on the above question and the found matching knowledge, please provide a complete query plan, with no more than 5 subqueries in the query plan. Output format: Query plan: ". Here, {query} refers to the location where the query information needs to be filled, and {content} refers to the location where the matching knowledge needs to be filled.
[0144] The eighth prompt template includes at least the locations where query information needs to be filled and the locations where matching knowledge needs to be filled. The eighth prompt template can also define the format, content, and examples of the response framework. By filling the corresponding locations in the eighth prompt template with the input query information and the retrieved matching knowledge, the corresponding eighth prompt information can be obtained. The eighth prompt information is used to prompt the auxiliary generation model to generate a response framework based on the query information and the retrieved matching knowledge. The eighth prompt information is input into the auxiliary generation model, which then generates the response framework based on the eighth prompt information.
[0145] For example, an example of the eighth prompt template is as follows: "When answering a question, there is usually a complete answer framework, which allows for a more complete and organized response. The current question is: "{query}", and the matching knowledge found is: "{content}". Please provide an answer framework based on the above question and the found matching knowledge. Output format: Answer framework:". Here, {query} refers to the location where the query information needs to be filled in, and {content} refers to the location where the matching knowledge for the query information needs to be filled in.
[0146] The method in this embodiment first searches and / or retrieves query information to obtain matching knowledge, and then generates a query plan and a response framework based on the query information and the matching knowledge. This can improve the quality of the generated query plan and response framework, making the generated query plan more comprehensive and richer, and the generated response framework more reasonable.
[0147] Figure 4 This is a flowchart illustrating the generation of response information based on query results and a response framework, provided as an exemplary embodiment of this application. Based on any of the foregoing embodiments, in this embodiment, as... Figure 4 As shown, in the aforementioned step S203, the response information for the query information is generated based on the query results and response framework of multiple subqueries. Specifically, this can be achieved using the following steps S2031-S2033:
[0148] Step S2031: Based on the maximum input length of the human-computer interaction model, divide the multiple subqueries and query results into multiple groups, with each group containing at least one subquery and its corresponding query result.
[0149] Considering that the human-computer interaction model has a limit on the input length, in this embodiment, multiple subqueries and query results are divided into multiple groups according to the maximum input length of the human-computer interaction model, and each group contains at least one subquery and its corresponding query result.
[0150] In this step, when dividing multiple subqueries and query results into multiple groups, the maximum input length of the human-computer interaction model and the length of each subquery and query result are considered. It is required that the subqueries and corresponding query results of any group be filled into the corresponding positions in the third prompt template along with the query information and the response framework. The length of the resulting third prompt information meets the input length requirements of the human-computer interaction model, that is, the length of the third prompt information is less than or equal to the maximum input length of the human-computer interaction model.
[0151] Step S2032: Input the response framework, subqueries and query results of each group into the human-computer interaction model. The human-computer interaction model generates the corresponding response information for each group based on the response framework, subqueries and query results of each group.
[0152] In this step, for each group of subqueries and query results, the query information, response framework, and subqueries and query results within that group are populated into the third prompt template to generate the corresponding third prompt information. This third prompt information is used to prompt the human-computer interaction model to generate response information based on the response framework, the query information, and the subqueries and query results within that group. Further, the third prompt information is input into the human-computer interaction model, which then generates the corresponding response information. An example of the third prompt template can be found in the aforementioned embodiment and will not be repeated here.
[0153] In this embodiment, each group of corresponding response information serves as a response fragment of the final response information, thus obtaining the response fragments for each group. Furthermore, the response fragments for each group are integrated to obtain the response information for the query.
[0154] In one optional embodiment, based on the response framework, the response fragments corresponding to each group are concatenated and duplicate content is removed to obtain the response information for the query. Optionally, the response fragments corresponding to each group can also be sorted according to the time order of the content in each group, concatenated in sequence, and duplicate content is removed to obtain the response information for the query.
[0155] Step S2033: Input the query information and the corresponding response information of each group into the human-computer interaction model, and integrate and sort the corresponding response information of each group through the human-computer interaction model to generate the response information of the query information.
[0156] In this embodiment, a human-computer interaction model is used to integrate and sort the corresponding response information of each group to generate response information for the query information.
[0157] Specifically, based on the pre-configured fourth prompt template, the query information and the corresponding response information for each group are filled into the fourth prompt template to generate the fourth prompt information. The fourth prompt information is used to prompt the human-computer interaction model. Based on the integration and sorting of the corresponding response information for each group, the response information for the query information is generated. Further, the fourth prompt information is input into the human-computer interaction model, which then generates the response information for the query information based on the fourth prompt information.
[0158] The fourth prompt template is used to generate a fourth prompt message for the human-computer interaction model, which generates a complete response message based on the query information and multiple response fragments. The fourth prompt template includes the positions where the query information needs to be filled and the positions where multiple response fragments need to be filled. By using the corresponding response information as response fragments and filling the query information and multiple response fragments into the corresponding positions in the fourth prompt template, the fourth prompt message can be obtained. The fourth prompt message refers to a type of prompt message generated based on the fourth prompt template; filling different information into the fourth prompt template will produce different fourth prompt messages.
[0159] Additionally, the fourth prompt template can also include rules for integrating multiple response fragments. For example, it can summarize and merge multiple response fragments for the same event / event, and then output them in chronological order. The fourth prompt template can also include length limits for generated response information, examples of integrating multiple response fragments, etc. The specific content of the fourth prompt template can be configured and adjusted according to actual application needs, and no specific limitations are made here.
[0160] For example, an example of the fourth prompt template is as follows:
[0161] "You need to answer the question: \"{query}\". Your summary history is as follows: \n\n{abslog}\n\nPlease find the information that answers \"{query}\" based on the above information, and output it in a structured and organized manner in ascending chronological order. Please summarize and merge the same event at the same time, and output multiple events separately. The total word count should not exceed 1800."
[0162] Here, {query} refers to the location where query information needs to be filled, and {abslog} refers to the location where multiple response fragments need to be filled.
[0163] The method in this embodiment divides multiple subqueries and query results into multiple groups according to the maximum input length of the human-computer interaction model. Each group contains at least one subquery and its corresponding query result. The response framework and the subqueries and query results of each group are input into the human-computer interaction model. The human-computer interaction model generates response information corresponding to each group based on the response framework and the subqueries and query results of each group. The query information and the response information corresponding to each group are input into the human-computer interaction model. The human-computer interaction model integrates and sorts the response information corresponding to each group to generate response information for the query information. This method can solve the problem of input length limitation of the human-computer interaction model. No matter how long the query results of multiple subqueries are, high-quality response information can be generated based on the query results of multiple subqueries and the response framework.
[0164] Figure 5 A general framework diagram of a human-computer interaction method provided for an exemplary embodiment of this application is shown below. Figure 5 As shown, the human-computer interaction method provided in this application consists of three stages. In the first stage, a query plan and a response framework are generated based on an auxiliary generation model. In the first stage, the auxiliary generation model can generate a query plan and a response framework based on the query information. Based on the query plan, the content needed to answer the query information can be searched, obtaining comprehensive, rich, and real-time information related to the query information. The response framework can provide an overall framework for the response information to be generated, which is organized and hierarchical.
[0165] The second phase involves decomposing and executing the query plan. Specifically, the query plan is broken down into multiple subqueries. Based on the dependencies between these subqueries, subqueries without dependencies are executed in parallel, while those with dependencies are executed according to those dependencies. For each subquery, the retrieval / search module executes the query (performing a search and / or retrieval) to obtain the query results for each subquery.
[0166] In the third stage, response information is generated based on the query results and response framework. Each subquery is aligned with the query results, and a human-computer interaction model is used to generate response information based on each subquery, its corresponding query results, response framework, and query information.
[0167] The framework of this embodiment generates a query plan and a response framework based on the input query information to pre-plan the query plan and response framework to be executed. Furthermore, it executes queries based on multiple sub-queries included in the query plan to obtain the query results of multiple sub-queries, which can search for rich information required to respond to the query information in real time and fully. Furthermore, it generates response information based on multiple sub-queries, query results, and the response framework, which can ensure the real-time nature and richness of the content of the generated response information, and conform to the pre-planned response framework, with good organization, thereby comprehensively improving the response quality in terms of real-time nature, richness, and organization.
[0168] Figure 6 A flowchart illustrating the training process of an auxiliary generative model provided for an exemplary embodiment of this application. (See attached diagram.) Figure 6 As shown, the auxiliary generation model used to generate query plans and response frameworks in the aforementioned embodiments can be trained through the following steps:
[0169] Step S601: Obtain a lightweight pre-trained model.
[0170] In this embodiment, a lightweight (i.e., smaller-scale) pre-trained model is obtained as the base model for fine-tuning the training to obtain the auxiliary generative model. The auxiliary generative model is then obtained by fine-tuning the base model using training data.
[0171] Among them, the pre-trained model can be any type of pre-trained language model, such as the pre-trained Large Language Model (LLM), BERT (Bidirectional Encoder Representations from Transformers) model, etc.
[0172] In this embodiment, selecting or constructing a small-scale pre-trained model as the base model and fine-tuning it to obtain an auxiliary generation model can improve the efficiency of the auxiliary generation model in generating query plans and response frameworks, thereby improving the efficiency of human-computer interaction.
[0173] For example, a smaller pre-trained model can be obtained by compressing a pre-trained large language model (LLM), which can then be used as the base model for the auxiliary generative model. Model compression of the LLM can be achieved through methods such as model pruning, knowledge distillation, and parameter quantization; specific limitations are not specified here.
[0174] Step S602: Construct training data based on domain knowledge. The training data includes query samples, matching knowledge corresponding to the query samples, query plans, and response frameworks.
[0175] To train an auxiliary generative model capable of generating query plans and response frameworks based on query information and matching knowledge, a training set is constructed, comprising query samples and corresponding matching knowledge, query plans, and response frameworks, for fine-tuning. The query plans and response frameworks serve as supervision data.
[0176] Specifically, the process involves acquiring domain knowledge of the target domain to which the human-computer interaction model belongs, extracting multiple query information from the domain knowledge as query samples, performing searches and / or retrievals based on the query samples to obtain matching knowledge for the query samples, filling the query samples and related details into the first prompt template to obtain the first prompt information corresponding to the query samples, and inputting the first prompt information corresponding to the query samples into a pre-trained large language model to generate a query plan and response framework corresponding to the query samples through the pre-trained large language model.
[0177] Each obtained query sample, the matching knowledge of the query sample, the query plan and response framework corresponding to the query sample constitute a training data, and the collection of obtained training data is used as the training set.
[0178] In acquiring domain knowledge for the target domain of the human-computer interaction model, a large amount of unlabeled and structured data can be obtained based on the target domain involved in the applied human-computer interaction system. This data includes, but is not limited to, documents, web page content, PDF files, and data from knowledge bases. By preprocessing this data, the corresponding text content can be obtained as domain knowledge for the target domain. The preprocessing process includes, but is not limited to, converting non-text formats such as images and tables into text, and removing useless information such as hyperlinks and symbols without semantic information. The specific rules for preprocessing can be configured and adjusted according to actual application needs; no specific limitations are imposed here.
[0179] When extracting multiple query information as query samples from domain knowledge, a question generation model can be used to generate relevant questions based on the acquired domain knowledge, which can then be used as query information. For example, questions can be generated based on the entire content of a document, or the content of multiple documents, or the content of randomly selected paragraphs from a document; or questions can be generated based on at least one piece of knowledge from a knowledge base, etc., with each question being generated as a query sample.
[0180] The question generation model used to generate questions can be any model that can generate questions based on given text, such as a Self-QA-based model or other models built based on Transformer that have the ability to generate questions based on given text. This embodiment does not make any specific limitations here.
[0181] Furthermore, when acquiring matching knowledge of the query sample, a third-party search engine can be used to search for matching knowledge of the query information, or matching knowledge with high similarity to the query information can be retrieved from the internal knowledge base of the human-computer interaction system. The specific retrieval method can adopt any augmented retrieval method in the human-computer interaction scheme based on augmented retrieval, and no specific limitation is made here.
[0182] Furthermore, after obtaining query samples and matching knowledge, the powerful generation capabilities of existing pre-trained large language models with large-scale parameters are leveraged to generate query plans and response frameworks based on query samples and matching knowledge.
[0183] Specifically, based on a pre-configured first prompt template, the query sample and its matching knowledge are populated into the first prompt template to generate first prompt information. This first prompt information is used to prompt the pre-trained large language model to generate a query plan and response framework based on the query sample and its matching knowledge. Further, the first prompt information is input into the pre-trained large language model, which then generates the query plan and response framework based on the first prompt information, thus obtaining the query plan and response framework corresponding to the query sample.
[0184] The first prompt template includes the locations where query information needs to be filled and the locations where matching knowledge needs to be filled. By filling the query sample into the locations where query information needs to be filled in the first prompt template, and filling the matching knowledge of the query sample into the locations where matching knowledge needs to be filled in the first prompt template, the corresponding first prompt information can be obtained. Examples and related explanations of the first prompt template can be found in the aforementioned embodiments, and will not be repeated here.
[0185] Step S603: Fine-tune the pre-trained model using training data to obtain an auxiliary generative model for generating query plans and response frameworks.
[0186] After obtaining training data including query samples, corresponding matching knowledge, query plans, and response frameworks, the pre-trained model is fine-tuned using the training data based on a pre-configured fine-tuning strategy, thus obtaining an auxiliary generative model for generating query plans and response frameworks.
[0187] Specifically, based on a pre-configured first prompt template, the query sample and its matching knowledge are populated into the first prompt template to generate first prompt information. This first prompt information is used to instruct the pre-trained model to generate a query plan and response framework based on the query sample and its matching knowledge. Further, the first prompt information is input into the pre-trained model, which then generates the query plan and response framework based on the first prompt information.
[0188] Furthermore, based on the query plan and response framework generated by the pre-trained model, and the query plan and response framework corresponding to the query samples in the training data, the cross-entropy loss is calculated, and the parameters of the pre-trained model are optimized through backpropagation.
[0189] Optionally, a cross-line loss can be calculated as the first loss based on the difference between the query plan generated by the pre-trained model and the query plan corresponding to the query sample in the training data; a cross-entropy loss can be calculated as the second loss based on the difference between the response framework generated by the pre-trained model and the response framework corresponding to the query sample in the training data; the first and second losses are weighted and summed to obtain the comprehensive loss, and the parameters of the pre-trained model are optimized through backpropagation based on the comprehensive loss. The weight coefficients of the first and second losses can be configured and adjusted according to actual application needs and empirical values, and are not specifically limited here.
[0190] Optionally, the concatenation results of the query plan and response framework generated by the pre-trained model, as well as the concatenation results of the query plan and response framework corresponding to the query samples in the training data, are obtained. Based on the difference between these two concatenation results, the cross-entropy loss is calculated, and the parameters of the pre-trained model are optimized through backpropagation based on the cross-entropy loss.
[0191] In addition, the fine-tuning strategies used for fine-tuning the pre-trained model, including the optimization algorithm, loss function, and test evaluation metrics used, as well as how to adjust hyperparameters such as learning rate and batch size, can be designed, configured, and adjusted by relevant technical personnel according to the needs of the actual application scenario. This embodiment does not impose specific limitations here.
[0192] The method in this embodiment acquires domain knowledge, extracts multiple query information as query samples from the domain knowledge, searches based on the query samples to obtain matching knowledge of the query samples, and leverages the powerful capabilities of a pre-trained large language model to generate query plans and response frameworks corresponding to the query samples. Training data for fine-tuning is constructed, allowing for supervised training without manual annotation. Furthermore, a lightweight pre-trained model is acquired, and fine-tuned based on the constructed training data to obtain an auxiliary generation model for generating query plans and response frameworks. This auxiliary generation model possesses the ability to generate query plans and response frameworks based on query information and matching knowledge, thereby improving the quality of the generated query plans and response frameworks.
[0193] In another alternative embodiment, another training set may be constructed, which includes query samples and query plans and response frameworks for the query samples, but does not contain matching knowledge of the query samples.
[0194] When constructing the training set, the methods for acquiring domain knowledge and query samples are described in steps S601-S602 above. Further, according to the configured second prompt template, the query samples are filled into the second prompt template to generate second prompt information. The second prompt information is used to prompt the large-scale pre-trained model to generate a query plan and response framework based on the query samples. The second prompt information is input into the large-scale pre-trained model, which then generates the query plan and response framework based on the second prompt information. Each obtained query sample, along with its corresponding query plan and response framework, constitutes a training data set, and the collection of obtained training data serves as the training set.
[0195] Furthermore, based on a pre-configured fine-tuning strategy, the pre-trained model is fine-tuned using the training set, enabling the auxiliary generative model to generate query plans and response frameworks based on query information (without knowing the matching knowledge of the query information).
[0196] Specifically, based on a pre-configured second prompt template, the query sample is populated into the second prompt template to generate second prompt information. This second prompt information is used to prompt the pre-trained model to generate a query plan and response framework based on the query sample. Further, the second prompt information is input into the pre-trained model, which then generates the query plan and response framework based on the second prompt information.
[0197] Furthermore, based on the query plan and response framework generated by the pre-trained model, and the query plan and response framework corresponding to the query samples in the training data, the cross-entropy loss is calculated, and the parameters of the pre-trained model are optimized through backpropagation. The loss calculation method and optimization method are detailed in step S603 above and will not be repeated here.
[0198] The method in this embodiment constructs a training set including query samples, query plans and response frameworks corresponding to the query samples, and fine-tunes the pre-trained model based on the training set to obtain an auxiliary generation model for generating query plans and response frameworks. This enables the auxiliary generation model to generate query plans and response frameworks based on query information, thereby improving the generation quality of query plans and response frameworks.
[0199] Figure 7 A flowchart illustrating a human-computer interaction method provided as another exemplary embodiment of this application. Based on any of the foregoing embodiments, this embodiment can combine the human-computer interaction scheme based on query plans and response frameworks provided in the foregoing embodiments with a scheme that directly generates response information for query information through a human-computer interaction model, and a human-computer interaction scheme based on enhanced retrieval. For example... Figure 7 As shown, the specific steps of this method are as follows:
[0200] Step S701: Identify the category of the input query information.
[0201] Among them, query information refers to the input information (query) when users use the human-computer interaction system. It can be questions, instructions, etc., submitted by users, which describe the user's interaction needs.
[0202] For example, in the context of intelligent customer service in e-commerce, information retrieval could involve users asking questions about products, such as how to use a product, its target audience, or how to store it. In the context of intelligent dialogue in finance, information retrieval could involve users providing instructions or asking questions about a specific financial service, such as how to activate SMS notifications or the specific details of a project.
[0203] After receiving the input query information, the server first uses a classification model to identify the category of the query information. The categories of query information include, but are not limited to, complex queries, knowledge queries, and real-time fact queries.
[0204] Among these, knowledge-based queries refer to common-sense queries with low timeliness requirements. For example, "Are whales mammals?" Real-time fact-based queries refer to queries closely related to time and with high timeliness requirements. For example, "What will the weather be like tomorrow? Will it rain?" Complex queries refer to queries with high timeliness requirements and complex reasoning logic. For example, "How can high school students make money?"
[0205] In this embodiment, the classification model is a pre-trained model, which can be obtained by fine-tuning various pre-trained language models, text classification models, etc. The classification model can classify and identify the input query information and determine the category corresponding to the query information.
[0206] Based on the classification model, the category of the obtained query information is identified. If the category of the query information is a complex query, steps S702-S704 are executed. Using a pre-planned query plan and response framework, multiple subqueries are searched based on the query plan. Based on the query results and response framework of the multiple subqueries, response information for the query information is generated.
[0207] Based on the classification model, the category of the obtained query information is identified. If the category of the query information is knowledge, step S705 is executed to directly generate the response information of the query information using the human-computer interaction model.
[0208] Based on the category of the query information identified by the classification model, if the category of the query information is real-time fact, steps S706-S707 are executed, and a human-computer interaction scheme based on augmented search is adopted. First, the corresponding candidate knowledge is searched based on the query information. The query information and candidate knowledge are input into the human-computer interaction model, and the human-computer interaction model generates the response information of the query information based on the candidate knowledge.
[0209] Step S702: If the type of query information is a complex query, generate a query plan and response framework based on the input query information.
[0210] The query plan contains multiple subqueries used to search for the information needed to answer the query.
[0211] In this embodiment, when the category of the query information is determined to be a complex query, the server generates a query plan and a response framework based on the input query information. The query plan contains multiple sub-queries used to search for the information needed to respond to the query information. For specific implementation principles and technical effects, please refer to the relevant content of step S201 in the aforementioned embodiment, which will not be repeated here.
[0212] Step S703: Execute the query based on the multiple subqueries contained in the query plan to obtain the query results of the multiple subqueries.
[0213] The specific implementation principle and technical effect of this step can be found in the relevant content of step S202 in the aforementioned embodiment, and will not be repeated here.
[0214] Step S704: Generate response information for the query information based on the query results and response framework of multiple subqueries.
[0215] The specific implementation principle and technical effect of this step can be found in the relevant content of step S203 in the aforementioned embodiment, and will not be repeated here.
[0216] Step S705: If the category of the query information is knowledge, input the query information into the human-computer interaction model, and generate the response information directly through the human-computer interaction model.
[0217] In this embodiment, when the category of the query information is determined to be knowledge-based, the query information is input into the human-computer interaction model, and the response information for the query information is directly generated through the human-computer interaction model.
[0218] Step S706: If the category of the query information is real-time fact, retrieve candidate knowledge that matches the query information based on the query information.
[0219] In this embodiment, when the category of the query information is determined to be real-time factual information, an augmented retrieval-based human-computer interaction scheme is used to retrieve candidate knowledge matching the query information. For details, please refer to the augmented retrieval processing flow in any augmented retrieval-based human-computer interaction scheme; no specific limitations are made here.
[0220] In this step, when performing enhanced retrieval to obtain candidate knowledge corresponding to the query, a third-party search engine can be used to search for matching knowledge of the query information, or matching knowledge with high similarity to the query information can be retrieved from the internal knowledge base of the human-computer interaction system. The specific retrieval method can be any enhanced retrieval method in the human-computer interaction scheme based on enhanced retrieval, and no specific limitation is made here.
[0221] Step S707: Input the query information and candidate knowledge into the human-computer interaction model, and generate the response information of the query information based on the candidate knowledge through the human-computer interaction model.
[0222] After obtaining candidate knowledge corresponding to the query information through enhanced retrieval, the query information and candidate knowledge are input into the human-computer interaction model. The human-computer interaction model then generates a response to the query information based on the candidate knowledge. For details, please refer to the processing flow of any retrieval-enhanced human-computer interaction scheme; specific limitations are not provided here.
[0223] This embodiment combines the human-computer interaction scheme based on query plan and response framework provided in the previous embodiments with the scheme of directly generating response information for query information through a human-computer interaction model, and a human-computer interaction scheme based on enhanced retrieval. When the query information is a complex query type, it pre-plans the query plan and response framework, executes multiple sub-queries based on the query plan, and generates response information based on the query results and response framework. This ensures the real-time nature, coherence, and richness of the generated response information, improving response quality. When the query information is a knowledge type, directly generating response information using a human-computer interaction model improves response efficiency. When the query information is a real-time fact type, an enhanced search-based human-computer interaction scheme is used. First, corresponding candidate knowledge is searched based on the query information. The query information and candidate knowledge are input into the human-computer interaction model, and the model generates response information based on the candidate knowledge, ensuring both real-time response and improved response efficiency.
[0224] Figure 8 This is a flowchart illustrating the human-computer interaction method provided in an exemplary embodiment of this application. Figure 8 As shown, the interaction flow of the human-computer interaction method is as follows:
[0225] Step S801: The end device sends a user question to the server.
[0226] User questions refer to the questions (queries) raised by users when using the human-computer interaction system.
[0227] For example, in the context of intelligent customer service in the e-commerce field, information retrieval can be a user's inquiry about product-related issues, such as asking how to use a product, the target audience, storage methods, etc.
[0228] In intelligent dialogue scenarios within the financial sector, information queries can include user inquiries about specific financial services, such as how to activate SMS notifications or the detailed contents of a particular project.
[0229] Step S802: The server receives user questions from the receiving device.
[0230] Step S803: The server generates a query plan and a response framework based on the user's question. The query plan contains multiple subqueries used to search for the information needed to answer the user's question.
[0231] The specific implementation principle and technical effect of this step can be found in the relevant content of step S202 above, and will not be repeated here.
[0232] Step S804: The server executes a query based on multiple subqueries to obtain the query results of the multiple subqueries.
[0233] The specific implementation principle and technical effect of this step can be found in the relevant content of step S203 above, and will not be repeated here.
[0234] Step S805: The server generates answer information for the user's question based on the query results of multiple subqueries and the answer framework.
[0235] The specific implementation principle and technical effect of this step can be found in the relevant content of step S204 above, and will not be repeated here.
[0236] Step S806: The server returns the answer information for the user's question to the end device.
[0237] Step S807: The end-side device outputs a response message.
[0238] When the method of this embodiment is applied to a downstream human-computer interaction system, a query plan and a response framework are generated based on the user questions sent by the terminal device. The query plan includes multiple sub-queries for searching the information needed to answer the user questions. This allows for the pre-planning of the query plan and response framework to be executed. Furthermore, by executing the queries based on the multiple sub-queries and obtaining the query results of the multiple sub-queries, the rich information needed to answer the user questions can be searched in real time and fully. Furthermore, by generating response information for the user questions based on the query results of the multiple sub-queries and the response framework, the real-time and rich content of the generated response information can be guaranteed, and it conforms to the pre-planned response framework and has good organization. Thus, the response quality is comprehensively improved in terms of real-time performance, richness, and organization.
[0239] Figure 9 This is a flowchart illustrating the human-computer interaction method provided in an exemplary embodiment of this application. Figure 9 As shown, the interaction flow of this human-computer interaction method is as follows:
[0240] Step S901: Obtain the input query information.
[0241] Among them, query information refers to the input information (query) when users use the human-computer interaction system. It can be questions, instructions, etc., submitted by users, which describe the user's interaction needs.
[0242] For example, in the context of intelligent customer service in e-commerce, information retrieval could involve users asking questions about products, such as how to use a product, its target audience, or how to store it. In the context of intelligent dialogue in finance, information retrieval could involve users providing instructions or asking questions about a specific financial service, such as how to activate SMS notifications or the specific details of a project.
[0243] When human-computer interaction is required, the user submits query information through a client device, which then sends the query information to the server. The server receives the query information sent by the client device.
[0244] Step S902: Input the query information into the retrieval decision model, and determine whether a retrieval is needed through the retrieval decision model.
[0245] After receiving the query information, the server will input the query information into the retrieval decision model, which will then determine whether a retrieval is necessary to respond to the current query information.
[0246] For relatively simple knowledge-based queries, or queries with low timeliness requirements, augmented retrieval is generally unnecessary; the human-computer interaction model can directly generate high-quality responses. However, for more complex queries, or queries with high real-time requirements, responses generated directly by the human-computer interaction model may not adequately answer the user's question, or the dataset used to train the model may not meet timeliness requirements, resulting in poor response quality. In these cases, augmented retrieval is necessary. The human-computer interaction model then uses the candidate knowledge retrieved through augmented retrieval to generate higher-quality responses.
[0247] The retrieval decision model can be obtained by fine-tuning a pre-trained model using a pre-built dataset, and it can make decisions on whether to perform enhanced retrieval based on the input query information. The pre-trained model can be various pre-trained language models, such as pre-trained Large Language Models (LLM), BERT (Bidirectional Encoder Representations from Transformers) models, etc.
[0248] Based on the judgment results of the retrieval decision model, if a retrieval is required, steps S903-S904 are executed to first enhance the retrieval to obtain candidate knowledge, and then generate response information based on the candidate knowledge through a human-computer interaction model. If a retrieval is not required, step S905 is executed to directly generate response information through the human-computer interaction model.
[0249] Step S903: Based on the judgment result, if retrieval is required, retrieve candidate knowledge that matches the query information based on the query information.
[0250] Based on the judgment results of the retrieval decision model, when a retrieval is required, enhanced retrieval is performed based on the query information to obtain candidate knowledge that matches the query information. The specific retrieval method can adopt any enhanced retrieval method in the human-computer interaction scheme based on enhanced retrieval, and no specific limitation is made here.
[0251] When performing enhanced retrieval, third-party search engines can be used to search for relevant information of the query information, or relevant information with high similarity to the query information can be retrieved from the internal knowledge base of the human-computer interaction system. The specific retrieval method can be any enhanced retrieval method in the human-computer interaction scheme based on enhanced retrieval, and no specific limitation is made here.
[0252] Step S904: Input the query information and candidate knowledge into the human-computer interaction model, and generate the response information of the query information based on the candidate knowledge through the human-computer interaction model.
[0253] Step S905: If no retrieval is required, input the query information into the human-computer interaction model, and generate a response to the query information through the human-computer interaction model.
[0254] In this embodiment, the solution uses a retrieval decision model to determine whether enhanced retrieval is needed to generate query information. Based on the determination result, if no retrieval is needed, the query information is input into a human-computer interaction model, which generates a response to the query information, thereby improving the timeliness of the response. If retrieval is needed, candidate knowledge matching the query information is retrieved based on the query information. The query information and candidate knowledge are then input into the human-computer interaction model, which generates a response to the query information based on the candidate knowledge, thereby improving the timeliness of the response.
[0255] Figure 10 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Figure 10 As shown, the server includes a memory 1001 and a processor 1002. The memory 1001 stores computer-executable instructions and can be configured to store various other data to support operations on the server. The processor 1002 is communicatively connected to the memory 1001 and executes the computer-executable instructions stored in the memory 1001 to implement the technical solutions provided in any of the above method embodiments. Their specific functions and the technical effects they achieve are similar and will not be repeated here.
[0256] Optional, such as Figure 10 As shown, the server also includes other components such as a firewall 1003, a load balancer 1004, a communication component 1005, and a power supply component 1006. Figure 10 The diagram only shows some components and does not mean that the server only includes... Figure 10 The components shown. Figure 10 This example uses a cloud server deployed in the cloud as an example, but the server can also be deployed locally. This embodiment does not make any specific limitations here.
[0257] This application also provides a computer-readable storage medium storing computer-executable instructions. When a processor executes the computer-executable instructions, it implements the method of any of the foregoing embodiments. The specific functions and technical effects to be achieved are not described here.
[0258] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method of any of the foregoing embodiments. The computer program is stored in a readable storage medium, and at least one processor of the server can read the computer program from the readable storage medium. The execution of the computer program by the at least one processor causes the server to perform the technical solution provided in any of the above method embodiments. The specific functions and the technical effects that can be achieved are not described here.
[0259] This application provides a chip, including a processing module and a communication interface. The processing module is capable of executing the technical solution of the server in the aforementioned method embodiments. Optionally, the chip further includes a storage module (e.g., a memory), which stores instructions. The processing module executes the instructions stored in the storage module, and the execution of the instructions stored in the storage module causes the processing module to execute the technical solution provided in any of the aforementioned method embodiments.
[0260] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.
[0261] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules from at least one processor.
[0262] The memory may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk or optical disc, etc.
[0263] The aforementioned storage device can be object storage service (OSS).
[0264] The aforementioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0265] The aforementioned communication components are configured to facilitate wired or wireless communication between the device containing the communication components and other devices. The device containing the communication components can access wireless networks based on communication standards, such as mobile hotspots (WiFi), second-generation (2G), third-generation (3G), fourth-generation (4G) / Long Term Evolution (LTE), fifth-generation (5G), or combinations thereof. In one exemplary embodiment, the communication components receive broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication components also include a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be based on Radio Frequency Identification (RFID), infrared, Ultra Wide Band (UWB), Bluetooth, and other technologies.
[0266] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.
[0267] The aforementioned storage medium can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium accessible to general-purpose or special-purpose computers.
[0268] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. The processor and storage medium can reside within an application-specific integrated circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components within an electronic device or host device.
[0269] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0270] The order of the embodiments described above is merely for illustrative purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, some processes described in the above embodiments and accompanying drawings include multiple operations appearing in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The sequence numbers are merely used to distinguish different operations, and the sequence numbers themselves do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types. "Multiple" means two or more, unless otherwise explicitly specified.
[0271] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0272] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0273] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A human-computer interaction method, characterized in that, include: Based on the input query information, a query plan and a response framework are generated. The query plan contains multiple subqueries for searching the information needed to respond to the query information. Execute the query based on the multiple subqueries to obtain the query results of the multiple subqueries; Based on the query results of the multiple subqueries and the response framework, a response to the query information is generated.
2. The method according to claim 1, characterized in that, The step of generating a query plan and response framework based on the input query information includes: Search and / or retrieve based on the input query information to obtain matching knowledge for the query information; Based on the query information and the matching knowledge of the query information, a query plan and a response framework are generated.
3. The method according to claim 2, characterized in that, The step of generating a query plan and response framework based on the query information and the matching knowledge of the query information includes: The query information and the matching knowledge of the query information are filled into the first prompt template to generate the first prompt information. The first prompt information is used to prompt the auxiliary generation model to generate a query plan and a response framework based on the query information and the matching knowledge of the query information. The first prompt information is input into the auxiliary generation model, which then generates a query plan and a response framework based on the first prompt information.
4. The method according to claim 1, characterized in that, The step of generating a query plan and response framework based on the input query information includes: The query information is filled into the second prompt template to generate the second prompt information. The second prompt information is used to prompt the auxiliary generation model to generate a query plan and response framework based on the query information. The second prompt information is input into the auxiliary generation model, which then generates a query plan and response framework based on the second prompt information.
5. The method according to claim 1, characterized in that, The step of executing a query based on the multiple subqueries to obtain the query results of the multiple subqueries includes: Based on the dependencies between multiple subqueries in the query plan, determine the subqueries that have no dependencies and the subqueries that have dependencies; Subqueries without dependencies are executed in parallel. For subqueries with dependencies, the queries are executed in the order of the dependencies to obtain the query results of each subquery.
6. The method according to claim 1, characterized in that, The step of generating response information for the query information based on the query results of the multiple sub-queries and the response framework includes: The query information, the response framework, the multiple sub-queries, and the query results are input into the human-computer interaction model. The human-computer interaction model then generates response information based on the response framework according to the query information, the response framework, the multiple sub-queries, and the query results.
7. The method according to claim 6, characterized in that, The step of inputting the query information, the response framework, the multiple sub-queries, and the query results into a human-computer interaction model, and generating response information based on the response framework through the human-computer interaction model, includes: The query information, the response framework, the multiple sub-queries, and the query results are filled into the third prompt template to generate the third prompt information. The third prompt information is used to prompt the human-computer interaction model to generate response information based on the response framework according to the query information, the response framework, the multiple sub-queries, and the query results. The third prompt information is input into the human-computer interaction model, and the human-computer interaction model generates a response based on the response framework according to the third prompt information.
8. The method according to claim 6, characterized in that, The step of inputting the response framework, the multiple sub-queries, and the query results into the human-computer interaction model, and generating response information based on the response framework through the human-computer interaction model, includes: Based on the maximum input length of the human-computer interaction model, the multiple subqueries and query results are divided into multiple groups, and each group contains at least one subquery and its corresponding query result. The response framework, subqueries and query results of each group are input into the human-computer interaction model. The human-computer interaction model generates the corresponding response information for each group based on the response framework, subqueries and query results of each group. The query information and the corresponding response information for each group are input into the human-computer interaction model. The human-computer interaction model integrates and sorts the response information for each group to generate the response information for the query information.
9. The method according to claim 8, characterized in that, The step of inputting the query information and the corresponding response information of each group into the human-computer interaction model, and integrating and sorting the corresponding response information of each group through the human-computer interaction model to generate response information for the query information includes: The query information and the corresponding response information of each group are filled into the fourth prompt template to generate the fourth prompt information. The fourth prompt information is used to prompt the human-computer interaction model to generate the response information of the query information by integrating and sorting the corresponding response information of each group. The fourth prompt information is input into the human-computer interaction model, and the human-computer interaction model generates a response to the query information based on the fourth prompt information.
10. The method according to claim 3 or 4, characterized in that, The training process for the generative model includes: Obtain lightweight pre-trained models; Training data is constructed based on domain knowledge, and the training data includes query samples, matching knowledge, query plans, and response frameworks corresponding to the query samples. The pre-trained model is fine-tuned using the training data to obtain an auxiliary generative model for generating query plans and response frameworks.
11. The method according to claim 10, characterized in that, The training data constructed based on domain knowledge includes: Acquire domain knowledge, and extract multiple query information from the domain knowledge as query samples; Search and / or retrieve based on the query sample to obtain matching knowledge of the query sample; The query sample and matching knowledge are populated into the first prompt template to obtain the first prompt information corresponding to the query sample; The first prompt information corresponding to the query sample is input into the pre-trained large language model, and the query plan and response framework corresponding to the query sample are generated through the pre-trained large language model.
12. The method according to any one of claims 1-9, characterized in that, The step of generating a query plan and response framework based on the input query information includes: Based on the input query information, a classification model is used to identify the category of the query information; If the query information is classified as a complex query, a query plan and a response framework are generated based on the input query information.
13. The method according to claim 12, characterized in that, Also includes: If the category of the query information is knowledge-based, the query information is input into the human-computer interaction model, and the human-computer interaction model directly generates the response information for the query information. If the category of the query information is real-time fact, a search is performed based on the query information to obtain candidate knowledge that matches the query information. The query information and candidate knowledge are then input into the human-computer interaction model, which generates a response to the query information based on the candidate knowledge.
14. A human-computer interaction method, characterized in that, include: User issues sent by the receiving device; A query plan and a response framework are generated based on the user question. The query plan contains multiple sub-queries for searching the information needed to answer the user question. Execute the query based on the multiple subqueries to obtain the query results of the multiple subqueries; Based on the query results of the multiple sub-queries and the response framework, generate response information for the user's question; The system returns a response to the user's question to the terminal device.
15. A human-computer interaction method, characterized in that, include: Retrieve the input query information; The query information is input into the retrieval decision model, which then determines whether a retrieval is necessary. Based on the judgment result, when retrieval is required, candidate knowledge matching the query information is retrieved based on the query information, the query information and the candidate knowledge are input into the human-computer interaction model, and the human-computer interaction model generates response information for the query information based on the candidate knowledge. If no retrieval is required, the query information is input into the human-computer interaction model, and the human-computer interaction model generates a response to the query information.
16. A human-computer interaction method, characterized in that, include: Identify the category of the input query information; When the type of query information is a complex query, a query plan and a response framework are generated based on the input query information. The query is executed based on the multiple sub-queries contained in the query plan to obtain the query results of the multiple sub-queries. The response information for the query information is generated based on the query results of the multiple sub-queries and the response framework. If the category of the query information is knowledge-based, the query information is input into the human-computer interaction model, and the human-computer interaction model generates the response information for the query information. When the category of the query information is real-time fact, candidate knowledge matching the query information is retrieved based on the query information, the query information and the candidate knowledge are input into the human-computer interaction model, and the human-computer interaction model generates the response information of the query information based on the candidate knowledge.
17. A server, characterized in that, include: At least one processor; as well as A memory that is communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, cause the server to perform the method according to any one of claims 1-16.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1-16.
19. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-16.