Human-computer interaction method, computing device, server, storage medium and program product

By asynchronously executing subprocesses in the RAG framework in the server side and gradually outputting the results, the problem of too long response time in the RAG framework is solved, improving user experience and system response speed.

CN120123501AInactive Publication Date: 2025-06-10ALIBABA CLOUD FEITIAN (HANGZHOU) CLOUD COMPUTING TECH CO LTD

Patent Information

Application Number
CN202510600733.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The current RAG framework faces a long time from user input questions to getting answers in practical applications, which leads to poor user experience and is difficult to meet users' requirements for real-time and rapid response.

Method used

By asynchronously executing multiple subprocesses of the enhanced generation process on the server side, and outputting their execution results after the first subprocess is completed, gradually providing intermediate results, and finally outputting the final answer after the last subprocess is completed.

Benefits of technology

It significantly reduces the time from input query to obtaining a preliminary response, improves user satisfaction and interactive experience, reduces user perception of latency, improves system response speed, and effectively improves system throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123501A_ABST
    Figure CN120123501A_ABST
Patent Text Reader

Abstract

The invention provides a man-machine interaction method, computing equipment, a server, a storage medium and a program product. The method relates to the field of man-machine interaction, and can receive input query information; executing a plurality of sub-processes of a retrieval enhancement generation process according to the query information, and asynchronously outputting execution results of the plurality of sub-processes; wherein the execution result of the last sub-process of the retrieval enhancement generation process is reply information of the query information, a user can see feedback of the system more quickly, the time from query input to preliminary response obtaining is remarkably shortened, and the user experience is improved by gradually providing intermediate results. The user can obtain useful information in the process of waiting for the final answer, the instant feedback can improve the satisfaction and interaction experience of the user, even if the final answer needs to be generated for a long time and the user obtains part of information in the period, the information flow mode reduces the perception of the user on delay, and the user experience is improved. And the system response speed is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and particularly to a human-computer interaction method, a computing device, a server, a storage medium, and a program product. Background Art

[0002] RAG (Retrieval-Augmented Generation) is a method that combines information retrieval with a generation model, aiming to improve the accuracy and credibility of answers generated by a large language model (LLM) by integrating external knowledge sources. The core idea of this method is to use information retrieval technology to extract relevant information from a large external database, and then input this information as context into the generation model to generate more accurate and credible answers.

[0003] The current RAG framework faces some challenges in practical applications. One of them is that the time from the user inputting a question to obtaining an answer is relatively long, which has a negative impact on the user experience and is difficult to meet the requirements of users for real-time and quick response. Summary of the Invention

[0004] The present application provides a human-computer interaction method, a computing device, a server, a storage medium, and a program product to solve the problems of high end-to-end latency and poor real-time performance.

[0005] In a first aspect, the present application provides a human-computer interaction method, including:

[0006] Receiving the input query information; performing multiple subprocesses of a retrieval-augmented generation process according to the query information, and asynchronously outputting the execution results of the multiple subprocesses; wherein, the execution result of the last subprocess of the retrieval-augmented generation process is the reply information to the query information.

[0007] In a second aspect, the present application provides a human-computer interaction method based on a document knowledge base, including:

[0008] Receiving the input query information; performing multiple subprocesses of a retrieval-augmented generation process according to the query information and the document knowledge base, and asynchronously outputting the execution results of the multiple subprocesses; wherein, the multiple subprocesses include multi-round rewriting, knowledge word recall, knowledge rewriting, document retrieval, document ranking, and generating a reply.

[0009] In a third aspect, the present application provides a computing device, including:

[0010] A memory and a processor; the memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the methods described in any of the foregoing aspects are implemented.

[0011] In a fourth aspect, the present application provides a server, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the server is caused to execute the methods provided in any of the foregoing aspects.

[0012] In a fifth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the methods provided in any of the foregoing aspects are implemented.

[0013] In a sixth aspect, the present application provides a computer program product, including a computer program, which when executed by a processor implements the methods provided in any of the foregoing aspects.

[0014] For the human-computer interaction method, computing device, server, storage medium and program product provided by the present application, after the server completes the first sub-process, it outputs its execution result to the user. Except for the last sub-process, the execution results of other sub-processes are regarded as intermediate results. The server outputs these intermediate results asynchronously, and finally outputs its corresponding execution result after the last sub-process is completed, that is, the final answer to the user's query information. From the user's perspective, the user can see the system's feedback faster, significantly reducing the time from inputting the query to obtaining a preliminary response. And by gradually providing intermediate results, the user can obtain useful information during the waiting for the final answer. This kind of instant feedback can improve the user's satisfaction and interaction experience. Even if the final answer takes a long time to generate, the user has obtained some information during this period. This way of information flow reduces the user's perception of latency, improves the system response speed, and avoids the resource load pressure caused by a large amount of data calculation and transmission in a short time, which can effectively improve the system throughput, reduce the system's first token output time and end-to-end delay time. Description of the Drawings

[0015] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0016] Figure 1 It is a schematic diagram of the system architecture of the human-computer interaction system provided by the present application;

[0017] Figure 2 Flowchart of a human - machine interaction method provided by an exemplary embodiment of the present application;

[0018] Figure 3 Flowchart of asynchronously outputting execution results of multiple sub - processes provided by an exemplary embodiment of the present application;

[0019] Figure 4 Schematic diagram of progress information of a retrieval - enhanced generation process provided by an exemplary embodiment of the present application;

[0020] Figure 5 Flowchart of multiple sub - processes for executing a retrieval - enhanced generation process provided by an exemplary embodiment of the present application;

[0021] Figure 6 Schematic diagram of the process of a human - machine interaction method provided by an exemplary embodiment of the present application;

[0022] Figure 7 Schematic diagram of the process of extracting message data provided by an exemplary embodiment of the present application;

[0023] Figure 8 Flowchart of a human - machine interaction method based on a document knowledge base provided by an exemplary embodiment of the present application;

[0024] Figure 9 Block diagram of a computing device according to an embodiment of the present application;

[0025] Figure 10 Schematic diagram of a server provided by an embodiment of the present application.

[0026] Through the above - mentioned drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0027] Here, exemplary embodiments will be described in detail, and examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0028] It should be noted that the user information (including but not limited to user device information, user attribute information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0029] First, the nouns involved in this application are explained as follows:

[0030] Message queue: A mechanism for inter-process communication, mainly used to provide asynchronous communication in a distributed system. A message queue can temporarily store messages and ensure the reliable transfer of messages between different system components or services. This approach can enhance the scalability and reliability of the system and is commonly used to decouple the relationship between producers and consumers. Common message queue implementations include Kafka, RabbitMQ, RocketMQ, etc.

[0031] VLM (Vision Language Model): A multimodal model that can simultaneously process vision (images / videos) and text. By jointly learning the information of the two modalities, it realizes cross-modal understanding and generation.

[0032] RAG (Retrieval-Augmented Generation): A technical architecture that integrates information retrieval and text generation, mainly addressing the factual bias problem in the content generated by large language models. Its core principle is to retrieve external knowledge bases (such as databases, document sets) in real time and use relevant text fragments as context inputs for the generation model, enabling the model output to have factual accuracy and domain pertinence. The typical workflow is divided into three stages: ① Retrieve the top-K relevant documents according to the user's Query; ② Concatenate the retrieval results with the original input; ③ The generation model outputs the final result based on the enhanced context. It is commonly found in scenarios that require dynamic knowledge updates, such as intelligent customer service and knowledge base Q&A.

[0033] LLM (Large Language Model): An ultra-large-scale language model trained based on massive text data, usually based on the Transformer architecture. By learning the statistical laws and semantic associations in the text, it has the ability to generate, understand, reason, and translate natural language.

[0034] Large models refer to deep learning models with a large number of model parameters, usually containing hundreds of millions, tens of billions, or even hundreds of billions of model parameters. Large models can also be referred to as Foundation Models (FM). Through the pre-training of large models with a large amount of unlabeled corpus, a pre-trained model with over hundreds of millions of parameters is produced. Such a model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.

[0035] When large models are applied in practice, only a small number of samples are needed to fine-tune the pre-trained model for use in different tasks. Large models can be widely applied in fields such as natural language processing (NLP), computer vision, etc. Specifically, they can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), image generation, etc., as well as tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.

[0036] With the rapid development of artificial intelligence technology, large language models (LLMs) have demonstrated powerful language recognition, understanding, and reasoning abilities in the field of natural language processing. However, LLMs also face many challenges in practical applications, such as the limitations, lag, hallucination problems, and data security of knowledge. To address these issues, Retrieval-Augmented Generation (RAG) technology has emerged. RAG is a method that combines information retrieval with a generation model, aiming to improve the accuracy and credibility of the answers generated by LLMs by integrating external knowledge sources.

[0037] The emergence of RAG technology has brought revolutionary changes to the application of LLMs. Especially in knowledge-intensive tasks, RAG allows for continuous knowledge updates and the integration of domain-specific information, thus greatly enhancing the practicality and accuracy of LLMs. However, despite the significant theoretical advantages of RAG, it still faces some technical problems in practical applications, especially in terms of response time and user experience.

[0038] In a typical RAG scenario of a document knowledge base, the RAG process usually includes multiple stages, such as multi-round rewriting, knowledge word recall, knowledge rewriting, document recall, document ranking, summary answering, and so on. The complexity and diversity of these stages lead to a high computational load and a long processing time. The current RAG framework usually outputs an answer to the customer's question only after the last stage is completed. This process design directly results in a long waiting time for the user to see the answer after inputting the question, causing a high end-to-end latency and a poor user experience.

[0039] To address the above technical problems, the present application provides a human-computer interaction method. After receiving the query information input by the user, the server performs processing based on the RAG process. Similar to the prior art, this RAG process includes multiple stages, each stage contains at least one sub-process, and each sub-process generates a corresponding execution result after completion. However, the innovation of the present application is that the server no longer waits until the last sub-process of the last stage is completed before outputting the final answer to the user's query information. Instead, after the first sub-process is completed, the server outputs its execution result to the user. Except for the last sub-process, the execution results of other sub-processes are regarded as intermediate results. The server outputs these intermediate results asynchronously and finally outputs the corresponding execution result of the last sub-process, that is, the final answer to the user's query information, after the last sub-process is completed.

[0040] Currently, the stages and sub-processes in the RAG process can be divided differently according to the actual scenario, and the present application does not make any restrictions on this.

[0041] Exemplarily, the query information input by the user is "What about City A?" The "multi-round rewriting" stage of the RAG process will perform multi-round rewriting based on the user's previous query information. For example, if the user's previous query information was "What's the weather like in City B?", the rewritten result might be "What's the weather like in City A", and the server directly outputs "What's the weather like in City A" to the user without waiting to retrieve the detailed information about the weather in City A before outputting to the user.

[0042] Therefore, from the user's perspective, the user can see the system's feedback faster, significantly reducing the time from inputting the query to obtaining a preliminary response. And by gradually providing intermediate results, the user can obtain useful information while waiting for the final answer. This kind of instant feedback can improve the user's satisfaction and interaction experience. Even if the final answer takes a long time to generate, the user has already obtained some information during this period. This way of information flow reduces the user's perception of latency, improves the system response speed, and avoids the resource load pressure caused by a large amount of data calculation and transmission in a short time.

[0043] Figure 1 This is a schematic diagram of the system architecture of the human-computer interaction system provided by this application. As Figure 1 shown, the system architecture includes a server and client devices. Among them, there is a communicable communication link between the server and the client devices, which can realize the communication connection between the server and the client devices.

[0044] Among them, the server is a device with computing power deployed in the cloud or locally, such as a cloud cluster, etc. The server is the server device in the question-and-answer system and is responsible for generating corresponding reply information based on the query information input by the user.

[0045] The client device can be an electronic device running the client of the human-computer interaction system, specifically a hardware device with network communication functions, computing functions, and information display functions, including but not limited to smartphones, tablets, desktop computers, local servers, cloud servers, etc. The user conducts human-computer interaction with the server through the client device used to achieve human-computer intelligent dialogue / question answering. Among them, the human-computer interaction system can be a question-and-answer system, a smart assistant, a smart robot, etc.

[0046] In this embodiment, the user inputs query information through the client, the client sends the query information input by the user to the server, the server receives the query information sent by the client, and performs a retrieval enhancement generation process according to the query information. The retrieval enhancement generation process includes multiple sub-processes. Exemplarily, in the RAG scenario of the document knowledge base, the multiple sub-processes included in the retrieval enhancement generation process can be: multi-round rewriting, knowledge word recall, knowledge rewriting, document recall, document ranking, summary answering. The execution results of the sub-processes can be output to the client after they are completed respectively, without waiting for all sub-processes to be completed and then output to the client together. Among them, the execution result corresponding to multi-round rewriting is the first rewriting result, the execution result corresponding to knowledge word recall is the first recall result, the execution result corresponding to knowledge rewriting is the second rewriting result, the execution result corresponding to document recall is the second recall result, and the execution result corresponding to document ranking is the ranking result; the execution result of the last sub-process, that is, "summary answering", is the reply information of the query information.

[0047] Furthermore, after receiving any execution result, the client displays the execution result on the display interface.

[0048] The human-computer interaction method provided by this application can be applied to question-and-answer systems in various fields to reduce the end-to-end latency, improve the system response speed, meet the user's requirements for real-time and fast response, and improve the user experience.

[0049] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problems. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below in conjunction with the accompanying drawings.

[0050] Figure 2 This is a flowchart of a human-computer interaction method provided for an exemplary embodiment of this application. The execution subject of this embodiment is the server in the aforementioned system architecture. As Figure 2 shown, the specific steps of this method are as follows:

[0051] Step S201: Receive the input query information.

[0052] Among them, the query information is usually the question or the information requested input by the user. These query information can be questions in natural language form, such as "What's the weather like today?" or "What is Einstein's theory of relativity?" etc. The purpose of the query information is to enable the system to understand the user's needs in order to provide relevant answers or information.

[0053] Specifically, the user can input the query information through the client, and the client sends the query information input by the user to the server, and the server receives the query information input by the user.

[0054] Step S202: Execute multiple subprocesses of the retrieval-augmented generation process according to the query information and asynchronously output the execution results of the multiple subprocesses.

[0055] Among them, the execution result of the last subprocess of the retrieval-augmented generation process is the reply information of the query information.

[0056] Retrieval-Augmented Generation (RAG) is a technology that combines information retrieval and generation models to improve the performance of question-answering systems, dialogue systems, and other natural language processing applications. The core idea of RAG is to enhance the ability of the generation model by retrieving relevant information, so as to provide more accurate and context-related answers.

[0057] The stages and subprocesses in the RAG process can be divided differently according to the actual scenario, and this application does not make any restrictions on this. In an optional implementation manner, the division of the stage and the subprocess is the same, that is, one stage corresponds to one subprocess; in another optional implementation manner, each stage includes at least one subprocess, and the at least one subprocess included is executed in sequence; in another optional implementation manner, each stage includes at least one subprocess, and the at least one subprocess included is executed in parallel. Exemplarily, the retrieval stage can include multi-way retrieval, the subprocess corresponding to each way of retrieval, and the multiple subprocesses corresponding to the multi-way retrieval are executed in parallel.

[0058] RAG may include the following stages: query rewriting, retrieval, ranking, and generation. Each stage may include one or more sub - processes.

[0059] The following is a detailed description of each stage and its possible sub - processes:

[0060] Query rewriting: The purpose is to optimize the query input by the user to improve the retrieval effect. The rewritten query may be more in line with the format and structure of the information source.

[0061] The sub - processes corresponding to query rewriting may include:

[0062] Multi - round rewriting: The purpose is to gradually improve the quality and relevance of the query by iteratively rewriting the query multiple times. Specific process: Make a preliminary rewrite of the original query, which may include simple synonym replacement or grammar adjustment, or may also be rewritten or content - expanded according to the context.

[0063] Knowledge word recall: The purpose is to identify knowledge words or terms related to the query. Specific process: Retrieve relevant terms from the knowledge base or domain - specific dictionaries, identify the terms already in the query, and recall relevant knowledge words.

[0064] Knowledge rewriting: The purpose is to rewrite the query using the recalled relevant indicator words to enrich the query content.

[0065] Retrieval: The purpose is to find the most relevant documents or data fragments from the information source.

[0066] The sub - processes corresponding to retrieval may include:

[0067] Index search: Use an inverted index or other search techniques to quickly find relevant content in the document collection.

[0068] Boolean retrieval: Apply Boolean logic (such as AND, OR, NOT) to filter and select documents.

[0069] Vector retrieval: Utilize the vector space model or embedding technology to retrieve relevant documents according to semantic similarity.

[0070] Optionally, retrieval may include multi - path retrieval. For each path of retrieval, index search, Boolean retrieval, and vector retrieval can be executed sequentially to obtain the search result, or different retrieval strategies can be set for each path of retrieval. For example, the retrieval strategy corresponding to the first path of retrieval is index search, the retrieval strategy corresponding to the second path of retrieval is Boolean retrieval, and the retrieval strategy corresponding to the third path of retrieval is vector retrieval. Then, the retrieval results corresponding to each path of retrieval are merged to obtain the final retrieval result.

[0071] Sorting: The purpose is to sort the retrieved documents or data fragments to determine their relevance and priority.

[0072] The sub-processes corresponding to sorting may include:

[0073] Relevance scoring: Use TF-IDF (Term Frequency-Inverse Document Frequency), BM25 (Best Matching 25), or a neural network model to score the documents.

[0074] Sorting algorithm: Apply a sorting algorithm (such as PageRank or learning to rank) to rank the documents, where PageRank is a web page ranking algorithm.

[0075] Deduplication and filtering: Remove duplicate or irrelevant entries to improve the quality of the results.

[0076] Generation: The purpose is to use a generation model to create answers in natural language form, combining the retrieved information and the user query.

[0077] The sub-processes corresponding to generation may include:

[0078] Context integration: Combine the retrieved information with the user query to form the input of the generation model.

[0079] Answer generation: Use a generation model to generate natural language answers, where the generation model can be any one of the LLMs.

[0080] Answer optimization: Post-process the generated answers to improve their fluency and accuracy.

[0081] The sub-processes of each stage can be adjusted according to specific applications and system designs to meet specific requirements and performance requirements. This application does not limit the division of stages and sub-processes. Further, the sub-processes can also be divided into finer-level sub-processes.

[0082] Execute multiple sub-processes of the retrieval-augmented generation process according to the query information and asynchronously output the execution results of the multiple sub-processes.

[0083] Among them, asynchronous output is a way of processing and outputting data, where the execution and result output of the sub-processes are non-blocking, that is, there is no need to wait for all related tasks to complete before outputting. Instead, asynchronous output allows each task to output its result after completion without waiting for other tasks to complete.

[0084] After receiving the query information input by the user, the server starts the Retrieval-Augmented Generation (RAG) process, which consists of multiple sub-processes. First, the server may enter the query rewriting stage, optimizing the query by performing sub-processes such as multi-round rewriting, knowledge word recall, and knowledge rewriting. Next, the server enters the information retrieval stage, retrieving content related to the query from relevant information sources. Subsequently, the server ranks the retrieved content to determine its relevance and priority. Finally, the server uses a generation model to generate an answer in natural language form.

[0085] Throughout the process, the server outputs the execution results of each sub-process asynchronously without waiting for all sub-processes to complete. Optionally, when the server outputs the execution results of each sub-process asynchronously, it can output the result corresponding to each sub-process after each sub-process is completed, or it can configure whether to output the result of each sub-process according to requirements, either output or not output.

[0086] Among them, the execution result of the last sub-process of the Retrieval-Augmented Generation process is the reply information to the query information input by the user. The execution results of sub-processes other than the last one are all intermediate results. That is, each intermediate result or the final reply information of this application will be output after being generated, without waiting to be output together after the final reply is generated.

[0087] In this embodiment, after the first sub-process is completed, the server outputs its execution result to the user. The execution results of other sub-processes except the last one are regarded as intermediate results. The server outputs these intermediate results asynchronously, and finally outputs the corresponding execution result after the last sub-process is completed, that is, the final answer to the user's query information. From the user's perspective, the user can see the system's feedback faster, significantly reducing the time from inputting the query to obtaining a preliminary response. And by gradually providing intermediate results, the user can obtain useful information during the waiting for the final answer. This kind of instant feedback can improve the user's satisfaction and interaction experience. Even if the final answer takes a long time to generate, the user has already obtained some information during this period. This way of information flow reduces the user's perception of delay, making the system respond quickly, and avoiding the resource load pressure caused by a large amount of data calculation and transmission in a short time, which can effectively improve the system throughput and reduce the first token output time and end-to-end delay time.

[0088] Figure 3 This is a flowchart for asynchronously outputting the execution results of multiple sub-processes provided for an exemplary embodiment of this application. As Figure 3 shown, the specific steps of this method are as follows:

[0089] Step S301: Execute multiple subprocesses of the retrieval enhancement generation process according to the query information, encapsulate the execution results of each subprocess into a message, and insert the message into an asynchronous message queue; wherein, the message includes an output flag, and the output flag indicates whether the execution result needs to be output.

[0090] Specifically, for any subprocess in the retrieval enhancement generation process, after executing the subprocess, obtain the execution result corresponding to the subprocess, encapsulate the execution result corresponding to the subprocess into a message, and insert the encapsulated message into the asynchronous message queue. The encapsulated message includes an output flag, and the output flag can indicate whether the execution result needs to be output to the user.

[0091] Among them, the asynchronous message queue is a mechanism for implementing asynchronous communication in a distributed system. It allows different components or services to communicate through messages without directly connecting to each other or waiting synchronously. This mechanism decouples the execution processes of the sender and the receiver through asynchronous message passing, improving the flexibility and scalability of the system.

[0092] Optionally, in addition to the asynchronous message queue, other methods can also be used to implement asynchronous message passing. This application does not limit this. Exemplarily, it can be middleware for message streams, event buses, task queues, service buses, etc. Among them, the event bus can be used in an event-driven architecture; the task queue is suitable for task distribution and asynchronous processing, and the service bus can provide enterprise-level message passing and integration services.

[0093] The output flag is an identifier used to indicate whether a certain execution result needs to be output. It can have various forms, specifically depending on the design and implementation requirements of the system.

[0094] In an optional implementation manner, a field (for example, `outputFlag`) can be added to the metadata of the message.

[0095] This field can be a boolean value, a string, or other types, used to indicate the output requirement.

[0096] In another optional implementation manner, a boolean flag can be used, in the form of a simple boolean value (`true` or `false`), where `true` indicates that the execution result needs to be output, and `false` indicates that it does not need to be output.

[0097] In yet another optional implementation manner, a status code can be used, in the form of an integer or a string status code. Whether it needs to be output can be represented by a specific status code. For example, `200` indicates that it needs to be output, and `204` indicates that it does not need to be output.

[0098] Step S302: Obtain the message from the asynchronous message queue through an asynchronous output thread. If the output flag in the message indicates that output is required, output the execution result in the message.

[0099] The asynchronous output thread is an independently running thread used to obtain messages from the asynchronous message queue. For any message in the asynchronous message queue, after obtaining the message, determine the output flag corresponding to the message. If the output flag in the message indicates that output is required, output the execution result in the message to the client for display on the client's display interface. If the output flag in the message indicates that output is not required, the execution result in the message will not be output to the client.

[0100] In this way, by encapsulating the execution results of each subprocess into messages, inserting the encapsulated messages into the asynchronous message queue, and using the asynchronous output thread to obtain and output messages from the asynchronous message queue, asynchronous message transmission is achieved, improving the system's response speed and throughput. Moreover, by including an output flag in the message, the system can flexibly control which execution results need to be output, avoiding unnecessary information output and improving the system's output efficiency.

[0101] In addition, before encapsulating the execution results of each subprocess into messages, the subprocesses that need to output the execution results to the client can be configured. In this application, the subprocesses that need to output the execution results are called the first subprocesses, and the subprocesses that do not need to output the execution results are called the second subprocesses, that is, the second subprocesses are the subprocesses other than the first subprocesses.

[0102] In an optional implementation, for any subprocess, an identifier can be defined in the code corresponding to the subprocess to indicate the output requirement of the subprocess. For example, a boolean value or an enumeration type can be used. The identifiers in the code corresponding to the first subprocesses can all be "true", indicating that the execution results need to be output, and the identifiers in the code corresponding to the second subprocesses can all be "false", indicating that the execution results do not need to be output.

[0103] When encapsulating the execution results of each subprocess into messages and inserting the messages into the asynchronous message queue, the following two cases can be divided:

[0104] For any first subprocess, the server encapsulates the execution result of the first subprocess into a first message, where the first message includes a first output flag, and the first output flag can indicate that the execution result needs to be output. The server inserts the first message into the asynchronous message queue.

[0105] For any second sub - process, the server encapsulates the execution result of the second sub - process into a second message, where the second message includes a second output flag, and the second output flag can indicate that the execution result does not need to be output. The server inserts the second message into the asynchronous message queue.

[0106] The first output flag and the second output flag are different. The form of the output flag is as described in the above embodiments and will not be elaborated here.

[0107] In this way, by configuring the first sub - process that needs to output the execution result, the first message encapsulated by the execution result corresponding to the first sub - process will contain a first output identifier, and the first output identifier can indicate to output the execution result. The second message encapsulated by the execution result corresponding to the second sub - process will contain a second output identifier, and the second output identifier can indicate not to output the execution result. It is possible to freely decide according to requirements which need to be output to the client and which do not. Different sub - processes can be configured according to needs whether to output their execution results without large - scale modification of the entire system. This makes the system more scalable and able to more easily adapt to future demand changes.

[0108] Optionally, after configuring the sub - process that needs to output the execution result to the client, it is also possible to only encapsulate the execution result of the first sub - process into a message and insert the encapsulated message into the asynchronous message queue. Through the asynchronous output thread, the message is retrieved from the asynchronous message queue and output. It should be noted that in the scenario where only the execution result of the first sub - process is encapsulated into a message, the message encapsulated by the execution result of the first sub - process may not include an output flag. After retrieving the message from the asynchronous message queue, there is no need to judge the output flag anymore, and it can be directly output.

[0109] In this way, by directly encapsulating the execution result of the sub - process that needs to be output into a message and inserting it into the asynchronous message queue without including an output flag, the logic of message processing is simplified. In the asynchronous output thread, after retrieving the message from the queue, it can be directly output without additional judgment steps. This simplification helps to reduce code complexity and improve code readability and maintainability.

[0110] The retrieval - enhanced generation process includes multiple sub - processes. Some sub - processes are executed by a large - model. Since the inference time of the large - model is usually relatively long, if the execution result is output after all inferences are completed, the system response time is long and the experience is poor, which in turn affects the user's satisfaction and trust in the system.

[0111] To solve the above - mentioned technical problems, the following methods can be adopted:

[0112] For any sub - process, if the sub - process is executed by a large - model, then the result fragments output by the large - model each time are encapsulated into messages, and the messages are inserted into the asynchronous message queue.

[0113] The definition of the result fragment can be adjusted according to the requirements and context of the specific application.

[0114] In one example, the result fragment can be a complete sentence. That is, every time the large - model outputs a complete sentence, the sentence is encapsulated into a message and the message is inserted into the asynchronous message queue.

[0115] In another example, the result fragment can be a unit with complete meaning (such as a complete idea or argument).

[0116] In yet another example, the result fragment can be all the content output by the large - model within a fixed time interval.

[0117] Since the large - model generates text token - by - token instead of outputting the entire paragraph at once, a streaming output strategy can be adopted. Encapsulating the result fragments output each time into messages and inserting them into the asynchronous message queue can significantly reduce latency and improve the system's response speed.

[0118] When encapsulating the result fragments output by the large - model each time into messages, for the last result fragment output by the large - model, the message encapsulated from this result fragment contains an end marker, and for the non - last result fragments output by the large - model, the messages encapsulated from these result fragments contain a generating marker.

[0119] In an optional implementation, the end marker or the generating marker can be added to the beginning or end of the corresponding result fragment. The generating marker indicates that the generation is not yet complete, and the end marker indicates that the generation is complete. Finally, the result fragment with the added end marker or generating marker is encapsulated into a message.

[0120] In another optional implementation, a fixed field can be added to the message. The value corresponding to this fixed field includes: generating and end. If the value of this fixed field is generating, it means the generation is not yet complete; if the value of this fixed field is end, it means the generation is complete.

[0121] The server, through an asynchronous output thread, retrieves messages from the asynchronous message queue, extracts the result fragments and the generating marker (or end marker) from the messages. The client can not only receive the result fragments but also determine the progress information based on whether it is a generating marker or an end marker. Real - time feedback on the progress information can significantly improve the user experience, especially in applications that require immediate response.

[0122] This application does not limit the specific display form of the progress information of the retrieval-augmented generation process.

[0123] In an alternative implementation, Figure 4 is a schematic diagram of the progress information of the retrieval-augmented generation process provided by an exemplary embodiment of this application. As Figure 4 shown, the retrieval-augmented generation process includes 5 subprocesses, namely subprocess 1, subprocess 2, subprocess 3, subprocess 4, and subprocess 5. The completed subprocess can be represented by "√", the subprocess in execution can be represented by "◎", and the subprocess not yet started can be represented by "×". Therefore, subprocess 1 and subprocess 2 are completed, subprocess 3 is in execution, and subprocess 4 and subprocess 5 have not yet started.

[0124] Optionally, for any subprocess, after the subprocess is fully executed, the obtained execution result can be encapsulated into a message, and the encapsulated message can be inserted into the asynchronous message queue.

[0125] Since each subprocess generates only one message, the message processing logic becomes simpler. This reduces the management requirements for the message order and integrity, reduces the system complexity, and reducing the number of messages can reduce the load of the message queue and the overhead of the system in message transmission and processing.

[0126] Optionally, the encapsulated message also includes the information of the subprocess where the execution reaches. When outputting the execution result in the message, the information of the subprocess where the execution reaches should be output.

[0127] Exemplarily, after subprocess 2 is executed and the execution result is obtained, the execution result is encapsulated into a message. The encapsulated message includes the information of the subprocess, that is, subprocess 2. When outputting the execution result, the information of subprocess 2 should also be output. If the execution result is encapsulated into a message only after each subprocess finishes execution, when the client receives the execution result and the information of subprocess 2, it can be determined that subprocess 2 has been executed, and the progress information corresponding to subprocess 2 can be modified to "√".

[0128] If a certain subprocess adopts the method of streaming output, that is, each output result segment is encapsulated into a message, when outputting the execution result corresponding to the message, in addition to outputting the information of subprocess 2, a status flag should also be output, where the status flag can be "generating" or "ended". If the client receives the execution result, the information of subprocess 2, and "generating", it means that subprocess 2 is being executed, but subprocess 2 has not been executed yet. Therefore, the progress information of subprocess 2 can be modified to "◎".

[0129] If the client receives the execution result, the information of sub - process 2, and at the end, it indicates that sub - process 2 has not been executed completely yet, even though it seems that the client has received relevant information. Therefore, the progress information of sub - process 2 can be modified to "√".

[0130] In this way, by outputting the information of the executed sub - processes, users can obtain detailed information about the system progress in real - time. This kind of instant feedback can reduce users' anxiety and sense of uncertainty, and enhance the overall user experience.

[0131] Figure 5 It is a flowchart of multiple sub - processes for executing a retrieval - enhanced generation process provided by an exemplary embodiment of the present application. As Figure 5 shown, the specific steps of the method are as follows:

[0132] Step S501: Asynchronously execute the multiple sub - processes according to the query information.

[0133] Step S502: When executing any one of the sub - processes, obtain the dependency data of the current sub - process from the asynchronous message queue, execute the current sub - process based on the dependency data, and insert the execution result of the current sub - process into the asynchronous message queue. The dependency data of the current sub - process includes: the execution results of at least one sub - process executed before the current sub - process.

[0134] For any first sub - process except the first sub - process, during the execution of this first sub - process, the execution results corresponding to one or more sub - processes before this first sub - process are required. In the present application, the execution results corresponding to one or more sub - processes before this first sub - process are referred to as the dependency data of this first sub - process.

[0135] Exemplarily, the multiple sub - processes of the retrieval - enhanced process include: sub - process 1, sub - process 2, sub - process 3, sub - process 4, and sub - process 5. Among them, during the execution of sub - process 2, the execution result corresponding to sub - process 1 may be required, then the execution result corresponding to sub - process 1 is the dependency data of sub - process 2. During the execution of sub - process 3, the execution result corresponding to sub - process 1 may be used only, then the execution result of sub - process 1 is the dependency data of sub - process 3, or the execution result corresponding to sub - process 2 may be used only, then the execution result of sub - process 2 is the dependency data of sub - process 3, or the execution results corresponding to both sub - process 1 and sub - process 2 may be used, then the execution results of sub - process 1 and sub - process 2 are the dependency data of sub - process 3.

[0136] In the prior art, multiple sub - processes included in the retrieval - enhanced generation process are executed sequentially. However, sequential execution requires that each sub - process must start after the previous sub - process is completed. This means that if a certain sub - process takes a long time, it will cause delays in the entire process and reduce the overall efficiency. Also, in sequential execution, system resources (such as the central processing unit and memory) may be idle at certain times because they can only be used for the currently executing sub - process and cannot be used for other sub - processes simultaneously, resulting in waste of system resources.

[0137] Although the execution of the first sub - process of the first type depends on the execution results of the previous sub - processes during the execution process, it does not require the execution results of the previous sub - processes at the very beginning. Instead, a pre - processing process needs to be carried out first. During the pre - processing process, the execution results of the previous sub - processes are not required. After the pre - processing is completed, the execution results of the previous sub - processes will be used. When needed, obtain them from the asynchronous message queue and then execute the current first sub - process based on the obtained dependent data.

[0138] Therefore, this application adopts an asynchronous execution method for multiple sub - processes. Asynchronous execution allows multiple sub - processes to run simultaneously, thus significantly improving the overall processing efficiency. Each sub - process can run in parallel without mutual blocking, reducing the total execution time; and asynchronous execution can make better use of system resources. When multiple sub - processes run in parallel, the system can make full use of multi - core CPUs (Central Processing Units) and multi - threading technology to improve resource utilization; through parallel processing, asynchronous execution can significantly reduce the response time, enabling the system to return results faster and enhancing the user experience.

[0139] Optionally, obtaining the dependent data of the current sub - process from the asynchronous message queue and executing the current sub - process based on the dependent data includes:

[0140] Obtain the message of the target sub - process on which it depends from the asynchronous message queue, and obtain the execution result of the target sub - process from the message;

[0141] If the message contains a generating flag, store the obtained execution result fragment of the target sub - process, and continue to obtain the message of the target sub - process from the asynchronous message queue until a message containing a generation - end flag is obtained, and then obtain the execution result fragment of the target sub - process in the message containing the generation - end flag;

[0142] Integrate the obtained execution result fragments of the target sub - process to obtain the complete execution result of the target sub - process;

[0143] Execute the current subprocess according to the complete execution result of the target subprocess.

[0144] Specifically, the target subprocess can be one or more. Since the message contains the execution result and the information of the subprocess corresponding to the execution result, the message of the target subprocess on which it depends can be obtained from the asynchronous message queue.

[0145] For any message of the target subprocess obtained:

[0146] If the message contains a generating flag, store the execution result segment in the message and continue to obtain the message of the target subprocess from the asynchronous message queue; if the message contains an end flag, store the execution result segment in the message and integrate the multiple execution result segments corresponding to the target subprocess to obtain the complete execution result corresponding to the target subprocess. If the message still contains a generating flag, continue to store the execution result segment in the message and continue to obtain the message of the target subprocess from the asynchronous message queue until the obtained message contains an end flag, and then integrate the multiple execution results corresponding to the target subprocess.

[0147] If the message does not contain a generating flag nor an end flag, the execution result in the message is the complete execution result of the corresponding subprocess.

[0148] In the case where there are multiple target subprocesses, execute the current subprocess according to the complete execution results corresponding to the multiple target subprocesses respectively.

[0149] Optionally, in the scenario where the target subprocess corresponds to multiple messages, the message corresponding to the target subprocess further includes: sequence information, which is used to indicate the order of the execution result among all the execution results corresponding to the target subprocess. The execution results of the target subprocess obtained can be integrated according to the sequence information to obtain the complete execution result of the target subprocess.

[0150] In this way, by obtaining the execution result of the target subprocess from the asynchronous message queue, the system can ensure that the current subprocess obtains complete and accurate dependent data before execution, and this mechanism reduces errors caused by incomplete or inconsistent data.

[0151] Figure 6 It is a schematic flowchart of the human-computer interaction method provided by an exemplary embodiment of the present application. Figure 7A flowchart for extracting message data provided by an exemplary embodiment of the present application. The human-computer interaction method provided by the present application is a RAG (Retrieval-Augmentation-Generation) framework that realizes asynchronous output based on a message queue. Through decoupling design using a message queue and an incremental streaming output protocol, the retrieval, sorting, and generation stages in the RAG process are separated from the data calculation and transmission process, thereby improving the scalability of the system. Among them, the incremental streaming output protocol is a network protocol for efficiently transmitting data streams, especially suitable for scenarios that require real-time or near-real-time updates. Its main goal is to reduce bandwidth consumption and improve transmission efficiency by only transmitting the incremental change part of the data instead of the entire data set.

[0152] This method distributes the calculation and data transmission over a longer period of time instead of processing a large number of data requests instantaneously, which can manage resources and loads more effectively. At the same time, this method can provide real-time feedback on intermediate results, significantly reducing the output time of the first token and the end-to-end latency time, thereby enhancing the user's product experience.

[0153] The specific steps are as follows:

[0154] Step 1: Data production. Taking the typical RAG scenario of a document knowledge base as an example, the main sub - processes it executes are the "multi - round rewriting" sub - process, the "knowledge word recall" sub - process, the "knowledge rewriting" sub - process, the "document recall" sub - process, the "document sorting" sub - process, and the "summary answering" sub - process. The execution of each sub - process will generate corresponding execution results and information about the currently executed sub - process, that is, the status information of which sub - process is being executed. For any sub - process, if it is necessary to display the execution result of this sub - process in the final conversation, then encapsulate the execution result corresponding to this sub - process into a message and push this message to the asynchronous message queue; if it is necessary to display the execution status and execution result of this sub - process in the final conversation, then encapsulate the execution status and execution result corresponding to this sub - process into a message and push this message to the asynchronous message queue (here, the asynchronous message queue is used as an example for description, mainly to introduce the implementation idea of this solution. In a distributed system, using middleware such as Kafka, RabbitMQ, RocketMQ can also achieve the same effect). For sub - processes such as multi - round rewriting, knowledge rewriting, and summary answering that need to call large language models (LLMs) or multi - modal models (VLMs), in the case of streaming output, each token generated by the model will be pushed to the asynchronous message queue. If it is a non - streaming call, then it will be pushed to the asynchronous message queue after obtaining the complete execution result. During the pushing process, subsequent other sub - processes will continue to execute without being blocked. Since the execution results of different sub - processes have different uses, and considering the issue of parallel execution of the same type of sub - processes (such as the sub - processes corresponding to multi - path retrieval), it is necessary to define reasonable and unique identification information in the data protocol for differentiation. Specifically, refer to Step 3 for data protocol definition. Also, when asynchronously executing multiple sub - processes according to the query information, when any sub - process is executed, it obtains the required dependent data from the asynchronous message queue. This acquisition is asynchronous, which means that the sub - process does not need to block and wait for the data to arrive, but can be notified when the data is ready or actively check the queue. This mechanism supports asynchronous updates because it allows data to be prepared and transmitted in the background.

[0155] Step 2, Data consumption. At the data consumption level of the asynchronous message queue, a message processing mechanism based on the event loop (EventLoop) is adopted to continuously obtain messages from the asynchronous message queue until the flag bit for aborting the task is returned. The flag bit for aborting the task is generated after the last sub-process is completed and is encapsulated into the message together with the execution result corresponding to the last sub-process. Queue.get_nowait() is a method to non-blockingly obtain a message from the asynchronous message queue. When this method is called, it attempts to immediately obtain a message from the asynchronous message queue. If there is a message available in the asynchronous message queue, it will return the message; if the asynchronous message queue is empty, it will raise a `queue.Empty` exception. Queue.task_done() is used to notify the asynchronous message queue that a message previously obtained from Queue.get_nowait() has been completed. Whenever a message acquisition task is completed, this method should be called to reduce the counter of uncompleted tasks in the asynchronous message queue. Then, it is judged whether there is a flag bit for aborting the task in the obtained message. If there is a flag bit for aborting the task, stop obtaining messages from the asynchronous message queue; if there is no flag bit for aborting the task, continue to obtain messages from the asynchronous message queue. After obtaining the message data, the messages are dynamically aggregated or encapsulated to generate structured data according to the policy differences in different stages. Among them, dynamically aggregating messages according to the policy differences in different stages is applicable to the case where sub-processes of the same type are executed in parallel. Exemplarily, there may be multiple document recall sub-processes executing simultaneously. After each document recall sub-process is completed, there will be a corresponding document recall result. The multiple document recall results are aggregated to be used in the subsequent sub-process execution process.

[0156] In the streaming output mode, data is sent to the user via the HTTP Server-Sent Events or WebSocket protocol to answer the user's questions. Here, a unique identifier, Last-Event-ID, is generated for each piece of data. In case of network anomalies on the client side, an error recovery mechanism (resume from breakpoint) is provided. By detecting whether the Last-Event-ID is included in the client request, if it exists, data is sent starting from the specified ID (Identification). In the non-streaming output mode, the data is aggregated and output to the user via the HTTP Post protocol to answer the user's questions. Among them, eventsourceresponse is a response object used to implement Server-Sent Events (SSE). It is used to achieve unidirectional real-time data flow over the HTTP protocol, pushing data from the server to the client. eventsourceresponse allows the server to push data to the client immediately when the data is available, without the client polling the server.

[0157] Step 3: Data protocol definition. For asynchronous output, a standardized message protocol needs to be defined, including message type, specific content, other auxiliary information, etc. Table 1 shows the data protocol definition:

[0158] Table 1

[0159]

[0160] Table 2 shows the format definition of Message:

[0161] Table 2

[0162]

[0163] Table 3 shows the format definition of MessageFeatures:

[0164] Table 3

[0165]

[0166] Table 4 shows the format definition of DocCitation:

[0167] Table 4

[0168]

[0169] The human-computer interaction method in this application can significantly reduce the user waiting time. The first token (text) time is shortened from the original 6 - 7 seconds to within 2 seconds, reducing the end-to-end latency time, enhancing the user's product experience, distributing the computing and data transmission over a longer time period instead of instantaneously processing a large number of data requests, being able to better manage resources and loads, alleviating the system resource and load pressure, and increasing the number of concurrent requests that can be processed by more than 20%.

[0170] Figure 8 The flowchart of the human-computer interaction method based on a document knowledge base provided for an exemplary embodiment of this application. The execution subject of this embodiment is the server in the aforementioned system architecture. As Figure 8 shown, the specific steps of this method are as follows:

[0171] Step S801: Receive the input query information.

[0172] Step S802: According to the query information and the document knowledge base, execute multiple sub-processes of the retrieval enhancement generation process and asynchronously output the execution results of the multiple sub-processes; wherein, the multiple sub-processes include multi-round rewriting, knowledge word recall, knowledge rewriting, document retrieval, document ranking, and generating a reply.

[0173] Optionally, the step of according to the query information and the document knowledge base, executing multiple sub-processes of the retrieval enhancement generation process and asynchronously outputting the execution results of the multiple sub-processes includes:

[0174] According to the query information and the document knowledge base, execute multiple sub-processes of the retrieval enhancement generation process, encapsulate the execution results of each sub-process into a message, and insert the message into an asynchronous message queue; wherein, the message includes an output flag, and the output flag indicates whether the execution result needs to be output;

[0175] Through an asynchronous output thread, obtain the message from the asynchronous message queue. If the output flag in the message indicates that output is required, output the execution result in the message.

[0176] Optionally, the step of according to the query information and the document knowledge base, executing multiple sub-processes of the retrieval enhancement generation process includes:

[0177] Asynchronously execute the multiple sub-processes according to the query information and the document knowledge base;

[0178] Wherein, when executing any one of the sub-processes, obtain the dependency data of the current sub-process from the asynchronous message queue, execute the current sub-process based on the dependency data, and insert the execution result of the current sub-process into the asynchronous message queue;

[0179] The dependent data of the current subprocess includes: the execution results of at least one subprocess executed before the current subprocess.

[0180] For the implementation principle and technical effects of this embodiment, please refer to the relevant content of the foregoing embodiment, which will not be elaborated here.

[0181] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0182] Figure 9 is a structural block diagram of a computing device according to an embodiment of the present application. As Figure 9 shown, the computing device 900 may include: one or more (only one is shown in the figure) processors 901 and a memory 902. Among them, the memory 902 is used to store computer programs / instructions, and the processor 901 is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor 901, the technical solutions provided in any of the foregoing method embodiments are implemented. Their specific functions and achievable technical effects are similar and will not be elaborated here.

[0183] The above computing device can be understood as an integrated intelligent terminal, including but not limited to servers, desktop computers, PCs (Personal Computers), all-in-one models, mobile phones, tablet computers or other portable intelligent terminals, etc. And the computing device may be pre-installed with the models in the foregoing embodiments of the present application.

[0184] Specifically, the computing device can pre - set multiple types of models, including but not limited to models in the fields of natural language processing, visual processing, speech processing, code processing, multi - modal task processing, etc., so as to provide diverse model selection. In different product forms, the computing device can support one or more model usage methods, including but not limited to model training, model invocation, model fine - tuning, model deployment, model inference and application, etc. In some product forms, the computing device also supports model management, including but not limited to multi - type model management (supporting the management of multiple types of models such as discriminative and generative models), model version control (supporting the control of different model versions), model evaluation (evaluating the performance and effectiveness of the model based on model evaluation tools), etc. In other product forms, the computing device can also create applications based on models, provide API (Application Programming Interface) invocation capabilities, and can call the model into the created application through the API interface, while providing application management tools to achieve the management and monitoring of applications.

[0185] Furthermore, the computing device can also include data management (supporting the creation and management of model - tuning data sets), a training center (providing rich training resources to help users learn and master AI technologies), and basic control capabilities (providing enterprise - level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, a comprehensive and integrated AI development, training, deployment, and application device is provided.

[0186] Figure 10 This is a schematic structural diagram of a server provided by an embodiment of the present application. As Figure 10 shown, the server includes: a memory 1001 and a processor 1002. The memory 1001 is used to store computer - executable instructions and can be configured to store various other data to support operations on the server. The processor 1002 is communicatively connected to the memory 1001 and is used to execute the computer - executable instructions stored in the memory 1001 to implement the technical solutions provided by any of the above - mentioned method embodiments. Its specific functions and achievable technical effects are similar and will not be elaborated here.

[0187] Optionally, as Figure 10 shown, the server further includes: other components such as a firewall 1003, a load balancer 1004, a communication component 1005, and a power supply component 1006. Figure 10 Only some components are schematically shown, and it does not mean that the server only includes Figure 10 the components shown. Figure 10 Here, only the cloud server deployed in the cloud is taken as an example for illustrative purposes. The server can also be deployed locally, and no specific limitation is made here in this embodiment.

[0188] An embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the method described in any of the foregoing embodiments is implemented. The specific functions and achievable technical effects are not elaborated herein.

[0189] An embodiment of the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in any of the foregoing embodiments is implemented. The computer program is stored in a readable storage medium. At least one processor of the server can read the computer program from the readable storage medium. Executing the computer program by at least one processor enables the server to execute the technical solution provided in any of the foregoing method embodiments. The specific functions and achievable technical effects are not elaborated herein.

[0190] An embodiment of the present application provides a chip, including: a processing module and a communication interface. The processing module can execute the technical solution of the server in the foregoing method embodiment. Optionally, the chip further includes a storage module (such as a memory). The storage module is used to store instructions, and the processing module is used to execute the instructions stored in the storage module. Executing the instructions stored in the storage module enables the processing module to execute the technical solution provided in any of the foregoing method embodiments.

[0191] The above-mentioned integrated module implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above-mentioned software functional module is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in the various embodiments of the present application.

[0192] It should be understood that the above-mentioned processor may be a central processing unit (CPU for short), a graphics processing unit (GPU for short), or other general-purpose processors, a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by the execution of the hardware processor, or can be implemented by the combination of the hardware and software modules in at least one processor.

[0193] The memory may include high-speed random access memory (Random Access Memory, RAM), and may also include non-volatile storage, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.

[0194] The above-mentioned memory may be an object storage (Object Storage Service, OSS). The above-mentioned memory may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read Only Memory, EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, EPROM), programmable read-only memory (Programmable Read Only Memory, PROM), read-only memory (Read Only Memory, ROM), magnetic memory, flash memory, magnetic disk or optical disc.

[0195] The above-mentioned communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on a communication standard, such as a mobile hotspot (WiFi), a second-generation mobile communication system (2G), a third-generation mobile communication system (3G), a fourth-generation mobile communication system (4G) / Long Term Evolution (LTE), a fifth-generation mobile communication system (5G), etc., or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a near field communication (Near Field Communication, NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (Radio Frequency Identification, RFID) technology, infrared technology, ultra-wideband (Ultra Wide Band, UWB) technology, Bluetooth technology, and other technologies.

[0196] The above-mentioned power supply component provides power for various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.

[0197] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0198] An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and insert information into the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application-specific integrated circuit. Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a master device.

[0199] It should be noted that in this document, the terms "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including such element.

[0200] The order of the above embodiments of the present application is for description only and does not represent the superiority or inferiority of the embodiments. Additionally, in some of the processes described in the above embodiments and the accompanying drawings, there are multiple operations that occur in a specific order. However, it should be clearly understood that these operations can be performed not in the order in which they appear in this document or in parallel, and are only used to distinguish different operations. The serial numbers themselves do not represent any order of execution. Additionally, these processes can include more or fewer operations, and these operations can be performed in sequence or in parallel. It should be noted that the descriptions such as "first", "second", etc. in this document are used to distinguish different messages, devices, modules, etc., do not represent a sequence, and do not limit that "first" and "second" are of different types. The meaning of "a plurality" is two or more, unless otherwise specifically and clearly defined.

[0201] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0202] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.

[0203] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A human-computer interaction method, characterized in that: include: Receive input query information; Executing multiple sub-processes of the retrieval enhancement generation process according to the query information, and asynchronously outputting the execution results of the multiple sub-processes; The execution result of the last sub-process of the retrieval enhancement generation process is the reply information of the query information.

2. The method according to claim 1, characterized in that The step of executing multiple sub-processes of the retrieval enhancement generation process according to the query information and asynchronously outputting the execution results of the multiple sub-processes includes: Execute multiple sub-processes of the retrieval enhancement generation process according to the query information, encapsulate the execution results of each sub-process into a message, and insert the message into an asynchronous message queue; wherein the message includes an output tag, and the output tag indicates whether the execution result needs to be output; The message is obtained from the asynchronous message queue through an asynchronous output thread, and if the output mark in the message indicates that output is required, the execution result in the message is output.

3. The method according to claim 2, characterized in that The step of encapsulating the execution results of each sub-process into a message and inserting the message into an asynchronous message queue comprises: Configure the first subprocess that needs to output the execution result; Encapsulating the execution result of the first sub-process into a first message, and inserting the first message into the asynchronous message queue, wherein the first message includes a first output tag, and the first output tag indicates that the execution result needs to be output; The execution result of the second sub-process is encapsulated into a second message, and the second message is inserted into the asynchronous message queue, wherein the second message includes a second output tag, the second output tag indicates that the execution result does not need to be output, and the second sub-process refers to a sub-process other than the first sub-process.

4. The method according to claim 2, characterized in that: The step of encapsulating the execution results of each sub-process into a message and inserting the message into an asynchronous message queue comprises: For any of the sub-processes, if the sub-process is executed through the large model, the result fragments outputted by the large model each time are encapsulated into the message, and the message is inserted into the asynchronous message queue.

5. The method according to claim 4, characterized in that The step of encapsulating the result fragment outputted by the large model each time into a message comprises: The result fragments that are not the last output of the large model are encapsulated into messages containing a generating mark, and the result fragments that are the last output of the large model are encapsulated into messages containing an end mark.

6. The method according to claim 2, characterized in that The message also includes: information about the sub-process executed, When outputting the execution result in the message, information of the executed sub-process in the message is output.

7. The method according to any one of claims 1 to 6, characterized in that The multiple sub-processes of performing the retrieval enhancement generation process according to the query information include: asynchronously executing the plurality of sub-processes according to the query information; Wherein, when executing any of the sub-processes, the dependent data of the current sub-process is obtained from the asynchronous message queue, the current sub-process is executed based on the dependent data, and the execution result of the current sub-process is inserted into the asynchronous message queue; The dependency data of the current sub-process includes: an execution result of at least one sub-process executed before the current sub-process.

8. The method according to claim 7, characterized in that The obtaining the dependency data of the current sub-process from the asynchronous message queue and executing the current sub-process based on the dependency data includes: Obtaining a message of the target subprocess on which the process depends from the asynchronous message queue, and obtaining an execution result of the target subprocess from the message; If the message includes a generating mark, the obtained execution result fragment of the target sub-process is stored, and the message of the target sub-process is continuously obtained from the asynchronous message queue until a message including a generating end mark is obtained, and then the execution result fragment of the target sub-process in the message including the generating end mark is obtained; Integrate the acquired execution result fragments of the target sub-process to obtain the complete execution result of the target sub-process; The current sub-process is executed according to the complete execution result of the target sub-process.

9. A human-computer interaction method based on a document knowledge base, characterized in that: include: Receive input query information; According to the query information and the document knowledge base, multiple sub-processes of the retrieval enhancement generation process are executed, and the execution results of the multiple sub-processes are output asynchronously; wherein the multiple sub-processes include multiple rounds of rewriting, knowledge word recall, knowledge rewriting, document retrieval, document sorting and generating replies.

10. The method according to claim 9, characterized in that The step of executing multiple sub-processes of the retrieval enhancement generation process according to the query information and the document knowledge base, and asynchronously outputting the execution results of the multiple sub-processes, includes: According to the query information and the document knowledge base, multiple sub-processes of the retrieval enhancement generation process are executed, and the execution results of each sub-process are encapsulated into a message, and the message is inserted into an asynchronous message queue; wherein the message includes an output tag, and the output tag indicates whether the execution result needs to be output; The message is obtained from the asynchronous message queue through an asynchronous output thread, and if the output mark in the message indicates that output is required, the execution result in the message is output.

11. The method according to claim 9 or 10, characterized in that: The multiple sub-processes of performing the retrieval enhancement generation process according to the query information and the document knowledge base include: asynchronously executing the plurality of sub-processes according to the query information and the document knowledge base; Wherein, when executing any of the sub-processes, the dependent data of the current sub-process is obtained from the asynchronous message queue, the current sub-process is executed based on the dependent data, and the execution result of the current sub-process is inserted into the asynchronous message queue; The dependency data of the current sub-process includes: an execution result of at least one sub-process executed before the current sub-process.

12. A computing device, characterized in that: include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the method according to any one of claims 1 to 11 is implemented.

13. A server, characterized in that: include: at least one processor; as well as a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the server to execute the method described in any one of claims 1-11.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the method according to any one of claims 1 to 11 is implemented.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Intestinal microorganism intelligent question-answering system based on knowledge graph and large model

    CN119226487A

  • Retrieval enhancement generation method and device, equipment and storage medium

    CN119903079A

  • A system for generating answers to multiple questions using rag-based generative artificial intelligence technology

    KR102710159B1

  • Large language model-based information retrieval for large datasets

    US20250086215A1

Cited By

  • Streaming output method of model and electronic equipment

    CN121050892A