Streaming return method and device for knowledge questions and answers and computer equipment

By gradually generating and retrieving incremental reply content and document information in the knowledge Q&A system, the problem of delayed traceability response in the existing technology is solved, and the parallel processing of question and answer content and traceability information is realized, improving the system's interaction efficiency and user experience.

CN120258154AActive Publication Date: 2025-07-04SHENZHEN LANLING SOFTWARE CO LTD

Patent Information

Application Number
CN202510733847.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

In the existing knowledge Q&A system, in long text generation or interactive Q&A, there is a significant delay in the traceability response, which makes it difficult for users to judge the reliability and accuracy of the Q&A content in a timely manner, limiting the system's interaction efficiency and traceability feedback capabilities.

Method used

By inputting the user input text into the large language model, incremental reply content is gradually generated, and matching target document information is retrieved in real time in the pre-stored knowledge base, and then sent to the client in a streaming return mode, realizing parallel processing of question-and-answer content and traceability information.

Benefits of technology

It significantly improves the interpretability of Q&A results and document positioning efficiency, shortens the delay time for users to obtain content traceability information, improves the system's interaction continuity and information transparency, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258154A_ABST
    Figure CN120258154A_ABST
Patent Text Reader

Abstract

The invention relates to a streaming return method and device for knowledge questions and answers and computer equipment. The method comprises the following steps: inputting a user input text received from a client into a large language model to obtain incremental reply content for gradually returning to the client; after the incremental reply content needed each time is obtained, based on the incremental reply content needed this time, retrieval is conducted in a pre-stored knowledge base, and target document information matched with the incremental reply content needed this time is obtained; and combining the incremental reply content required at this time and the target document information to obtain streaming return content returned at this time for the text input by the user, and sending the streaming return content returned at this time to the client in a streaming return mode. By adopting the method, the traceability instantaneity of knowledge questions and answers can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a method, device, computer device, computer-readable storage medium, and computer program product for streaming return of knowledge Q&A. Background Art

[0002] Although some existing knowledge Q&A systems have the ability to return content in a streaming manner, their content traceability often depends on unified execution after the completion of the generation of the complete Q&A content. This post-processing mechanism results in a significant delay in the traceability response. Users usually need to wait until the entire answer is generated before they can obtain the relevant document source information, making it difficult to judge the reliability and accuracy of the Q&A content in a timely manner. Especially in application scenarios with high requirements for response speed and interpretability, such as long text generation or interactive Q&A, this mechanism severely limits the interaction efficiency and traceability feedback ability of the system, causing a delay in the association between the Q&A content and its source, and ultimately resulting in poor real-time traceability of the system. Summary of the Invention

[0003] Based on this, in view of the above technical problems, it is necessary to provide a method, device, computer device, computer-readable storage medium, and computer program product for streaming return of knowledge Q&A that can improve the real-time traceability of knowledge Q&A.

[0004] In a first aspect, this application provides a method for streaming return of knowledge Q&A, including:

[0005] Input the user input text received from the client into a large language model to obtain incremental reply content for gradually returning to the client;

[0006] After obtaining the incremental reply content required each time, based on the incremental reply content required this time, retrieve in a pre-stored knowledge base to obtain target document information that matches the incremental reply content required this time;

[0007] Merge the incremental reply content required this time and the target document information to obtain the streaming return content returned this time for the user input text, and send the streaming return content returned this time to the client in a streaming return manner.

[0008] In one of the embodiments, the step of retrieving in a pre-stored knowledge base based on the incremental reply content required this time to obtain target document information that matches the incremental reply content required this time includes:

[0009] Input the incremental reply content required this time into a semantic parsing model to obtain the semantic vector corresponding to the incremental reply content required this time;

[0010] Perform a similarity match between the semantic vector and the semantic vectors of the pre-stored document information in the pre-stored knowledge base to determine the target document information with the highest similarity.

[0011] In one embodiment, the step of inputting the user input text received from the client into the large language model to obtain incremental reply content for gradually returning to the client includes:

[0012] Input the user input text received from the client into the large language model to obtain the reply content continuously output by the large language model in a streaming manner;

[0013] Based on the continuously output reply content, determine the incremental reply content for gradually returning to the client.

[0014] In one embodiment, the step of determining the incremental reply content for gradually returning to the client based on the continuously output reply content includes:

[0015] Continuously count the number of tokens corresponding to the continuously output reply content;

[0016] When the number of tokens is greater than a preset number threshold, start detecting and determining the latest text end identifier in the continuously output reply content;

[0017] Determine the reply content ending with the latest text end identifier as the required incremental reply content for one time, and clear the number of tokens.

[0018] In one embodiment, after obtaining the required incremental reply content each time, based on the required incremental reply content for this time, retrieving in the pre-stored knowledge base to obtain the target document information matching the required incremental reply content for this time, further includes:

[0019] Store the target document information in a preset document cache;

[0020] The step of retrieving in the pre-stored knowledge base based on the required incremental reply content for this time to obtain the target document information matching the required incremental reply content for this time includes:

[0021] First, obtain the target document information matching the required incremental reply content for this time from the preset document cache;

[0022] In the case that there is no target document information matching the required incremental reply content for this time in the preset document cache, then obtain the target document information matching the required incremental reply content for this time from the pre-stored knowledge base.

[0023] In one embodiment, the method further includes:

[0024] When all the streaming return contents for the user input text have been sent to the client, clear the preset document cache.

[0025] In a second aspect, the present application further provides a streaming return device for knowledge Q&A, including:

[0026] A large model application module, configured to input the user input text received from the client into a large language model to obtain incremental reply contents for gradually returning to the client;

[0027] A document matching module, configured to, after obtaining each required incremental reply content, retrieve in a pre-stored knowledge base based on the current required incremental reply content to obtain target document information matching the current required incremental reply content;

[0028] A streaming return module, configured to merge the current required incremental reply content and the target document information to obtain the streaming return content for the current return for the user input text, and send the current returned streaming return content to the client in a streaming return manner.

[0029] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0030] Input the user input text received from the client into a large language model to obtain incremental reply contents for gradually returning to the client;

[0031] After obtaining each required incremental reply content, retrieve in a pre-stored knowledge base based on the current required incremental reply content to obtain target document information matching the current required incremental reply content;

[0032] Merge the current required incremental reply content and the target document information to obtain the streaming return content for the current return for the user input text, and send the current returned streaming return content to the client in a streaming return manner.

[0033] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0034] Input the user input text received from the client into a large language model to obtain incremental reply contents for gradually returning to the client;

[0035] After obtaining the incremental response content required each time, based on the incremental response content required this time, retrieve in the pre-stored knowledge base to obtain the target document information that matches the incremental response content required this time;

[0036] Merge the incremental response content required this time and the target document information to obtain the streaming return content returned this time for the user input text, and send the streaming return content returned this time to the client in a streaming return manner.

[0037] In a fifth aspect, the present application also provides a computer program product, including a computer program, which when executed by a processor implements the following steps:

[0038] Input the user input text received from the client into a large language model to obtain incremental response content for gradually returning to the client;

[0039] After obtaining the incremental response content required each time, based on the incremental response content required this time, retrieve in the pre-stored knowledge base to obtain the target document information that matches the incremental response content required this time;

[0040] Merge the incremental response content required this time and the target document information to obtain the streaming return content returned this time for the user input text, and send the streaming return content returned this time to the client in a streaming return manner.

[0041] The above-mentioned streaming return method, device, computer device, computer-readable storage medium, and computer program product for knowledge Q&A. First, the user input text received from the client is input into a large language model to obtain incremental reply content for gradually returning to the client. Through the large language model, it is possible to quickly parse and process the user's natural language questions, and based on the language understanding ability of the large language model, gradually generate corresponding Q&A content, providing a basis for the subsequent real-time splitting, associated retrieval, and streaming output, helping to decouple Q&A generation from related information processing and improving the overall response speed of the system. Then, after obtaining the incremental reply content required each time, based on the incremental reply content required this time, a search is performed in the pre-stored knowledge base to obtain target document information that matches the incremental reply content required this time, enabling fast semantic matching of the currently generated content and obtaining relevant document basis for this part of the content in real time, which can significantly shorten the latency of the user to obtain content traceability information and effectively improve the interpretability of the Q&A results and the document location efficiency. Finally, the incremental reply content required this time and the target document information are merged to obtain the streaming return content returned this time for the user input text, and the streaming return content returned this time is sent to the client in a streaming return manner, which can synchronously push the Q&A results and their source information, enabling the user to understand the source basis of the corresponding content in real time during the process of receiving the answer, improving the interaction continuity and information transparency of the system, and enhancing the user's actual usage experience. In the above method, by dynamically performing traceability retrieval during the Q&A generation process and synchronously returning the incremental Q&A content and its corresponding document information, parallel processing and real-time output of the Q&A content and traceability information are achieved, significantly improving the response speed of the traceability information, overcoming the latency problem caused by the need to wait for unified traceability after complete generation in the traditional method, and optimizing the content feedback efficiency of the system and the user interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained without creative efforts.

[0043] Figure 1 It is a flowchart of the streaming return method for knowledge Q&A in an embodiment;

[0044] Figure 2 It is a flowchart of the determination steps of incremental reply content in an embodiment;

[0045] Figure 3Schematic flowchart of the streaming return method for knowledge Q&A in another embodiment;

[0046] Figure 4 Interaction flowchart of the streaming return method for knowledge Q&A in one embodiment;

[0047] Figure 5 Block diagram of the structure of the streaming return device for knowledge Q&A in one embodiment;

[0048] Figure 6 Internal structure diagram of a computer device in one embodiment. Detailed implementation manners

[0049] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0050] In one embodiment, as Figure 1 shown, a streaming return method for knowledge Q&A is provided. In this embodiment, it is exemplified that this method is applied to the server side. It can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, etc. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. In this embodiment, the method includes the following steps:

[0051] Step S101: Input the user input text received from the client into a large language model to obtain incremental reply content for gradually returning to the client.

[0052] Among them, the user input text refers to questions, instructions, or other information requests input by the user in natural language at the client, and this text constitutes the basic corpus for the large language model to understand and generate. The incremental reply content refers to multiple partial reply results sequentially, stage by stage, and gradually generated by the large language model during the reasoning process, and each part of the content can be independently transmitted and displayed.

[0053] Exemplarily, the server side can receive the natural language question input by the user through the interface module, encapsulate the question in a request format, and then pass it to the large language model server side; the large language model performs inference calculations on the text and outputs the reply content in a streaming manner. The server side continuously listens for its return results and divides the received content into multiple incremental reply contents according to the set segmentation mechanism (for example, detecting that the number of tokens reaches the threshold or the appearance of a semantic end symbol). The server side can cache, package, or structure these incremental reply contents to prepare for subsequent matching and merging with the traceability information.

[0054] Step S102, after obtaining the incremental reply content required each time, based on the incremental reply content required this time, retrieve in the pre-stored knowledge base to obtain the target document information that matches the incremental reply content required this time.

[0055] Among them, the pre-stored knowledge base refers to a collection of structured or unstructured documents pre-constructed and stored by the system, and this collection includes but is not limited to text information such as Internet online content, knowledge graphs, regulatory documents, and technical materials; the target document information refers to one or more paragraphs of document content with reference value determined by semantic relevance matching with the current incremental reply content, and usually includes the document title, the body paragraph, and its adjacent context.

[0056] Exemplarily, after the server side obtains a certain incremental reply content, it first converts the content into a semantic vector representation through an embedding model, and then compares the semantic vector with the document vectors in the pre-stored knowledge base (for example, using algorithms such as cosine similarity and Euclidean distance); the server side can select one or more document paragraphs with the highest similarity as the target document information of the current incremental reply content according to the comparison result. In addition, to improve the user's understanding efficiency, the server side can further extract the title information of the document paragraph and the adjacent context paragraphs to form a complete reference traceability content.

[0057] Step S103, merge the incremental reply content required this time and the target document information to obtain the streaming return content returned this time for the user input text, and send the streaming return content returned this time to the client in a streaming return manner.

[0058] Among them, the streaming return content refers to the output formed by structurally combining the incremental reply content with the corresponding target document information, which can be gradually pushed to the client for display. This content is usually returned in the form of data packets or data segments and has the characteristic of being generated and sent while in progress. The streaming return manner means that the server side, without waiting for the entire Q&A process to complete, pushes the currently available partial results to the client in sequence in real time, thereby improving the response speed and interaction coherence.

[0059] Exemplarily, after the server side completes the association matching between the incremental reply content and the target document information, it encapsulates the two into a structured data object, for example, constructs a data body containing a "reply content field" and a "traceability information field" in JSON format. Subsequently, the server side pushes this data body as a streaming message to the client based on a long connection protocol (such as WebSocket, HTTP / 2, or other protocols that support streaming communication). After receiving this content, the client can immediately display the current Q&A result and its document basis, while the server side continues to listen for the generation of the next round of incremental content and repeats the above process to achieve continuous and traceable knowledge Q&A interaction.

[0060] In the above method for streaming return of knowledge Q&A, first, the user input text received from the client is input into a large language model to obtain incremental reply content for gradually returning to the client. Through the large language model, it is possible to quickly parse and process the user's natural language questions, and based on the language understanding ability of the large language model, gradually generate corresponding Q&A content, providing a basis for the subsequent real-time splitting, associated retrieval, and streaming output of content, which helps to decouple Q&A generation from related information processing and improve the overall response speed of the system. Then, after obtaining the incremental reply content required each time, based on the incremental reply content required this time, a search is performed in the pre-stored knowledge base to obtain target document information that matches the incremental reply content required this time, which can achieve fast semantic matching of the currently generated content and obtain the document basis related to this part of the content in real time, significantly shortening the latency time for the user to obtain content traceability information and effectively improving the interpretability of the Q&A result and the document positioning efficiency. Finally, the incremental reply content required this time and the target document information are merged to obtain the streaming return content returned this time for the user input text, and the streaming return content returned this time is sent to the client in a streaming return manner, which can achieve the synchronous push of the Q&A result and its source information, enabling the user to understand the source basis of the corresponding content in real time during the process of receiving the answer, improving the interaction continuity and information transparency of the system, and enhancing the user's actual usage experience. In the above method, by dynamically performing traceability retrieval during the Q&A generation process and synchronously returning the incremental Q&A content and its corresponding document information, parallel processing and real-time output of the Q&A content and traceability information are achieved, significantly improving the response speed of the traceability information, overcoming the latency problem caused by the need to wait for unified traceability after complete generation in the traditional method, and optimizing the content feedback efficiency of the system and the user interaction experience.

[0061] In an exemplary embodiment, the above step S102 retrieves in a pre-stored knowledge base based on the incremental reply content required this time to obtain target document information that matches the incremental reply content required this time, and further includes: inputting the incremental reply content required this time into a semantic parsing model to obtain a semantic vector corresponding to the incremental reply content required this time; performing a similarity matching between the semantic vector and the semantic vectors of the pre-stored document information in the pre-stored knowledge base to determine the target document information with the highest similarity.

[0062] Among them, the semantic parsing model refers to a deep learning model used to convert natural language text into low-dimensional or high-dimensional semantic vectors, such as BERT (Bidirectional Encoder Representation from Transformers), Sentence-BERT (Sentence Bidirectional Encoder Representations from Transformers), word2vec (Word to Vector), etc., which can realize information expression and alignment at the semantic level. The semantic vector refers to a numerical vector representation that encodes the input text and can be used for semantic similarity comparison in a vector space, and this vector representation can be used for subsequent matching and retrieval operations.

[0063] Exemplarily, the server first inputs the incremental reply content required currently into the semantic parsing model to obtain the semantic vector corresponding to this content; then, performs a similarity matching between this semantic vector and the semantic vectors corresponding to each document paragraph in the pre-stored knowledge base, and the similarity matching can be implemented by means such as cosine similarity, Euclidean distance, or Mahalanobis distance; then sorts the similarity scores, selects one or more document paragraphs with the highest scores therefrom, and determines them as the target document information that matches the incremental reply content this time. To improve the integrity of traceability and the context understanding ability, the server can further extract the title information of the document where this document paragraph is located and the adjacent paragraphs before and after this paragraph as auxiliary reference content, and include them in the target document information and return them together.

[0064] In this embodiment, by converting the incremental reply content into a semantic vector and performing a similarity matching with the semantic vectors of each document information in the pre-stored knowledge base, efficient and accurate semantic association retrieval can be realized during the question-and-answer generation process, quickly locate the document paragraph most relevant to the current question-and-answer content, significantly improve the response speed and pertinence of the traceability information, realize the parallel processing of content generation and content traceability, and optimize the feedback efficiency of the question-and-answer system and the interaction experience of users.

[0065] In an exemplary embodiment, as Figure 2As shown, the above step S101 inputs the user input text received from the client into the large language model to obtain incremental reply content for gradually returning to the client, and can also be implemented through the following steps:

[0066] Step S201: Input the user input text received from the client into the large language model to obtain the reply content continuously output by the large language model in a streaming return manner;

[0067] Step S202: Based on the continuously output reply content, determine the incremental reply content for gradually returning to the client.

[0068] Among them, the streaming return manner means that the large language model can continuously output content fragments during the generation process, rather than returning all at once after the whole is generated, and is usually implemented through interfaces that support streaming transmission (such as WebSocket, HTTP / 2, etc.). The reply content refers to the natural language output result generated by the large language model in real time according to the user input text, which can be represented as a series of tokens or statements output in sequence. The incremental reply content refers to the content blocks obtained by dividing the continuously output reply content according to the set rules and can be used for staged return.

[0069] Exemplarily, the server-side obtains the token stream continuously returned during its inference process by establishing a continuously connected call method with the large language model. During the process of receiving tokens, the server-side divides this streaming content in real time according to the preset content segmentation rules (for example, reaching a fixed number of tokens, or detecting semantic end markers such as full stops, question marks, etc.), so as to determine the incremental reply content to be returned each time. After the server-side determines the current incremental reply content, it enters the subsequent traceability retrieval and return stage, forming a complete question-answer - traceability linkage processing flow.

[0070] In this embodiment, by performing real-time monitoring and segmentation on the streaming output content of the large language model, it is possible to dynamically process the reply content during the generation process, effectively improving the response speed and interaction continuity of the question-answer system, providing a content basis with reasonable granularity for subsequent traceability operations, and enhancing the streaming processing ability and user experience of the overall system.

[0071] In an exemplary embodiment, the above step S202 based on the continuously output reply content to determine the incremental reply content for gradually returning to the client further includes: continuously counting the number of tokens corresponding to the continuously output reply content; when the number of tokens is greater than the preset number threshold, start detecting and determining the latest text end marker in the continuously output reply content; determine the reply content with the latest text end marker as the content end as an incremental reply content required once, and clear the number of tokens to zero.

[0072] Among them, a token refers to the smallest semantic unit output by a large language model during text generation, which can be a character, a word, a symbol, or a combination thereof, depending on the model tokenization mechanism adopted, such as BPE encoding (Byte Pair Encoding). The end-of-text marker refers to the symbol used to mark the end of a natural language sentence, including Chinese full stop (。), English full stop (.), question mark (?,?), exclamation mark (!,!), etc., and is often used to judge the semantic boundary.

[0073] Exemplarily, the server side can set a fixed token number threshold as a reference for the segmentation condition, and continuously record the number of received tokens during the process of receiving the output of the large language model. When the cumulative number of tokens reaches or exceeds the threshold, the server side will search forward in the currently received response content for the most recent occurrence of the end-of-text marker to delimit the boundary of the complete semantics. If the end-of-text marker is detected, the content before it will be used as an incremental response content; if the end-of-text marker is not found in the current content, the server side can also suspend the segmentation according to the policy or perform forced segmentation using a conservative strategy. After each division of the incremental response content is completed, the server side clears the token counter and continues to listen for the subsequent output content of the large language model, thereby achieving the phased segmentation of the streaming output content and the generation of the basic content for streaming return.

[0074] For example, the server side can set the token number threshold to 80. Suppose the large language model is continuously generating the response content "Newton proposed the classical mechanics system, which has important value in engineering practice. Modern physics has further developed the theory of relativity and quantum mechanics", and the server side has cumulatively received 85 tokens. At this time, the server side triggers the end-of-text marker detection mechanism and searches forward for the most recent occurrence of the end-of-text marker (such as the full stop "。"), and finds that the end of the sentence is after "has important value." The server side will use the previous content "Newton proposed the classical mechanics system, which has important value in engineering practice." ending with this full stop as a complete incremental response content. Subsequently, the token counter is cleared, and the server side continues to listen for the subsequent output content of the large language model and performs the next round of division.

[0075] In this embodiment, by combining the dual judgment strategies of token statistics threshold and semantic end marker, while ensuring the structural integrity of the output content, the length of the incremental block can be controlled, taking into account the response real-time performance and content readability, effectively improving the stability of content segmentation and the user perception quality in the streaming Q&A system.

[0076] In an exemplary embodiment, after obtaining the incremental reply content required each time in step S102 above, based on the incremental reply content required this time, retrieving in the pre-stored knowledge base to obtain the target document information matching the incremental reply content required this time, it further includes: storing the target document information in a preset document cache.

[0077] Further, in an exemplary embodiment, after obtaining the incremental reply content required each time in step S102 above, based on the incremental reply content required this time, retrieving in the pre-stored knowledge base to obtain the target document information matching the incremental reply content required this time, it further includes: preferentially obtaining, from the preset document cache, the target document information matching the incremental reply content required this time; in the case that there is no target document information in the preset document cache that matches the incremental reply content required this time, then obtaining, from the pre-stored knowledge base, the target document information matching the incremental reply content required this time.

[0078] Wherein, the preset document cache refers to a short-term document matching candidate set maintained by the server side in the context of the current Q&A session. The storage method can adopt a structured list or a vector index structure. The stored content includes information such as the titles of the target documents retrieved from the historical incremental content, the document paragraphs, and their contexts. The document information in the cache is not bound to a specific reply content, but serves as a global candidate document pool for subsequent incremental content to perform semantic matching reference.

[0079] Exemplarily, in the same Q&A context, there is usually strong semantic continuity between incremental reply contents, and their corresponding documents often focus on a certain theme. Therefore, reusing the matching results of the previous round as cache candidates can, on the premise of ensuring semantic relevance, speed up the traceability feedback speed, reduce the system load, and improve the overall user experience at the same time.

[0080] The server side can, each time when processing a new incremental reply content, first perform a similarity match between the semantic vector of the incremental content and the semantic vectors of the target documents already stored in the current preset document cache. If the matching score exceeds the preset threshold, it is regarded as hitting the cache, and the document information can be directly returned, skipping the knowledge base retrieval; if there is no document information in the cache that meets the conditions, then the server side accesses the pre-stored knowledge base again to obtain the target document information with the best semantic match. In the latter case, that is, when the server side obtains new target document information from the knowledge base, the document information should be synchronously written into the preset document cache as a reusable candidate in the current session.

[0081] In this way, the server forms a closed-loop process of "preferred cache matching - retrieval when necessary - result caching" throughout the knowledge Q&A session, enabling real-time processing of each incremental content while fully reusing the results of the previous round, greatly reducing the frequency of repeated access to the knowledge base.

[0082] For example, in a knowledge Q&A session, the first incremental reply content received by the server is "Newton proposed the three laws of motion". After semantic vector parsing, a document paragraph titled "Fundamental Principles of Classical Mechanics" is retrieved and hit in the knowledge base. The server stores the document title, the matching paragraph, and its adjacent paragraphs as the target document information in the preset document cache of the current session. Subsequently, the second incremental reply content generated by the large language model is "The first law is also known as the law of inertia". The server first performs semantic similarity matching between this content and the "Fundamental Principles of Classical Mechanics" document in the cache. Since the matching score is higher than the set threshold, the cached document is directly reused as the target document information and returned, skipping the knowledge base retrieval process. At this time, the system not only ensures the accuracy of semantic matching but also improves the processing efficiency, realizing the rapid linkage between content generation and document traceability.

[0083] In this embodiment, by introducing a document cache reuse mechanism based on context relevance, the frequency of repeated access to the knowledge base can be significantly reduced during multiple rounds of generation. By making full use of the semantic continuity of the current Q&A context, the response efficiency of the overall Q&A flow can be improved while ensuring the accuracy of document matching, which is particularly suitable for long text and semantically coherent knowledge Q&A application scenarios.

[0084] In an exemplary embodiment, the above method for streaming returns of knowledge Q&A further includes: when all the streaming return contents for the user input text have been sent to the client, clearing the preset document cache.

[0085] Exemplarily, the server can record the number of completed streaming returns in the current session during the process of receiving the output of the large language model, or listen for termination identifiers indicating the completion of model output, such as the EOS marker (End of Sequence) or the streaming transmission channel closing signal, to determine whether the complete response process for the user input text has ended. When the server confirms that all incremental reply contents and their corresponding target document information have been streamed and sent, the operation of clearing the preset document cache can be triggered. The cache can be cleared immediately or with a delay to ensure that it does not occupy storage space outside the session and prevent the cache content from being misused in non-current sessions. Ensure that the preset document cache is only valid within a single Q&A cycle, thereby improving the timeliness and isolation of the cache content, preventing misjudgment of cache hits caused by context invalidation, and releasing memory resources to provide a clean cache environment for the next round of knowledge Q&A tasks.

[0086] In this embodiment, by actively clearing the document cache at the end of the Q&A session, the cache structure has a clear life cycle boundary, improving the system's resource management ability and the context consistency of the cached content, and effectively ensuring the accuracy and controllability of the Q&A system in the scenario of multi-user and multi-task concurrency.

[0087] In another exemplary embodiment, as Figure 3 shown, the present application provides a method for streaming return of knowledge Q&A, which includes:

[0088] Step S301: Input the user input text received from the client into the large language model to obtain the reply content continuously output by the large language model in a streaming return manner.

[0089] Step S302: Continuously count the number of tokens corresponding to the continuously output reply content.

[0090] Step S303: When the number of tokens is greater than the preset number threshold, start detecting and determining the latest text end flag in the continuously output reply content.

[0091] Step S304: Determine the reply content with the latest text end flag as the content end as an incremental reply content required once, and clear the number of tokens to zero.

[0092] Step S305: After obtaining each incremental reply content required, preferentially obtain the target document information matching the incremental reply content required this time from the preset document cache.

[0093] Step S306: In the case that there is no target document information matching the incremental reply content required this time in the preset document cache, then obtain the target document information matching the incremental reply content required this time from the pre-stored knowledge base, and store the target document information in the preset document cache.

[0094] Step S307: Merge the incremental reply content required this time and the target document information to obtain the streaming return content returned this time for the user input text, and send the streaming return content returned this time to the client in a streaming return manner.

[0095] Step S308: When all the streaming return content for the user input text has been sent to the client, clear the preset document cache.

[0096] Exemplarily, as Figure 4 shown, taking the interaction between the client and the server as an example for illustration:

[0097] (1) Client: Input the question content;

[0098] (2)Server - Q&A Service: It is used to receive the question content input by the client, call the algorithm module for calculation, and return the result to the client;

[0099] (3)Server - Search Service: According to the incoming text information, it returns the content such as several document titles and document details with the most semantic matching.

[0100] The specific process is as follows:

[0101] (1)The client obtains the content input by the user.

[0102] The client uses the text content as input parameters and initiates a request to the server.

[0103] The data is transmitted using the HTTP (HyperText Transfer Protocol), HTTPS (HyperText Transfer Protocol Secure), TCP (Transmission Control Protocol), or UDP (User Datagram Protocol).

[0104] (2)The Q&A service interface receives the data.

[0105] The server parses the client - transmitted message and passes it to the algorithm service.

[0106] (3)Q&A Service - Stream - based Generation of Q&A Content.

[0107] The text received by the server - side interface is passed into the large - language model, and the content is incrementally returned in an asynchronous call manner.

[0108] (4)Search Service - Incremental Content Comparison.

[0109] The incrementally returned content in step (3) is passed into the search service for comparison to obtain the document titles, document paragraphs, and their adjacent paragraphs that are most semantically matching to the input content.

[0110] (5)The Q&A service interface returns the result.

[0111] The server transmits the data through HTTP, HTTPS, TCP, or UDP protocols, and combines the results of the algorithm service and the search service and returns them to the client.

[0112] The result should include: incremental Q&A results, document titles that match the content of the incremental Q&A results, document paragraphs, and their adjacent paragraphs.

[0113] (6) The client receives the returned information.

[0114] The client receives the structured information returned by the server side and displays it in a streaming manner: the incremental Q&A results, the document titles that match the content of the incremental Q&A results, the document paragraphs, and their adjacent paragraphs.

[0115] In this embodiment, compared with the traditional Q&A results, it can quickly respond to user requests. While incrementally returning the Q&A results, it can also quickly provide the document paragraph content that best matches the current content. By returning in parts, it reduces the user's waiting time and allows the user to obtain feedback quickly. In the traditional Q&A method, the user traces back to the document according to the Q&A results and processes it manually. This application can quickly let the user locate the document paragraph that best matches the current Q&A result, and at the same time provide adjacent document paragraphs, reducing the possibility of the user consulting the document through the search engine again and effectively improving the retrieval efficiency.

[0116] It should be understood that although the steps in the flowcharts involved in the above-mentioned embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0117] Based on the same inventive concept, the embodiments of the present application also provide a streaming return device for knowledge Q&A for implementing the above-mentioned streaming return method for knowledge Q&A. The implementation solutions provided by this device to solve problems are similar to those described in the above method. Therefore, the specific limitations in one or more embodiments of the following streaming return devices for knowledge Q&A can refer to the limitations on the streaming return method for knowledge Q&A in the above text, and will not be repeated here.

[0118] In an exemplary embodiment, as Figure 5 shown, a streaming return device for knowledge Q&A is provided, including: a large model application module 501, a document matching module 502, and a streaming return module 503, where:

[0119] The large model application module 501 is configured to input the user input text received from the client into the large language model to obtain incremental reply content for gradually returning to the client;

[0120] The document matching module 502 is configured to, after obtaining the incremental reply content required each time, retrieve in the pre-stored knowledge base based on the incremental reply content required this time, and obtain the target document information that matches the incremental reply content required this time;

[0121] The streaming return module 503 is configured to merge the incremental reply content required this time and the target document information, obtain the streaming return content returned this time for the user input text, and send the streaming return content returned this time to the client in a streaming return manner.

[0122] In one embodiment, the above-mentioned document matching module 502 is further configured to input the incremental reply content required this time into a semantic parsing model to obtain the semantic vector corresponding to the incremental reply content required this time; perform a similarity match between the semantic vector and the semantic vectors of the pre-stored document information in the pre-stored knowledge base to determine the target document information with the highest similarity.

[0123] In one embodiment, the above-mentioned large model application module 501 is further configured to input the user input text received from the client into a large language model to obtain the reply content continuously output by the large language model in a streaming return manner; determine the incremental reply content for gradually returning to the client based on the continuously output reply content.

[0124] In one embodiment, the above-mentioned large model application module 501 is further configured to continuously count the number of tokens corresponding to the continuously output reply content; when the number of tokens is greater than a preset number threshold, start detecting and determining the latest text end identifier in the continuously output reply content; determine the reply content with the latest text end identifier as the content end as the incremental reply content required once, and clear the number of tokens to zero.

[0125] In one embodiment, the above-mentioned streaming return device for knowledge Q&A further includes a document caching module for storing the target document information in a preset document cache.

[0126] In one embodiment, the above-mentioned document matching module 502 is further configured to preferentially obtain the target document information that matches the incremental reply content required this time from the preset document cache; in the case that there is no target document information in the preset document cache that matches the incremental reply content required this time, then obtain the target document information that matches the incremental reply content required this time from the pre-stored knowledge base.

[0127] In one embodiment, the above-mentioned streaming return device for knowledge Q&A further includes a cache cleaning module for emptying the preset document cache when all the streaming return content for the user input text has been sent to the client.

[0128] Each module in the above-mentioned streaming return device for knowledge Q&A can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0129] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a method for streaming return of knowledge Q&A.

[0130] Those skilled in the art can understand that Figure 6 the structure shown in

[0131] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0132] In an embodiment, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in the above-mentioned method embodiments.

[0133] In an embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, it implements the steps in the above-mentioned method embodiments.

[0134] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0135] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., and are not limited thereto.

[0136] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0137] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.

Claims

1. A method for streaming return of knowledge Q&A, characterized in that, Applied to the server side, the method includes: Input the user input text received from the client into a large language model to obtain incremental response content for gradually returning to the client; After obtaining the incremental response content required each time, based on the incremental response content required this time, retrieve in the pre-stored knowledge base to obtain target document information that matches the incremental response content required this time; Merge the incremental response content required this time and the target document information to obtain the streaming return content returned this time for the user input text, and send the streaming return content returned this time to the client in a streaming return manner.

2. The method according to claim 1, wherein The retrieving in the pre-stored knowledge base based on the incremental response content required this time to obtain target document information that matches the incremental response content required this time includes: Input the incremental response content required this time into a semantic parsing model to obtain the semantic vector corresponding to the incremental response content required this time; Perform similarity matching between the semantic vector and the semantic vectors of the pre-stored document information in the pre-stored knowledge base to determine the target document information with the highest similarity.

3. The method according to claim 1, wherein The inputting the user input text received from the client into a large language model to obtain incremental response content for gradually returning to the client includes: Input the user input text received from the client into a large language model to obtain the response content continuously output by the large language model in a streaming return manner; Based on the continuously output response content, determine the incremental response content for gradually returning to the client.

4. The method according to claim 3, wherein The determining the incremental response content for gradually returning to the client based on the continuously output response content includes: Continuously count the number of tokens corresponding to the continuously output response content; When the number of tokens is greater than a preset number threshold, start detecting and determining the latest text end identifier in the continuously output response content; Determine the response content with the latest text end identifier as the content end as the incremental response content required once, and clear the number of tokens.

5. The method according to claim 1, wherein After retrieving in the pre-stored knowledge base based on the incremental response content required each time to obtain target document information that matches the incremental response content required this time, it further includes: Store the target document information in a preset document cache; The retrieving in the pre-stored knowledge base based on the incremental response content required this time to obtain target document information that matches the incremental response content required this time includes: First, obtain the target document information that matches the incremental response content required this time from the preset document cache; In the case that there is no target document information in the preset document cache that matches the incremental response content required this time, then obtain the target document information that matches the incremental response content required this time from the pre-stored knowledge base.

6. The method according to claim 5, characterized in that, The method further includes: When all the streaming return content for the user input text has been sent to the client, clear the preset document cache.

7. A streaming return device for knowledge Q&A, characterized in that, The device includes: A large model application module, configured to input the user input text received from the client into a large language model to obtain incremental reply content for gradually returning to the client; A document matching module, configured to, after obtaining the incremental reply content required each time, retrieve in a pre-stored knowledge base based on the incremental reply content required this time to obtain target document information that matches the incremental reply content required this time; A streaming return module, configured to merge the incremental reply content required this time and the target document information to obtain the streaming return content returned this time for the user input text, and send the streaming return content returned this time to the client in a streaming return manner.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Question and answer and quotation generation method and system based on large language model

    CN119166774A

  • Construction method of real-time streaming voice intelligent question and answer service system

    CN119719438A

  • Reply generation system and method, and device

    WO2025039805A1

Cited By

  • Interactive response method and device based on large model, medium and equipment

    CN121480744A