Output result generation method and system, electronic equipment and storage medium
By determining the target search strategy in a large language model service and building a structured prompt message queue, the search strategy selection and result processing problems in a multi-knowledge base environment are solved, and efficient and accurate content generation is achieved.
Patent Information
- Application Number
- CN202510775878.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-11
AI Technical Summary
In the face of a multi-knowledge base environment, existing large-scale language model services cannot effectively select retrieval strategies, resulting in inefficient search efficiency and inaccurate results, and cannot efficiently integrate search results from multi-knowledge bases to generate high-quality output content.
By determining the target search strategy that matches the configuration of the inference application, select a single knowledge base or multiple knowledge base for searching, and construct a structured prompt message queue based on historical session information and system prompt words, and input it into a large language model to generate output results.
It realizes efficient and accurate search strategy selection and result integration in a multi-knowledge base environment, generates high-quality output results that are more in line with user needs, and improves search efficiency and information coverage.
Smart Images

Figure CN120277131A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and particularly to a method and system for generating output results, an electronic device, and a storage medium. Background Art
[0002] With the development of artificial intelligence technology, especially the progress in the field of natural language processing, services based on large language models (LLMs) have become an important means of information retrieval and content generation. However, when existing LLM services face complex queries, the retrieval strategies of their knowledge bases are often too single, unable to effectively handle the dynamic retrieval requirements and semantic understanding challenges in a multi-knowledge-base environment. Traditional methods, such as simple keyword retrieval or vector retrieval, often work independently in a single knowledge base. When encountering queries that require cross-knowledge-base retrieval, their efficiency and accuracy are greatly reduced. In addition, how to intelligently select the most appropriate retrieval strategy (single-base exact retrieval or multi-base parallel retrieval) according to the specific query content is also an unsolved problem. More importantly, a series of problems such as the processing and integration of retrieval results, and how to effectively convert these results into inputs for language models to generate high-quality content have not been well answered technically. Therefore, there is an urgent need for a more intelligent, efficient, and multi-knowledge-base-environment-adaptable retrieval strategy and result processing method to improve the performance of the RAG system in complex scenarios.
[0003] In related technologies, there is currently no effective solution to the problem of how to select a strategy of single knowledge base retrieval or multiple knowledge bases parallel retrieval and efficiently integrate retrieval results to generate high-quality output content. Summary of the Invention
[0004] This application provides a method and system for generating output results, an electronic device, and a storage medium, to at least solve the problem of how to select a strategy of single knowledge base retrieval or multiple knowledge bases parallel retrieval and efficiently integrate retrieval results to generate high-quality output content.
[0005] The present application provides a method for generating an output result, including: when a reasoning request of a target reasoning application is obtained, determining a target retrieval strategy according to the reasoning application configuration corresponding to the reasoning request, where the reasoning application configuration has configuration information corresponding to the target reasoning application; when the target retrieval strategy is a single knowledge base retrieval strategy, determining a target knowledge base from multiple knowledge bases and using the target knowledge base to determine the retrieval result of the reasoning content corresponding to the reasoning request; when the target retrieval strategy is a multi-knowledge base retrieval strategy, using multiple knowledge bases to jointly determine the retrieval result corresponding to the reasoning content; constructing a prompt message queue according to the prompt words corresponding to the reasoning application configuration, the historical session information corresponding to the reasoning request, the retrieval result, and the reasoning content; and generating the output result corresponding to the reasoning request by a target large language model according to the prompt message queue.
[0006] The present application further provides a system for generating an output result, including: a knowledge base retrieval routing component, configured to determine a target retrieval strategy according to the reasoning application configuration corresponding to the reasoning request when a reasoning request of a target reasoning application is obtained, where the reasoning application configuration has configuration information corresponding to the target reasoning application; when the target retrieval strategy is a single knowledge base retrieval strategy, determining a target knowledge base from multiple knowledge bases and using the target knowledge base to determine the retrieval result of the reasoning content corresponding to the reasoning request; when the target retrieval strategy is a multi-knowledge base retrieval strategy, using multiple knowledge bases to jointly determine the retrieval result corresponding to the reasoning content; a prompt word construction and context enhancement component, configured to construct a prompt message queue according to the prompt words corresponding to the reasoning application configuration, the historical session information corresponding to the reasoning request, the retrieval result, and the reasoning content; and a question and answer execution component, configured to generate the output result corresponding to the reasoning request by a target large language model according to the prompt message queue.
[0007] The present application further provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any one of the above methods for generating an output result when executing the computer program.
[0008] The present application further provides a computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the steps of any one of the above methods for generating an output result.
[0009] The present application further provides a computer program product including a computer program, where the computer program, when executed by a processor, implements the steps of any one of the above methods for generating an output result.
[0010] Through this application, by determining the target retrieval strategy that matches the inference application configuration, the deficiency of fixed retrieval strategies in traditional retrieval methods is solved. Specifically, when the inference request corresponds to a single knowledge base retrieval strategy, it can select the most suitable target knowledge base for precise retrieval from them, avoiding the retrieval of irrelevant knowledge bases and improving the retrieval efficiency and pertinence. When faced with a multi-knowledge base retrieval strategy, this method can use multiple knowledge bases for joint retrieval to ensure more comprehensive information coverage and overcome the information blind spots that may exist in a single knowledge base. Further, by combining the retrieval results with historical session information, system prompt words, and the user's current inference content, a structured prompt message queue is constructed and used as the input to the large language model. This multi-information source fusion strategy enriches the input context of the model, ensures the consistency of the input and the effective utilization of the model, and thus generates output results that are more in line with user needs and of high quality. That is, through intelligent selection of retrieval strategies, efficient integration of retrieval results, and optimization of model input, this application provides a more accurate and efficient content generation method, effectively solving the problems of retrieval strategy selection and result processing in a multi-knowledge base environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0012] Figure 1 is a hardware structure block diagram of a method for generating output results according to an embodiment of the present application;
[0013] Figure 2 is a schematic diagram of a RAG retrieval process according to an embodiment of the present application;
[0014] Figure 3 is an architecture diagram of a RAG inference system according to an embodiment of the present application;
[0015] Figure 4 is a flowchart of a method for generating output results according to an embodiment of the present application;
[0016] Figure 5 is a working principle diagram of a knowledge base retrieval routing component according to an embodiment of the present application;
[0017] Figure 6 is a structure block diagram of a system for generating output results according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the protection scope of the present application.
[0019] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0020] In order to enable those skilled in the art of this technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0021] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the output result generation method depends, the specific application environment architecture or specific hardware architecture is described herein.
[0022] The method embodiments provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 1 is a hardware structure block diagram of an output result generation method according to an embodiment of the present application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in Figure 1 processors 102 (the processors 102 may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only illustrative and does not limit the structure of the above server device. For example, the server device may further include more or fewer components than Figure 1 shown, or have a different configuration from
[0023] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the startup method of the operating system in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include memories remotely provided with respect to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.
[0024] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the server device. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0025] The following explains the professional terms that appear in this application:
[0026] RAG (Retrieval-Augmented Generation): An artificial intelligence method that combines information retrieval and text generation. First, relevant content is retrieved from an external knowledge source, and then the retrieval result and the user input are jointly injected into a large language model (LLM) to generate a response, which is widely used in scenarios such as question answering, dialogue, and content creation. It should be noted that Figure 2 Schematically shows a schematic diagram of a RAG retrieval process.
[0027] LLM (Large Language Model): Refers to a large-scale pre-trained language model based on the Transformer architecture, such as DeepSeek, ChatGPT, Claude, Gemini, etc., which has capabilities such as context understanding, language generation, and reasoning, and is the core component for performing RAG generation.
[0028] Prompt (prompt): A structured input text used to guide the generation of a large language model, usually composed of a system prompt, user input, historical context, and external knowledge, and is a key factor in the quality of RAG generation.
[0029] Knowledge Base: A collection of semantic content that stores structured or unstructured information. In this application, each knowledge base can be independently configured with a retrieval method, an embedding model, and a re-ranking strategy, serving as the basic information source in the RAG retrieval stage.
[0030] ReAct (Reasoning + Acting): A language model interaction paradigm that combines reasoning and actions (such as calling tools, selecting paths), which can be used to guide large language models to perform intelligent selection and planning in multi-knowledge base scenarios.
[0031] Function Calling: A mechanism that enables the model to autonomously select and call predefined functions or tools based on the context content. In the RAG system, it is used for the model to select the most suitable knowledge base or perform retrieval behaviors.
[0032] Embedding (Vector Representation): The process of converting text into a high-dimensional vector form for calculating semantic similarity. In this application, both knowledge documents and user queries are vectorized through an embedding model to achieve semantic retrieval.
[0033] pgvector: An open-source vector database plugin based on PostgreSQL, which is used to perform efficient vector retrieval and hybrid retrieval tasks in this application, supporting parallel queries and similarity calculations across multiple knowledge bases.
[0034] Hybrid Retrieval: An integrated retrieval method that combines vector retrieval and full-text retrieval simultaneously. It can balance semantic matching and keyword precision, improving the recall rate of relevant content. A common approach is to independently perform the two retrievals and then merge the results using simple strategies (such as score weighting, Reciprocal Rank Fusion - RRF). However, the result fusion strategy is usually fixed and lacks in-depth consideration and dynamic adaptability to query intent or knowledge base characteristics.
[0035] Multi-knowledge base processing: a) Parallel retrieval: Query all available knowledge bases simultaneously and aggregate the results. It is simple but inefficient and prone to introducing noise; b) Simple routing: Direct the query to a specific knowledge base based on predefined rules or metadata tags. It is more targeted than parallel retrieval, but the rules are rigid and difficult to handle complex queries. However, this multi-knowledge base processing method has deficiencies in terms of efficiency, flexibility, and the depth of processing query semantics.
[0036] LLM-based Initial Routing / Tool Selection: Leveraging the understanding ability of LLM to select appropriate knowledge bases (regarded as "tools" or "function calls"). Although it enhances the intelligence of routing, existing applications often stay at the level of "which library to select" and lack in-depth integration with the subsequent details of hybrid retrieval execution. They usually do not provide a mechanism for dynamically and intelligently switching between "single-library exact retrieval" and "multi-library parallel retrieval + fusion", nor do they have a systematic design for uniformly normalizing and structurally encapsulating the retrieved results to optimize the use of downstream LLM. LLM routing itself may also introduce additional latency.
[0037] Although related technologies have explored in aspects such as hybrid retrieval, multi-knowledge base query, and LLM-based initial routing, there is generally a lack of an integrated and systematic solution. Specifically, there is a lack of a mechanism that can tightly combine advanced hybrid retrieval algorithms with intelligent and dynamic knowledge base access strategies (including exact routing and parallel retrieval). Existing solutions are often discrete, either with simple hybrid retrieval strategies, or with fixed routing logic or decoupled from retrieval execution, making it difficult to flexibly adjust the overall retrieval process according to query characteristics and model capabilities (such as Function Calling). At the same time, there is also a lack of a unified and perfect design for how to effectively normalize, sort, and contextually encapsulate multi-source results after retrieval to optimally serve the RAG generation link. These deficiencies limit the retrieval efficiency, accuracy, and final generation quality of RAG systems in complex scenarios. This application aims to fill these technical gaps.
[0038] To solve the problems existing in related technologies, as Figure 3 shown, this application proposes a RAG inference system for inference applications, aiming to improve the stability, efficiency, and traceability of RAG inference services. The system includes multiple core modules: 1) Inference request concurrency management component, which manages the resource usage of inference services by setting a maximum concurrency threshold; 2) Inference application configuration management component, which realizes configuration version control and priority binding in multi-turn conversations to ensure the consistency and stability of conversations; 3) System observability and behavior tracking component, which collects and records core operations and performance data during the inference process in real time to support fault diagnosis and performance optimization; 4) Session and message management component, which ensures the context integrity and interaction status of multi-turn conversations; 5) Knowledge base retrieval routing component, which optimizes the fusion efficiency of knowledge base selection and retrieval results; 6) Prompt construction and context enhancement component, which improves the response accuracy of multi-turn conversations; 7) LLM Q&A execution component driven by knowledge base information, which efficiently connects prompt input with the question-answering ability of large language models. By integrating and optimizing each module, the system significantly improves the reliability and user experience of RAG inference applications.
[0039] It should be noted that in this embodiment, a method for generating an output result is provided, which is applied to the above-mentioned RAG inference system. Figure 4 It is a flowchart of a method for generating an output result according to an embodiment of the present application, as Figure 4 shown. The method includes the following steps S402 - S408:
[0040] Step S402: When a reasoning request of a target reasoning application is obtained, determine a target retrieval strategy according to the reasoning application configuration corresponding to the reasoning request, where the reasoning application configuration has configuration information corresponding to the target reasoning application;
[0041] Step S404: When the target retrieval strategy is a single knowledge base retrieval strategy, determine a target knowledge base from multiple knowledge bases, and use the target knowledge base to determine the retrieval result of the reasoning content corresponding to the reasoning request; when the target retrieval strategy is a multi - knowledge base retrieval strategy, use multiple knowledge bases together to determine the retrieval result corresponding to the reasoning content;
[0042] It should be noted that the above steps S402 and S404 can be executed by the knowledge base retrieval routing retrieval component.
[0043] Step S406: Construct a prompt message queue according to the prompt words corresponding to the reasoning application configuration, the historical session information corresponding to the reasoning request, the retrieval result, and the reasoning content;
[0044] Step S408: Generate the output result corresponding to the reasoning request through the target large - language model according to the prompt message queue.
[0045] The above steps solve the deficiency of fixed retrieval strategies in traditional retrieval methods by determining a target retrieval strategy that matches the reasoning application configuration. Specifically, when the reasoning request corresponds to a single knowledge base retrieval strategy, it can select the most suitable target knowledge base for precise retrieval from them, avoiding the retrieval of irrelevant knowledge bases and improving the retrieval efficiency and pertinence. When facing a multi - knowledge base retrieval strategy, the method can use multiple knowledge bases for joint retrieval to ensure more comprehensive information coverage and overcome the information blind spots that may exist in a single knowledge base. Further, by combining the retrieval result with the historical session information, system prompt words, and the user's current reasoning content, a structured prompt message queue is constructed, which is used as the input of the large - language model. This multi - information - source fusion strategy enriches the input context of the model, ensures the consistency of the input and the effective utilization of the model, and thus generates output results that are more in line with user needs and of high quality. That is, the present application provides a more accurate and efficient content generation method by intelligently selecting retrieval strategies, efficiently integrating retrieval results, and optimizing model input, effectively solving the problems of retrieval strategy selection and result processing in a multi - knowledge base environment.
[0046] In an exemplary embodiment, before determining the target retrieval strategy according to the inference application configuration corresponding to the inference request, the method further includes the following steps S11 - S13: Step S11: When an inference request of the target inference application is obtained, determine whether the current concurrent inference request count is greater than a preset threshold; Step S12: When the current concurrent inference request count is greater than the preset threshold, do not respond to the inference request; Step S13: When the current concurrent inference request count is less than or equal to the preset threshold, generate a request identifier for the inference request, and track and manage the entire life cycle of the inference request according to the request identifier;
[0047] Determining the target retrieval strategy according to the inference application configuration corresponding to the inference request includes: when the current concurrent inference request count is less than or equal to the preset threshold, determining the retrieval strategy according to the inference application configuration of the target inference application corresponding to the inference request.
[0048] It should be noted that the above steps S11 - S13 can be executed by the inference request concurrency management component. The functions of the inference request concurrency management component are described in detail below:
[0049] The inference request concurrency management component provides a concurrency request control mechanism for inference application services, aiming to ensure the stability and response efficiency of the large - model inference service in a conversational generation scenario with multi - user concurrent access. This mechanism judges the acceptance of each request entering the RAG inference system by setting the maximum concurrent request threshold at the application level (i.e., the above - mentioned preset threshold), and generates a unique identifier for tracking and managing the entire life cycle of the request.
[0050] In the normal conversation generation process, the inference request concurrency management component records the request identifier at the beginning of the request, and automatically triggers the resource release logic after the generation is completed or the data stream output ends. Through a unified recycling mechanism, the inference request concurrency management component removes the request identifier from the flow - limiting control structure, updates the current active request count, ensures that the RAG inference system can accept new requests in a timely manner, and avoids long - term resource occupation or abnormal concurrent counting.
[0051] In addition, the inference request concurrency management component also has an exception detection and resource recycling mechanism, which can automatically clean up the relevant resource status in abnormal scenarios such as request timeout, interruption, or execution failure, preventing "hanging" requests from affecting the overall service operation. As the core scheduling and protection component in the inference service, this mechanism effectively improves the system's concurrent bearing capacity and resource utilization efficiency, and is a key basic ability to ensure the deployment of highly reliable conversational generation services.
[0052] In an exemplary embodiment, the method further includes: during the creation of the target inference application, creating an inference application configuration for the target inference application according to the obtained configuration information, where the inference application configuration includes: a prompt, an associated knowledge base set, retrieval configuration information, an identifier of the large language model, and model parameters of the large language model; the retrieval configuration information includes a retrieval strategy.
[0053] It should be noted that the above steps can be executed by the inference application configuration management component. Specifically, the inference application configuration management component has a set of versioned inference application configuration management mechanisms for knowledge base inference applications based on RAG to achieve configuration stability and interaction consistency control during multi-round conversations. This mechanism focuses on the "inference application configuration" and ensures that users can obtain accurate and coherent model behavior responses when conducting dialogue retrieval in different scenarios by introducing configuration version binding and configuration priority determination.
[0054] When each inference application is created, the system will automatically generate a configuration entity for it, called the "inference application configuration". This configuration contains the following key information: System Prompt; associated knowledge base set (Knowledge Base); vector retrieval and re-ranking model and its parameter configuration; the dialogue model used and its specific parameters (such as temperature, Top-p, maximum number of tokens, etc.); other metadata related to inference control;
[0055] The inference application configuration has the ability to manage versions. Whenever the user updates the configuration content of the inference application (such as replacing the knowledge base, adjusting model parameters, modifying the prompt, etc.), the system will generate a new configuration version and retain the historical versions. This is convenient for backtracking, comparison, and debugging on the one hand, and can be used as a basic guarantee for maintaining semantic consistency in multi-round conversations on the other hand.
[0056] In an exemplary embodiment, before determining the target retrieval strategy according to the inference application configuration corresponding to the inference request, the method further includes the following steps S21 - S23:
[0057] Step S21: When the inference request is the first-round dialogue request of a new session of the target inference application and does not carry the inference application configuration, use the latest version of the inference application configuration associated with the target inference application as the inference application configuration corresponding to the inference request;
[0058] Step S22: When the inference request is a dialogue request of an existing target session of the target inference application and does not carry the inference application configuration, use the inference application configuration corresponding to the target session as the inference application configuration corresponding to the inference request;
[0059] Step S23: When the inference request carries the inference application configuration, use the carried inference application configuration as the inference application configuration corresponding to the inference request.
[0060] It should be noted that the above steps can be executed by the inference application configuration management component. Specifically, when the user initiates a conversation with the inference application, the system adopts a three-level priority configuration binding mechanism based on the session status and the user request content:
[0061] First priority: Explicitly passed-in configuration (immediate configuration);
[0062] When the user initiates a request in debug mode, a complete inference application configuration can be passed to the system along with the request. At this time, the system will give priority to using this immediate configuration to execute the conversation, ignoring other configurations bound to the current session or application.
[0063] Second priority: Session-bound configuration (historical consistency);
[0064] If the current is a subsequent dialogue turn of an existing session, the system will use the version of the inference application configuration bound in the first round of the session to ensure the consistency of the model behavior and the stability of the knowledge base context throughout the conversation.
[0065] Third priority: The latest configuration of the inference application;
[0066] If the current request is the first dialogue turn of a new session for a certain inference application and no configuration is explicitly passed in, the system will default to using the latest configuration version associated with the inference application as the session execution configuration and bind this configuration version to the session when the session is created.
[0067] This three-level priority mechanism ensures the stability and continuity of the inference service while guaranteeing flexibility. Especially in cases involving dynamic adjustment of the knowledge base, optimization of prompt words, and switching of model rearrangement strategies, it can avoid semantic drift or abnormal behavior in the conversation.
[0068] For better understanding, the following is an illustration of the application scenarios:
[0069] Formal usage scenario: The user initiates a conversation with a certain inference application through the UI or API without caring about the details of the underlying inference application configuration. The system will automatically select an appropriate inference application configuration according to the above priority rules to achieve seamless configuration management.
[0070] Debugging and testing scenario: Developers or advanced users can explicitly pass in a complete inference application configuration through the debug mode to quickly verify the impact of different knowledge base combinations, prompt word schemes, or model parameters on the generation effect, supporting efficient optimization.
[0071] Through the above mechanism, the system realizes the traceable, controllable, and extensible management capabilities of inference application configurations. While ensuring the consistency of the user experience, it takes into account the debugging flexibility and system governance capabilities, which is one of the key designs for the reliable operation of the RAG inference system in the actual production environment.
[0072] In an exemplary embodiment, determining a target retrieval strategy according to the inference application configuration corresponding to the inference request includes the following steps S31 - S33:
[0073] Step S31: Determine whether the target large language model used by the target inference application supports function - call inference capabilities or chain - of - thought inference capabilities; and obtain the retrieval strategy from the inference application configuration.
[0074] Step S32: When the target large language model supports function - call inference capabilities or chain - of - thought inference capabilities, and the retrieval strategy is a single knowledge - base retrieval strategy, determine the target retrieval strategy as the single knowledge - base retrieval strategy.
[0075] Step S33: When the target large language model does not support function - call inference capabilities or chain - of - thought inference capabilities, or the retrieval strategy is a multi - knowledge - base retrieval strategy, determine the target retrieval strategy as the multi - knowledge - base retrieval strategy.
[0076] It should be noted that the above steps are executed by the knowledge base retrieval routing component. In the process of determining the target retrieval strategy, the knowledge base retrieval routing component first intelligently identifies whether the large language model (LLM) adopted by the target inference application has advanced inference capabilities such as function calling or reasoning with actions (ReAct). This step is crucial because it determines the selection and execution method of the subsequent retrieval strategy. If the target LLM supports function calling or reasoning with actions, and the inference application configuration has clearly specified the single knowledge base retrieval mode, then the system will utilize the intelligent decision-making ability of the LLM to precisely select one target knowledge base that is most relevant to the current query from multiple preset knowledge bases, perform efficient and accurate retrieval, avoid waste of resources and interference from redundant information, and improve the retrieval efficiency and relevance of the results. On the contrary, when the target LLM does not support the above advanced inference capabilities, or the inference application configuration points to the multi-knowledge base retrieval strategy, the system adopts the parallel retrieval method, retrieves information from all knowledge bases simultaneously, and then processes the retrieval results through the normalization and fusion sorting mechanism to ensure that the information obtained from multiple sources can be effectively integrated, improve the comprehensiveness and accuracy of the retrieval from a higher global perspective, and provide richer and more comprehensive context information for subsequent content generation. This flexible strategy adjustment mechanism not only gives full play to the inference advantages of the LLM but also ensures that the RAG system can work in the optimal way in different scenarios, meet complex query requirements, optimize resource utilization and execution efficiency, and significantly improve the overall service quality and user experience.
[0077] In an exemplary embodiment, determining the target knowledge base from multiple knowledge bases includes: converting each knowledge base in the multiple knowledge bases into a structured tool description object to obtain multiple tool description objects, where each tool description object contains the unique identifier and usage description of the corresponding knowledge base; determining the target knowledge base from the multiple knowledge bases through the target large language model according to the multiple tool description objects and the inference content.
[0078] It should be noted that the above steps are executed by the knowledge base retrieval routing component. To achieve the optimal retrieval of a single knowledge base, this application adopts an innovative tool description object conversion mechanism. Specifically, the system first abstracts each knowledge base into a tool description object carrying a unique identifier and a clear usage description, aiming to enable the target large language model to understand the attributes and functions of each knowledge base, similar to providing the model with a "user manual" for the knowledge bases. Subsequently, based on the user's reasoning content and this "user manual", the model intelligently analyzes and selects the target knowledge base that best matches the query intention. This process effectively simulates the behavior of humans selecting the most appropriate data source when solving problems, making the retrieval more accurate and efficient. Through the semantic understanding and dynamic decision-making of the model, not only is the pertinence of the retrieval improved, but also the access to irrelevant knowledge bases is reduced, greatly saving system resources, optimizing the quality of the retrieval results, and laying a solid foundation for generating high-quality output content subsequently.
[0079] In an exemplary embodiment, using the target knowledge base to determine the retrieval result of the reasoning content corresponding to the reasoning request can be implemented through the following steps S41 - S42:
[0080] Step S41: Determine the configuration parameters of the target knowledge base, where the configuration parameters include: retrieval method, identifier of the embedding model, similarity scoring threshold, and number of returned segments;
[0081] Among them, the retrieval method: refers to the method used to find document segments. Common ones include vector retrieval (based on semantic similarity) and full-text retrieval (based on keyword matching). The system selects the most suitable retrieval method according to the characteristics of the knowledge base.
[0082] Identifier of the embedding model: represents the model type used to construct document vectors. Different embedding models affect the construction method of text vectors and the accuracy of similarity calculation.
[0083] Similarity scoring threshold: Set a score threshold. Only when the similarity score of the document segment to the query content reaches or exceeds this threshold will it be included in the retrieval results, used to filter out irrelevant information and improve the relevance of the results.
[0084] Number of returned segments: Specify the upper limit of the number of document segments returned by the system during retrieval, avoiding returning too much irrelevant content, and at the same time controlling the input scale of downstream processing to ensure the efficiency and accuracy of model generation.
[0085] Step S42: According to the reasoning content, use the configuration parameters to retrieve the target knowledge base, and normalize and encapsulate the retrieved content to obtain the retrieval result, where the retrieval result includes: retrieved segment text, similarity score of the retrieved segment text, and the position of the retrieved segment text in the target knowledge base.
[0086] It should be noted that the above steps are executed by the knowledge base retrieval routing component. When determining to use the target knowledge base for retrieval, the present application further refines the retrieval operation steps to improve the accuracy of retrieval and the usability of the results. First, the configuration parameters of the target knowledge base are read, including the retrieval method, the identifier of the embedding model, the similarity scoring threshold, and the number of returned fragments. These parameters jointly guide the execution of the retrieval task. Subsequently, based on these precise configurations, the system performs a retrieval on the target knowledge base, efficiently calculates the document similarity through the vector database, filters out the high-score paragraphs, and performs normalization processing and context encapsulation to form a structured retrieval result, which includes the retrieved text content, similarity score, and specific position information in the knowledge base. This structured output not only facilitates subsequent processing but also ensures the integrity and accuracy of the retrieval results, providing a solid foundation for the model to generate high-quality content.
[0087] In an exemplary embodiment, multiple knowledge bases are used to jointly determine the retrieval result corresponding to the inference content, including the following steps S51 - S52:
[0088] Step S51: Determine the configuration parameters of each of the multiple knowledge bases respectively, and perform retrievals on the multiple knowledge bases respectively based on the corresponding configuration parameters to obtain multiple retrieval sub-results. Among them, the configuration parameters include: retrieval method, identifier of the embedding model, similarity scoring threshold, number of returned fragments;
[0089] Optionally, the knowledge base retrieval routing component can perform retrievals on each knowledge base through different threads.
[0090] Step S52: Perform fusion sorting and re-screening on the multiple retrieval sub-results, and perform normalization and context encapsulation on the results of the fusion sorting and re-screening to obtain the retrieval result. Among them, the retrieval result includes: retrieved fragment text, similarity score of the retrieved fragment text, knowledge base corresponding to the retrieved fragment text, and position in the knowledge base.
[0091] It should be noted that the above steps are executed by the knowledge base retrieval routing component. For better understanding, Figure 5 schematically shows the working principle diagram of the knowledge base retrieval routing component. The following combines Figure 5 to elaborate in detail on the functions of the knowledge base retrieval routing component:
[0092] The knowledge base retrieval routing component supports automatically selecting the most relevant knowledge base or accessing multiple knowledge bases in parallel according to user queries in a multi-knowledge base environment, and performs score normalization and context encapsulation on the retrieval results to provide high-quality generation support information for large language models. This component is particularly suitable for language model inference service systems that support Function Calling or ReAct mechanisms, and has advantages such as flexible strategies, semantic-driven, and resource-efficient. The core functions of the component include the following aspects:
[0093] 1) Model ability perception and retrieval initialization mechanism:
[0094] At the initiation of each retrieval task, the system first identifies whether the large language model used in the current inference application supports function calling (Function Calling) or chain of thought (ReAct) inference capabilities, and based on this determines whether the retrieval strategy allows the model to participate in planning. On this basis, the system extracts the retrieval configuration of the inference application, including parameters such as top k, score threshold, re-ranking model, and embedding model, to establish a parameter basis for subsequent retrieval scheduling.
[0095] 2) Single knowledge base preferred retrieval mechanism:
[0096] When the inference application is configured in the "single knowledge base preferred mode", the system uses the semantic understanding ability of the large language model for the user query to assist in selecting the target knowledge base that best matches the current request, and only performs an efficient directional retrieval on this knowledge base. This strategy can effectively reduce redundant calculations and multi-source conflicts, and is suitable for scenarios with clear domains and concentrated knowledge. The specific process includes:
[0097] 1. The system converts the candidate knowledge bases into structured tool description objects, each object containing its unique identifier and usage description (such as knowledge type, domain label, etc.), for the large language model to identify and compare during the selection process.
[0098] For example, if the system has three candidate knowledge bases for "product usage instructions", "customer feedback records", and "sales data analysis" respectively, three corresponding tool descriptions will be generated, such as: "used to answer questions related to product functions and usage instructions", "contains user feedback and common question summaries", "stores sales trends and market analysis content".
[0099] 2. The system dynamically selects an adapted routing module based on the model's capabilities (supporting inference mechanisms such as Function Calling or ReAct). The model receives the user query and the knowledge base tool list, and combines semantic analysis and context judgment to output the recommended target knowledge base identifier.
[0100] For example: When a user asks a query such as "What is the battery life of this product?", the large language model can, based on semantic understanding, preferentially select a tool object related to "product description", and thus route to the most matching knowledge base for targeted retrieval.
[0101] 3. The system reads the local configuration parameters of the selected knowledge base (such as retrieval method, embedding model, score threshold, top k, etc.), and based on these configurations, uses a vector database (such as pgvector) to perform full-text retrieval or vector retrieval tasks. In a hybrid retrieval scenario, the system can initiate vector similarity queries and full-text retrieval queries separately, and merge the results of both to filter out the top k relevant paragraphs with scores higher than the threshold as the knowledge input basis for subsequent generation.
[0102] 4. The paragraph data returned by the retrieval will be encapsulated into a context object in a unified format, along with content, similarity scores, and location information. At the same time, the system records the knowledge base selection path, query behavior, and paragraph hit count for this time, forming a traceable link for subsequent analysis and optimization.
[0103] 3) Multi-knowledge base parallel retrieval and fusion sorting mechanism:
[0104] When the inference application is configured in the "multi-knowledge base concurrent retrieval mode", the system schedules all available knowledge bases in parallel, executes independent retrieval tasks according to their respective configurations, and unifies and fuses the output of the final results to form the structured context information required by the downstream generation module.
[0105] The system assigns an independent thread to each knowledge base, reads its local configuration (including retrieval method, top k, score threshold, etc.), and executes the query. The retrieval method is specified by the configuration, usually vector retrieval or vector + full-text combined retrieval, which is supported by the underlying vector database (such as pgvector).
[0106] Each thread only returns the top k paragraphs with scores higher than the threshold. The main thread merges and summarizes the results returned by all threads, and constructs a unified document segment structure, marking its source knowledge base, original score, and metadata information.
[0107] The system uniformly performs fusion sorting and re-screening on all candidate paragraphs according to the global configuration parameters of the inference application. The fusion logic supports multiple strategies (such as default score sorting, preset weights, etc.), and finally retains the top k paragraphs that meet the global score threshold for output for subsequent components to use.
[0108] The filtered paragraphs are encapsulated into a unified context object in a structured form, along with content, scores, sources, and location information, for prompt construction and large model generation tasks.
[0109] Through this mechanism, it is possible to efficiently access multiple knowledge sources under configuration drive, control retrieval concurrency, unify the output format, avoid upstream and downstream coupling, and provide key support for the modular deployment and scalability of the RAG system.
[0110] 4) Retrieval result normalization and context encapsulation mechanism:
[0111] All finally selected paragraph contents will be standardized and encapsulated into document context objects in a unified format. Each object contains metadata such as paragraph text, similarity score, source knowledge base, paragraph number, location information, etc., ensuring that the model input has good structural consistency and semantic interpretability. These objects support being directly used as input in generation tasks and also support being displayed as hit references on the front end.
[0112] 5) Index hit callback and observable feedback mechanism:
[0113] The component injects a hit callback processor during the initialization stage of the retrieval task, which is automatically called after the retrieval is completed to push the source information and metadata of the finally hit documents to the message queue or log system, realizing real-time reference feedback and behavior traceability. At the same time, the system will automatically update the hit count of the paragraphs and submit relevant tracking tasks to ensure the closed-loop transparency of the retrieval link.
[0114] 6) Concurrency safety and execution performance optimization mechanism:
[0115] The system uses a thread lock mechanism to protect the security of shared structure writes in a multi-threaded scenario to avoid concurrency conflicts; and records the key time nodes of each round of query process through a high-precision time-consuming tracking function to form performance monitoring metrics; in scenarios without reordering requirements, it automatically enables lightweight keyword matching and vector scoring paths to balance performance and accuracy.
[0116] In summary, the knowledge base retrieval routing component proposed in this application integrates the semantic planning ability of the large language model and the configurable retrieval ability of the knowledge base, supports two modes: single library optimization and multi-library concurrency, and combines with the vector database to achieve efficient query and unified encapsulation. The component has the capabilities of structured context generation, behavior tracking, and result normalization, significantly improving the retrieval efficiency, generation quality, and observability of the RAG inference system, and is a key support module for building an intelligent question-answering system with multiple knowledge sources.
[0117] In an exemplary embodiment, constructing the prompt message queue according to the corresponding prompt words of the inference application, the historical session information corresponding to the inference request, the retrieval results, and the inference content can be implemented through the following steps S61 - S63:
[0118] Step S61: Screen out target segments from the historical session information whose relevance to the inference content is higher than a preset relevance threshold;
[0119] Step S62: When the total number of tokens of the prompt, target segment, retrieval result, and inference content is greater than the preset number, crop the target segment and retrieval result according to the preset number, and construct a prompt message queue based on the cropped target segment, retrieval result, prompt, and inference content;
[0120] Step S63: When the total number of tokens of the prompt, target segment, retrieval result, and inference content is less than or equal to the preset number, directly construct a prompt message queue based on the target segment, retrieval result, prompt, and inference content.
[0121] It should be noted that the above steps S61 - S63 are executed by the prompt construction and context enhancement component,
[0122] This component is used to integrate multiple types of input content into a structured prompt message sequence for the chat large language model (LLM) to call before the inference request is executed. Its goal is to improve the response accuracy and context coherence in multi - turn dialogue scenarios, and it is an indispensable intermediate processing module in the RAG inference system.
[0123] It should be noted that the construction of the system prompt content focuses on the organization and integration of the following four core types of information:
[0124] System prompt: Specified by the inference application configuration, defining the model behavior boundary and dialogue context;
[0125] Historical conversation information: Extracting context content from previous messages to supplement dialogue continuity;
[0126] Knowledge base context: Injecting the paragraph content returned by the knowledge retrieval component as auxiliary information into the prompt;
[0127] User's current question: As the core input of the current round of dialogue, clarifying the generation intention.
[0128] This component supports the following key functions:
[0129] 1) Memory integration and dialogue history embedding mechanism:
[0130] According to the memory strategy and Token limit configured by the system, filter out highly relevant segments from historical messages, convert them into user and assistant role messages, and insert them into the prompt structure. This process retains the key path of the dialogue context and improves the model's ability to understand continuous semantics.
[0131] 2) Retrieval context and multi - modal information fusion mechanism:
[0132] Wrap the paragraphs retrieved from the knowledge base into text content with a unified structure and inject it into the prompt sequence as supplementary input for the user role; if there is multimodal information such as images and documents uploaded by the user, it can also be converted into structured prompt content and integrated uniformly to enable the model to comprehensively perceive external knowledge and context.
[0133] 3) Token budget evaluation and prompt structure pruning mechanism:
[0134] After the prompt structure is completed, the system will estimate the total number of tokens. When it exceeds the context window supported by the model, the system will preferentially prune the historical messages and knowledge base content to ensure that the system prompt words and user input are completely retained, thus ensuring the expression integrity and context compactness of the prompt.
[0135] Through the above mechanism, the prompt word organization and context fusion component provides a prompt input with structural norms, semantic coherence, and strong context awareness for the large model reasoning, which is the key support module for improving the generation effect and user experience.
[0136] In an exemplary embodiment, the above-mentioned method for generating an output result corresponding to an inference request by a target large language model based on a prompt message queue includes the following steps S71 - S72: Step S71: Construct a standardized request according to the model parameters of the target large language model and the prompt message queue in the inference application configuration corresponding to the inference request, where the standardized request adapts to the interface protocol of the target large language model; Step S72: Invoke the target large language model according to the standardized request to obtain the output result of the target large language model.
[0137] In an exemplary embodiment, after the above-mentioned method for generating an output result corresponding to an inference request by a target large language model based on a prompt message queue, the method further includes the following steps S81 - S82: Step S81: Package the output result into a structured output object, where the structured output object includes: response content, reference context, token consumption situation, and meta information of the generation task; the meta information includes: version, execution time, and status information of the target large language model; Step S82: Display the structured output object and record the structured output object in the session record corresponding to the inference request.
[0138] It should be noted that the above steps S71 - S72 and steps S81 - S82 are all executed by the LLM Q&A execution component (abbreviated as the Q&A execution component) driven by the knowledge base information. This component is used to submit the constructed structured prompt words input to the large language model (LLM) in the inference application to complete the Q&A generation task. This component is located at the end of the RAG inference process and is responsible for externally outputting a controllable, structured, and traceable generation response.
[0139] The Q&A execution component first receives the sequence of prompt messages output by the prompt construction and context enhancement component. This sequence has integrated the system prompt, historical conversation context, knowledge base retrieval content, and the user's current question, and has completed structural optimization and Token budget trimming according to the inference application configuration. Instead of directly processing document segments or knowledge retrieval results, the Q&A execution component makes model calls based on the encapsulated high-quality Prompt.
[0140] Subsequently, the Q&A execution component constructs a standardized generation request based on the model parameters specified in the inference application (such as temperature, top_p, Token limit, whether to enable streaming output, etc.), and adapts to the interface protocols of different model service providers. The system supports two call modes: synchronous and asynchronous (SSE). In the streaming mode, the generated content and reference information can be returned in real time, improving the response speed and user interaction experience.
[0141] During the execution of the generation task, the component has a sound exception handling mechanism, supporting the identification and fault tolerance of failure situations such as timeouts, interruptions, and format exceptions; at the same time, it is deeply integrated with the system observability tracking component, automatically recording meta-information such as Token consumption, latency, and model configuration used in this inference call, forming a complete and traceable record chain.
[0142] After generation is completed, the Q&A execution component encapsulates the model response into a structured output object, which includes the final answer content, the snapshot of the used Prompt, the mapping of the referenced knowledge base paragraphs (if applicable), and the generation-related meta-information. This response object will be synchronously written into the session and message management system, serving as the standard source for subsequent context continuation and response display.
[0143] In summary, the Q&A execution component realizes the efficient connection between prompt input and the Q&A ability of the large model, and is the key terminal execution module to ensure the response accuracy, context consistency, and tracking closed-loop integrity of the retrieval-enhanced generation process.
[0144] In an exemplary embodiment, the method further includes: creating a session record for the inference request when the inference request is the first-round dialogue request of a new session of the target inference application; updating the active time of the session record corresponding to the target session when the inference request is a dialogue request of an existing target session of the target inference application.
[0145] In an exemplary embodiment, the method further includes: generating a message record for each interaction during the interaction in the session, where the message record includes: the interaction input, the inference application configuration corresponding to the interaction, the relationship between the interaction and the previous interaction, and the placeholder for the preset response area; recording the message record in the session record of the session.
[0146] In an exemplary embodiment, the method further includes: when an attachment file is carried in an inference request, generating an attachment index of the attachment file and binding the attachment index to the corresponding message record.
[0147] It should be noted that the above steps are executed by the session and message management component, which can ensure that the inference application has stable context-bearing ability and interaction state recording ability during multi-round conversations, and dynamically creates, updates, and persistently manages the session records and message entities associated with each user request.
[0148] The core functions of the session and message management component focus on two levels: one is the session-level state management of the inference application, and the other is the input-output tracking at the message level. When the user triggers an inference request, the system first determines the user identity type (such as an end-user or a console account) and the invocation mode (such as a formal invocation or a debugging mode) based on the incoming inference application identifier and the invocation source. Subsequently, the system establishes or obtains the corresponding session record according to the current configuration version of the inference application (including information such as the model service provider, model identifier, and knowledge base configuration).
[0149] For the first initiated request, the system will create a new session record, the content of which includes the inference application identifier, model configuration information, session mode, initial input parameters, and user identity. If it is an existing session, only its last active time will be updated to maintain the timeliness of the conversation context.
[0150] Above the session layer, the system will also generate a new message record for each interaction, which is used to store key information such as the user input in this round, the inference configuration snapshot (referring to the specific configuration settings for this inference request, including model parameters (such as temperature, top_p), the selection of the knowledge base, the Token usage policy, etc.), the parent message association (used to build the context chain), and the placeholder in the preset response area. These data will be written into the database in a structured manner and be dynamically updated during the execution of the inference task, including the answer content generated by the model, the Token consumption statistics, the response latency, and the relevant price information.
[0151] In addition, when the user request contains an uploaded file (such as a document or an image), the system will automatically create an attachment metadata record and bind it to the message to ensure that the inference task has a complete resource reference path when it needs to call multi-modal information.
[0152] As a key support module in the dialogue process of the inference application, the session and message management component realizes the full process coverage of the organization of the multi-round session structure, the tracking of the message chain, and the state update, providing complete and reliable context data support for the retrieval-augmented generation task.
[0153] In an exemplary embodiment, the method further includes: when the target inference application is started, initializing a tracing manager for the target inference application; and recording events in the target inference application in real time through the tracing manager.
[0154] In an exemplary embodiment, recording events in the target inference application in real time through the tracing manager includes: collecting and recording the following information through the tracing manager: tool usage process information, input content review tracing, knowledge retrieval process tracing, reply generation link information, and conversation message processing tracing; wherein, the tool usage process information includes: the name of the tool called for each interaction, the input and output content, and the execution time; the input content review information includes: the review status of the input text and whether there is sensitive information; the knowledge retrieval process information includes: the query text, the retrieval trigger condition, the matched document content, and the result information; the reply generation link information includes: the intermediate steps in the large language model generation process, the response delay, and the generation result; the conversation message processing information includes: the time points of each stage from the reception, queuing, processing to return of the user message.
[0155] It should be noted that the above steps are executed by the system observability and behavior tracing component, which can improve the observability and behavior traceability of the conversational inference application constructed based on the RAG (Retrieval-Augmented Generation) mechanism during operation, and is used for collecting, scheduling, and persistent processing of the core operations and performance data in the whole process of the inference task. This component is centered around the concept of "application-level tracing context". By dynamically initializing the tracing manager when each inference application is started, it can record and delay-process multiple types of events during system operation without interfering with the main process.
[0156] 1) Application-level tracing context management: The system creates an independent tracing context for each inference application to ensure that each tracing task is bound to the corresponding application instance. Whenever a user starts a conversation task, the system establishes a tracing task channel for it, so that all behavior data flows and is centrally managed in this channel.
[0157] 2) Multi-scenario behavior data collection ability: This component supports automatically identifying and collecting the following core behavior information in the inference process: tool usage process information, input content review tracing, knowledge retrieval process tracing, reply generation link information, and conversation message processing tracing.
[0158] 3) Asynchronous task scheduling mechanism: The system designs a unified task buffer pool and a timed processing scheduling mechanism: all behavior data will first enter the task buffer pool, and the processing process will be automatically triggered in batches by the timed scheduler at set time intervals. The data processing process is executed asynchronously to avoid blocking or performance impact on the main inference process.
[0159] 4) Adaptive configuration ability: The system supports flexibly adjusting the batch processing frequency of tracking tasks and the amount of data per batch through configuration parameters to adapt to different business environments and server performance. For example, in high-concurrency request scenarios, the processing frequency can be increased and the batch interval can be reduced to ensure timely processing and complete recording of tracking data.
[0160] 5) Data persistence and visualization support: All collected tracking data will be uniformly stored on disk after processing to form structured log records. This data can be connected to a log analysis system or a visualization platform for performance evaluation, behavior reproduction, and problem tracing of inference services, and can also be used for model tuning and fault troubleshooting during the R & D phase.
[0161] It should be noted that the tracking and observability management component provides the ability to record the behavior of the inference application from input to generation, which is the key technical basic module to ensure the high stability, high maintainability, and high interpretability of the system.
[0162] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all embodiments. To better understand the above method, the following will describe the above process in conjunction with embodiments, but it is not used to limit the technical solutions of the embodiments of the present application. Specifically:
[0163] To implement a retrieval-augmented generation (RAG) application for multiple users and multiple scenarios, the RAG inference system of the present application provides end-to-end capability support in the links of inference request reception, configuration binding, knowledge retrieval, prompt construction, model generation, etc. The following will describe the key implementation steps in conjunction with the system core components and deployment processes:
[0164] Step 1: Inference application creation and configuration version generation:
[0165] When a user creates a new inference application on the system, the user needs to specify information such as the initial system prompt, associated knowledge base, and model parameters. The system automatically generates a configuration entity (inference application configuration), assigns a version number (config_version), and persists it for storage. For each subsequent configuration change (such as knowledge base update, prompt modification), the system generates a new version to support the configuration consistency backtracking and debugging analysis of historical sessions.
[0166] Step 2: Inference request reception and configuration binding judgment:
[0167] The user triggers an inference request through the front end or API. The system parses the request parameters and extracts the app_id and session ID (session_id). Subsequently, the system makes a configuration binding judgment according to a three-layer priority mechanism:
[0168] 1. If an immediate configuration is explicitly passed in the request, it is directly used;
[0169] 2. If it is a historical conversation turn, use the configuration version bound to the conversation;
[0170] 3. If it is a new conversation, read the current version of the inference application as the configuration bound to this conversation. This binding logic ensures the consistency of the model behavior and effectively avoids semantic drift or context mismatch.
[0171] Step 3: Initialization of conversation and message structure:
[0172] The system determines whether there is a bound conversation record. If not, create a new conversation entity and associate information such as app_id, config_version, and user identity. Subsequently, generate a message record, including the user's current question, request parameters, configuration snapshot, etc., and initialize the system response area. If the request contains attachments (such as documents or images), the system automatically generates an attachment index and binds it to the message for subsequent processing.
[0173] Step 4: Knowledge base retrieval task scheduling and routing selection:
[0174] According to the configured retrieval mode, the system enters the single-library optimization or multi-library concurrent process:
[0175] Single-library optimization mode: The system calls the Function Calling ability through the language model to perform semantic understanding and tool selection on multiple candidate knowledge bases, and initiates a targeted retrieval after selecting the target library.
[0176] Multi-library concurrent mode: The system creates independent threads for each candidate knowledge base to perform retrieval operations concurrently. All retrieval results are uniformly aggregated, then normalized scored and fusion sorted. The final result is uniformly encapsulated as a standardized context object, attached with meta-information such as source identifier, similarity score, paragraph number, etc., and synchronously written into the tracking record system.
[0177] Step 5: Prompt sequence construction and context pruning:
[0178] The system constructs a complete prompt input sequence based on the current conversation binding configuration and retrieval return results, including: system prompt; recent conversation context (pruned according to the Token budget); retrieved paragraph content (injected as user supplementary information); current user input question.
[0179] When the total Token length exceeds the model context window, the system executes a priority pruning strategy to ensure prompt semantic integrity and model compatibility.
[0180] Step 6: Model inference execution and streaming output management:
[0181] The system constructs a standard inference call request according to the configuration, supporting synchronous and asynchronous (SSE) modes. In the streaming mode, the system gradually transmits the response to the front end in segments and automatically records the following after completion: Token consumption; response latency; model version and parameter snapshot; reference paragraph mapping.
[0182] This process is completed by the standard execution component, and integrates exception handling logic to ensure the high availability and fault tolerance of the inference service.
[0183] Step 7: Request tracking and behavior data collection: When each round of inference task is started, the system initializes the tracking context channel and asynchronously records the following behaviors: tool call situation (including time consumption, input and output); input review result; retrieval query path and hit information; model generation link; dialogue message processing flow.
[0184] It should be noted that all tracking data will be batch-processed by the buffer pool scheduling mechanism regularly and stored in the log system or data platform for performance evaluation and problem reproduction.
[0185] Step 8: Response structure encapsulation and persistence: After generating the result output, the system encapsulates the response content, reference context, Token statistics information and generation metadata into a standard response object, writes it into the message record system, and returns it to the user. At the same time, update the last active time of the session to maintain the context state.
[0186] Step 9: Exception handling and status recovery mechanism: The system designs a unified exception handling mechanism, covering scenarios such as request interruption, call failure, and content format error. In the streaming task, the wrapper mechanism is used to automatically release the system status (such as flow control flag, active request count) to avoid resource hanging; in the non-streaming scenario, manual release through the interface is supported. There is also a scheduled task in the background to clean up the expired status every minute to ensure the availability of concurrent quotas and status consistency.
[0187] It should be noted that this application proposes a modular system architecture for RAG inference applications, which integrates key mechanisms such as inference request concurrency control, application configuration version management, context-aware knowledge retrieval, structured prompt construction, and model Q&A execution, and has the following beneficial effects:
[0188] 1. Implement a highly available inference service concurrency control mechanism: This application supports setting the maximum number of concurrent requests according to the inference application dimension through a concurrent request rate limiting mechanism based on Redis atomic operations. Combining the automatic release of streaming tasks and the regular cleaning of abnormal states, it effectively prevents the system from being overloaded due to request accumulation. This mechanism realizes state consistency in a multi-node environment and ensures the stable operation of the system in high-concurrency scenarios.
[0189] 2. Ensure the consistency and semantic stability of multi-round dialogue configuration: This application introduces a version management strategy for reasoning application configuration, and implements dynamic regulation and consistency maintenance of session configuration through a three-layer priority binding mechanism. This mechanism ensures that users use the same version of prompt words and search parameters throughout the entire dialogue process, avoiding model behavior drift and effectively improving the coherence and reliability of the dialogue.
[0190] 3. Improve the accuracy and adaptability of knowledge base retrieval: This application supports language model-assisted knowledge base selection and concurrent scheduling of multiple knowledge sources, and achieves efficient information aggregation across knowledge bases through unified normalization processing and fusion sorting strategies. Combined with Function Calling and ReAct mechanisms, the system can achieve more accurate semantic routing and improve the fit between retrieval and generation tasks.
[0191] 4. Enhance the context-awareness and structural standardization of prompt input: This application realizes the unified integration of historical dialogues, knowledge retrieval paragraphs, and multimodal information in the prompt word construction stage, and provides a token budget trimming mechanism to ensure that the model input structure is standardized and semantically coherent, effectively improve the generation quality and response accuracy, and adapt to multi-round interactions and complex task scenarios.
[0192] 5. Improve system observability and fault traceability: This application has a built-in asynchronous behavior tracking component that supports recording and analyzing key operations in the entire process of reasoning tasks, including request processing, tool calling, model execution, response output, etc. This mechanism provides complete data support for model optimization, system debugging and troubleshooting, and significantly improves the maintainability and explainability of the system.
[0193] 6. System management framework that takes into account both flexibility and governance capabilities: This application supports real-time debugging configuration injection, dynamic adjustment of inference parameters, and management of knowledge base and model versions by application dimension. It takes into account both user customization capabilities and platform governance capabilities, providing a good foundation for large-scale deployment, multi-team collaboration and application iteration, and has broad engineering adaptability and commercial promotion prospects.
[0194] In summary, this application has significant innovation and practical value in terms of structural design, operating mechanism and functional implementation, and provides key support for the large-scale deployment and high-quality service of the retrieval enhancement generation system.
[0195] For a better understanding, the following is a general description of the key technical points of this application. Specifically, this application proposes a retrieval enhancement generation (RAG) reasoning system driven by knowledge base information, which builds a highly available and scalable system architecture around the key links of concurrency control, configuration consistency, knowledge retrieval accuracy, context construction and response generation of reasoning tasks. Its main technical key points include:
[0196] 1. Inference Request Rate Limiting and State Management Mechanism: Through the atomic rate limiting logic implemented at the Redis layer, the system can precisely control the number of concurrent requests for each inference application. Combining the automatic release mechanism for streaming requests and the strategy of periodically cleaning up "zombie requests" effectively ensures the consistency and high availability of the system in a multi-node environment.
[0197] 2. Versioning and Binding Strategy for Inference Application Configuration: Each inference application configuration supports version management. The system adopts a "three-layer priority" strategy (immediate configuration > session binding > application default) to ensure the consistency of model behavior in multi-turn conversations and avoid semantic drift.
[0198] 3. Intelligent Knowledge Base Retrieval Routing Mechanism: Supporting two modes, single knowledge base optimization and multi-knowledge base concurrency, the system can utilize the capabilities of language models (such as Function Calling or ReAct) to assist in selecting the most matching knowledge base and performing retrieval. The retrieval results are uniformly encapsulated and normalized to improve accuracy and system efficiency.
[0199] 4. Prompt Construction and Context Fusion Mechanism: Combining system prompts, historical conversation content, retrieved paragraphs, and the user's current question, a structured prompt sequence is constructed, and a pruning strategy is executed based on the Token budget to improve the semantic integrity of the generated context and the response quality.
[0200] 5. Structured Session and Message Management Mechanism: The system automatically manages the session records and message entities for each round of interaction, including configuration snapshots, input and output content, attachment associations, and Token statistics, providing continuous context-bearing capabilities for multi-round inferences.
[0201] 6. Behavior Tracking and System Observability Component: Full-link tracking of the inference process, covering key data such as tool calls, retrieval execution, generation latency, and Token consumption, which are written to the logging system using an asynchronous scheduling mechanism for performance analysis and fault location.
[0202] In summary, through the integration of the above key technical points, this application realizes the comprehensive optimization of the RAG inference system in aspects such as concurrency control, configuration governance, retrieval scheduling, and generation management, with significant advantages of strong engineering practicability, high deployment adaptability, and good behavior interpretability. The key modules and implementation mechanisms are the main protected contents of the present invention.
[0203] From the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0204] The embodiments of this application also provide a system for generating output results.Figure 6 is a structural block diagram of a system for generating output results according to an embodiment of the present application. As Figure 6 shown, the system includes:
[0205] A knowledge base retrieval routing component 602, which is used to determine a target retrieval strategy according to the inference application configuration corresponding to the inference request when obtaining an inference request for a target inference application. Among them, the inference application configuration has configuration information corresponding to the target inference application; when the target retrieval strategy is a single knowledge base retrieval strategy, determine a target knowledge base from multiple knowledge bases, and use the target knowledge base to determine the retrieval result of the inference content corresponding to the inference request; when the target retrieval strategy is a multi-knowledge base retrieval strategy, use multiple knowledge bases together to determine the retrieval result corresponding to the inference content;
[0206] A prompt word construction and context enhancement component 604, which is used to construct a prompt message queue according to the prompt words corresponding to the inference application configuration, the historical session information corresponding to the inference request, the retrieval result, and the inference content;
[0207] A question and answer execution component 606, which is used to generate an output result corresponding to the inference request through a target large language model according to the prompt message queue.
[0208] The above system solves the deficiency of fixed retrieval strategies in traditional retrieval methods by determining a target retrieval strategy that matches the inference application configuration. Specifically, when the inference request corresponds to a single knowledge base retrieval strategy, it can select the most suitable target knowledge base for accurate retrieval from them, avoiding the retrieval of irrelevant knowledge bases and improving the retrieval efficiency and pertinence. When facing a multi-knowledge base retrieval strategy, this method can use multiple knowledge bases for joint retrieval to ensure more comprehensive information coverage and overcome the information blind spots that may exist in a single knowledge base. Further, combining the retrieval result with the historical session information, system prompt words, and the user's current inference content, a structured prompt message queue is constructed and used as the input to the large language model. This multi-information source fusion strategy enriches the input context of the model, ensures the consistency of the input and the effective utilization of the model, and thus generates a more user-demand-oriented and high-quality output result. That is, the present application provides a more accurate and efficient content generation method by intelligently selecting retrieval strategies, efficiently integrating retrieval results, and optimizing model inputs, effectively solving the problems of retrieval strategy selection and result processing in a multi-knowledge base environment.
[0209] In an exemplary embodiment, the output result generation system further includes: an inference request concurrency management component, configured to determine whether the current concurrent inference request count is greater than a preset threshold when obtaining an inference request for a target inference application, before determining a target retrieval strategy according to the inference application configuration corresponding to the inference request; if the current concurrent inference request count is greater than the preset threshold, not respond to the inference request; if the current concurrent inference request count is less than or equal to the preset threshold, generate a request identifier for the inference request, and perform full-life cycle tracking and management on the inference request according to the request identifier; a knowledge base retrieval routing component 602, configured to determine a retrieval strategy according to the inference application configuration of the target inference application corresponding to the inference request when the current concurrent inference request count is less than or equal to the preset threshold.
[0210] In an exemplary embodiment, the output result generation system further includes: an inference application configuration management component, configured to create an inference application configuration of a target inference application according to the obtained configuration information during the creation of the target inference application, where the inference application configuration includes: prompt words, associated knowledge base sets, retrieval configuration information, the identifier of the large language model, and the model parameters of the large language model; the retrieval configuration information includes a retrieval strategy.
[0211] In an exemplary embodiment, the inference application configuration management component is further configured to, before determining a target retrieval strategy according to the inference application configuration corresponding to the inference request, if the inference request is the first-round dialogue request of a new session of the target inference application and does not carry the inference application configuration, use the latest version of the inference application configuration associated with the target inference application as the inference application configuration corresponding to the inference request; if the inference request is the dialogue request of an existing target session of the target inference application and does not carry the inference application configuration, use the inference application configuration corresponding to the target session as the inference application configuration corresponding to the inference request; if the inference request carries the inference application configuration, use the carried inference application configuration as the inference application configuration corresponding to the inference request.
[0212] In an exemplary embodiment, the knowledge base retrieval routing component 602 is further configured to determine whether the target large language model used by the target inference application supports function call inference ability or chain of thought inference ability; and obtain the retrieval strategy from the inference application configuration; if the target large language model supports function call inference ability or chain of thought inference ability and the retrieval strategy is a single knowledge base retrieval strategy, determine the target retrieval strategy as a single knowledge base retrieval strategy; if the target large language model does not support function call inference ability or chain of thought inference ability, or the retrieval strategy is a multi-knowledge base retrieval strategy, determine the target retrieval strategy as a multi-knowledge base retrieval strategy.
[0213] In an exemplary embodiment, the knowledge base retrieval routing component 602 is further configured to convert each knowledge base in multiple knowledge bases into a structured tool description object, obtaining multiple tool description objects, where each tool description object includes a unique identifier and a usage description of the corresponding knowledge base; and determine a target knowledge base from the multiple knowledge bases according to the multiple tool description objects and the inference content through a target large language model.
[0214] In an exemplary embodiment, the knowledge base retrieval routing component 602 is configured to determine configuration parameters of a target knowledge base, where the configuration parameters include: a retrieval method, an identifier of an embedding model, a similarity scoring threshold, and the number of returned segments; retrieve the target knowledge base using the configuration parameters according to the inference content, and normalize and contextually encapsulate the retrieved content to obtain a retrieval result, where the retrieval result includes: a retrieved segment text, a similarity score of the retrieved segment text, and a position of the retrieved segment text in the target knowledge base.
[0215] In an exemplary embodiment, the knowledge base retrieval routing component 602 is configured to separately determine configuration parameters of each knowledge base in multiple knowledge bases, and retrieve the multiple knowledge bases based on the corresponding configuration parameters to obtain multiple retrieval sub-results, where the configuration parameters include: a retrieval method, an identifier of an embedding model, a similarity scoring threshold, and the number of returned segments; perform fusion sorting and re-screening on the multiple retrieval sub-results, and normalize and contextually encapsulate the results of the fusion sorting and re-screening to obtain a retrieval result, where the retrieval result includes: a retrieved segment text, a similarity score of the retrieved segment text, the knowledge base corresponding to the retrieved segment text, and a position in the knowledge base.
[0216] In an exemplary embodiment, the prompt construction and context enhancement component 604 is further configured to screen out target segments from the historical session information whose relevance to the inference content is higher than a preset relevance threshold; in the case where the total number of tokens of the prompt, the target segments, the retrieval result, and the inference content is greater than a preset number, crop the target segments and the retrieval result according to the preset number, and construct a prompt message queue according to the cropped target segments, the retrieval result, the prompt, and the inference content; in the case where the total number of tokens of the prompt, the target segments, the retrieval result, and the inference content is less than or equal to the preset number, directly construct a prompt message queue according to the target segments, the retrieval result, the prompt, and the inference content.
[0217] In an exemplary embodiment, the question and answer execution component 606 is further configured to construct a standardized request according to the model parameters of the target large language model in the inference application configuration corresponding to the inference request and the prompt message queue, where the standardized request adapts to the interface protocol of the target large language model; and call the target large language model according to the standardized request to obtain an output result of the target large language model.
[0218] In an exemplary embodiment, the Q&A execution component 606 is further configured to, after generating an output result corresponding to an inference request based on a prompt message queue through a target large language model, encapsulate the output result into a structured output object, where the structured output object includes: response content, reference context, token consumption, and meta information of the generation task; the meta information includes: the version of the target large language model, execution time, and status information; display the structured output object, and record the structured output object in the session record corresponding to the inference request.
[0219] In an exemplary embodiment, the output result generation system further includes: a session and message management component, configured to create a session record for an inference request in the case where the inference request is the first-round dialogue request of a new session of a target inference application; and update the active time of the session record corresponding to the target session in the case where the inference request is a dialogue request of an existing target session of the target inference application.
[0220] In an exemplary embodiment, the session and message management component is further configured to generate a message record for each interaction during the interaction in the session, where the message record includes: interaction input, the inference application configuration corresponding to the interaction, the relationship between the interaction and the previous interaction, and a placeholder for the preset response area; record the message record in the session record of the session.
[0221] In an exemplary embodiment, the session and message management component is further configured to generate an attachment index for the attachment file in the case where the attachment file is carried in the inference request, and bind the attachment index to the corresponding message record.
[0222] In an exemplary embodiment, the output result generation system further includes: a system observability and behavior tracking component, configured to initialize a tracking manager for the target inference application in the case where the target inference application is started; and record events in the target inference application in real time through the tracking manager.
[0223] In an exemplary embodiment, the system observability and behavior tracking component is further configured to collect and record the following information through the tracking manager: tool usage process information, input content review tracking, knowledge retrieval process tracking, reply generation link information, and dialogue message processing tracking; where the tool usage process information includes: the name of the tool called for each interaction, input and output content, and execution time; the input content review information includes: the review status of the input text and whether there is sensitive information; the knowledge retrieval process information includes: query text, retrieval trigger conditions, matching document content, and result information; the reply generation link information includes: intermediate steps, response latency, and generation results during the generation process of the large language model; the dialogue message processing information includes: time points of each stage from receiving, queuing, processing to returning of the user message.
[0224] For the description of the features in the embodiments corresponding to the output result generation system, reference can be made to the relevant descriptions in the embodiments corresponding to the output result generation method, which will not be elaborated here one by one.
[0225] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above embodiments of the output result generation method.
[0226] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above embodiments of the output result generation method when running.
[0227] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0228] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above embodiments of the output result generation method.
[0229] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above embodiments of the output result generation method.
[0230] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0231] The above has introduced in detail a method and device for generating an output result, an electronic device, a storage medium, and a computer program product provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can still be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for generating an output result, characterized in that, Including: When a reasoning request for a target reasoning application is obtained, determining a target retrieval strategy according to the reasoning application configuration corresponding to the reasoning request, where the reasoning application configuration has configuration information corresponding to the target reasoning application; When the target retrieval strategy is a single knowledge base retrieval strategy, determining a target knowledge base from multiple knowledge bases and using the target knowledge base to determine the retrieval result of the reasoning content corresponding to the reasoning request; when the target retrieval strategy is a multi-knowledge base retrieval strategy, using the multiple knowledge bases to jointly determine the retrieval result corresponding to the reasoning content; Constructing a prompt message queue according to the prompt words corresponding to the reasoning application configuration, the historical session information corresponding to the reasoning request, the retrieval result, and the reasoning content; Generating an output result corresponding to the reasoning request through a target large language model according to the prompt message queue.
2. The method for generating the output result according to claim 1, wherein Before determining the target retrieval strategy according to the reasoning application configuration corresponding to the reasoning request, the method further includes: When a reasoning request for a target reasoning application is obtained, determining whether the current concurrent reasoning request number is greater than a preset threshold; when the current concurrent reasoning request number is greater than the preset threshold, not responding to the reasoning request; when the current concurrent reasoning request number is less than or equal to the preset threshold, generating a request identifier for the reasoning request and performing full-life cycle tracking and management on the reasoning request according to the request identifier; Determining the target retrieval strategy according to the reasoning application configuration corresponding to the reasoning request includes: when the current concurrent reasoning request number is less than or equal to the preset threshold, determining the retrieval strategy according to the reasoning application configuration of the target reasoning application corresponding to the reasoning request.
3. The method for generating the output result according to claim 1, wherein The method further includes: During the creation of the target reasoning application, creating a reasoning application configuration of the target reasoning application according to the obtained configuration information, where the reasoning application configuration includes: prompt words, an associated knowledge base set, retrieval configuration information, an identifier of the large language model, and model parameters of the large language model; the retrieval configuration information includes a retrieval strategy.
4. The method for generating the output result according to claim 1, wherein Before determining the target retrieval strategy according to the reasoning application configuration corresponding to the reasoning request, the method further includes: When the reasoning request is the first-round dialogue request of a new session of the target reasoning application and does not carry a reasoning application configuration, using the latest version of the reasoning application configuration associated with the target reasoning application as the reasoning application configuration corresponding to the reasoning request; When the reasoning request is a dialogue request for an existing target session of the target reasoning application and does not carry a reasoning application configuration, using the reasoning application configuration corresponding to the target session as the reasoning application configuration corresponding to the reasoning request; When the reasoning request carries a reasoning application configuration, using the carried reasoning application configuration as the reasoning application configuration corresponding to the reasoning request.
5. The method for generating the output result according to claim 1, wherein Determining the target retrieval strategy according to the reasoning application configuration corresponding to the reasoning request includes: Determine whether the target large language model used by the target inference application supports function call inference ability or chain of thought inference ability; and obtain the retrieval strategy from the inference application configuration; In the case where the target large language model supports the function call inference ability or the chain of thought inference ability, and the retrieval strategy is a single knowledge base retrieval strategy, determine that the target retrieval strategy is the single knowledge base retrieval strategy; In the case where the target large language model does not support the function call inference ability or the chain of thought inference ability, or the retrieval strategy is a multi-knowledge base retrieval strategy, determine that the target retrieval strategy is the multi-knowledge base retrieval strategy.
6. The method for generating the output result according to claim 1, wherein, Determine the target knowledge base from multiple knowledge bases, including: Convert each knowledge base in the multiple knowledge bases into a structured tool description object to obtain multiple tool description objects, where each tool description object contains the unique identifier and usage description of the corresponding knowledge base; Through the target large language model, determine the target knowledge base from the multiple knowledge bases according to the multiple tool description objects and the inference content.
7. The method for generating the output result according to claim 1, wherein Use the target knowledge base to determine the retrieval result of the inference content corresponding to the inference request, including: Determine the configuration parameters of the target knowledge base, where the configuration parameters include: retrieval method, identifier of the embedding model, similarity scoring threshold, number of returned segments; According to the inference content, use the configuration parameters to retrieve the target knowledge base, and normalize and encapsulate the retrieved content to obtain the retrieval result, where the retrieval result includes: retrieved segment text, similarity score of the retrieved segment text, and position of the retrieved segment text in the target knowledge base.
8. The method for generating the output result according to claim 1, wherein Use the multiple knowledge bases to jointly determine the retrieval result corresponding to the inference content, including: Respectively determine the configuration parameters of each knowledge base in the multiple knowledge bases, and retrieve the multiple knowledge bases based on the corresponding configuration parameters to obtain multiple retrieval sub-results, where the configuration parameters include: retrieval method, identifier of the embedding model, similarity scoring threshold, number of returned segments; Perform fusion sorting and re-screening on the multiple retrieval sub-results, and normalize and encapsulate the results of the fusion sorting and re-screening to obtain the retrieval result, where the retrieval result includes: retrieved segment text, similarity score of the retrieved segment text, corresponding knowledge base of the retrieved segment text, and position in the knowledge base.
9. The method for generating the output result according to claim 1, wherein Construct a prompt message queue according to the prompt words corresponding to the inference application configuration, the historical session information corresponding to the inference request, the retrieval result, and the inference content, including: Filter out target segments from the historical session information whose relevance to the inference content is higher than a preset relevance threshold; In the case where the total number of tokens of the prompt words, the target segments, the retrieval result, and the inference content is greater than a preset number, crop the target segments and the retrieval result according to the preset number, and construct a prompt message queue according to the cropped target segments, the retrieval result, the prompt words, and the inference content; When the total number of tokens of the prompt, the target segment, the retrieval result, and the reasoning content is less than or equal to the preset number, directly construct a prompt message queue according to the target segment, the retrieval result, the prompt, and the reasoning content.
10. The method for generating the output result according to claim 1, wherein Generate the output result corresponding to the reasoning request through the target large language model, including: Construct a standardized request according to the model parameters of the target large language model in the reasoning application configuration corresponding to the reasoning request and the prompt message queue, where the standardized request adapts to the interface protocol of the target large language model; Call the target large language model according to the standardized request to obtain the output result of the target large language model.
11. The method for generating the output result according to claim 1, wherein After generating the output result corresponding to the reasoning request through the target large language model according to the prompt message queue, the method further includes: Package the output result into a structured output object, where the structured output object includes: response content, reference context, token consumption situation, and meta information of the generation task; the meta information includes: the version, execution time, and status information of the target large language model; Display the structured output object and record the structured output object in the session record corresponding to the reasoning request.
12. The method for generating the output result according to claim 1, wherein The method further includes: When the reasoning request is the first-round dialogue request of a new session of the target reasoning application, create a session record for the reasoning request; when the reasoning request is a dialogue request for an existing target session of the target reasoning application, update the active time of the session record corresponding to the target session.
13. The method for generating the output result according to claim 12, wherein The method further includes: During the interaction in the session, generate a message record for each interaction, where the message record includes: interaction input, the reasoning application configuration corresponding to the interaction, the relationship between the interaction and the previous interaction, and a placeholder for the preset response area; Record the message record in the session record of the session.
14. The method for generating the output result according to claim 13, wherein The method further includes: When an attachment file is carried in the reasoning request, generate an attachment index for the attachment file and bind the attachment index to the corresponding message record.
15. The method for generating the output result according to claim 1, wherein The method further includes: When the target reasoning application is started, initialize a tracing manager for the target reasoning application; Record the events in the target reasoning application in real time through the tracing manager.
16. The method for generating the output result according to claim 15, wherein, Recording the events in the target reasoning application in real time through the tracing manager includes: Collect and record the following information through the tracing manager: tool usage process information, input content review tracing, knowledge retrieval process tracing, response generation link information, and dialogue message processing tracing; Among them, the tool usage process information includes: the name of the tool called for each interaction, the input and output content, and the execution time; the input content review information includes: the review status of the input text and whether there is sensitive information; the knowledge retrieval process information includes: the query text, the retrieval trigger condition, the matched document content, and the result information; the response generation link information includes: the intermediate steps, response latency, and generation result during the generation process of the large language model; the conversation message processing information includes: the time points of each stage from receiving, queuing, processing to returning the user message.
17. A system for generating output results, characterized in that, including: A knowledge base retrieval routing component, configured to, when obtaining an inference request corresponding to a target inference application, determine a target retrieval strategy according to the inference application configuration corresponding to the inference request, where the inference application configuration has configuration information corresponding to the target inference application; when the target retrieval strategy is a single knowledge base retrieval strategy, determine a target knowledge base from multiple knowledge bases, and use the target knowledge base to determine the retrieval result of the inference content corresponding to the inference request; when the target retrieval strategy is a multi-knowledge base retrieval strategy, use the multiple knowledge bases to jointly determine the retrieval result corresponding to the inference content; A prompt construction and context enhancement component, configured to construct a prompt message queue according to the prompt corresponding to the inference application configuration, the historical session information corresponding to the inference request, the retrieval result, and the inference content; A question and answer execution component, configured to generate an output result corresponding to the inference request through a target large language model according to the prompt message queue.
18. An electronic device, characterized in that, including: A memory, configured to store a computer program; A processor, configured to implement the steps of the output result generation method according to any one of claims 1 to 16 when executing the computer program.
19. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program implements the steps of the output result generation method according to any one of claims 1 to 16 when executed by a processor.
20. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the output result generation method according to any one of claims 1 to 16 when executed by a processor.
Citation Information
Patent Citations
Dynamic multilevel retrieval enhanced knowledge base question and answer method and device, medium and equipment
CN118689976A
Information retrieval method and device, storage medium and computer program product
CN118939763A
Question answering system construction method based on large language model and question answering system
CN118964587A
Multi-knowledge-base retrieval and answer integration optimization method, retrieval enhancement generation system, equipment and medium
CN119312896A
Knowledge base retrieval method and system applied to large language model
CN119396946A
Cited By
Conversation message queue management system based on large model cluster
CN121187825A
A dialogue message queue management system based on a large model cluster
CN121187825B
Unified streaming processing method, system and device for multi-mode AI interactive content, medium and program product
CN121705057A
Techniques and architecture for securing large language model assisted interactions with a data catalog
US20250348480A1