A multi-model service-oriented intermediate result reuse and inference chain acceleration method
Patent Information
- Application Number
- CN202610912054.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-25
AI Technical Summary
[0008]针对现有方法中存在的缺陷,具体为现有技术在中间结果复用方面仍存在以下不足:缺乏对推理链内部阶段性结果的复用机制,大量重复计算无法有效避免;缺少针对阶段性推理结果的统一结构化描述与语义建模能力;无法实现跨任务、跨模型以及跨推理阶段的动态结果共享;难以支持复杂Agent协同、多阶段工作流及动态推理路径场景;无法根据已有阶段结果进行增量推理与局部跳转加速
[0026]本发明通过推理链阶段级中间结果复用机制,使系统能够对Embedding结果、检索结果、推理规划结果、工具调用结果、上下文摘要及阶段状态信息等阶段性结果进行统一存储与复用,避免完整推理链重复执行,显著降低系统资源消耗。
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and computer technology, and in particular to a method for reusing intermediate results and accelerating inference chains for multi-model services. Background Technology
[0002] With the development of artificial intelligence technology, large language models, multi-agent systems, and retrieval-enhanced generation technologies are widely used in scenarios such as intelligent question answering, code generation, complex task planning, and automated decision-making. To improve task processing capabilities, existing systems typically adopt a multi-model collaborative reasoning architecture, which uses multiple model nodes, tool modules, and external knowledge systems to form a reasoning chain to process complex tasks in stages.
[0003] In existing multi-model service systems, a complete task typically involves multiple steps, including task understanding, context building, knowledge retrieval, inference planning, tool invocation, staged inference, and result generation. Numerous intermediate results are generated between these stages, such as embedding vectors, retrieval results, inference states, context summaries, tool execution results, and staged inference content. Most of these intermediate results in the current system exist only temporarily within the lifecycle of a single request and are discarded after the task ends, making them unusable in subsequent requests. Furthermore, in real-world applications, many user requests exhibit high semantic similarity. Even when faced with highly similar requests, existing systems still need to re-execute the entire inference process, resulting in significant redundant computation.
[0004] In existing technologies, performance optimization for large language models and multi-model collaborative inference systems mainly includes the following solutions: One type is a caching solution based on the final result, where the previously generated result is directly returned when the user input is completely identical or similar to the previous request. While this type of solution can reduce the computational overhead caused by some duplicate requests, its cached objects are mainly the final output results, and it cannot achieve the reuse of results within the inference chain. Another type is an inference acceleration solution based on Key-ValueCache, which caches the Attention calculation results of historical tokens to reduce redundant calculations in the autoregressive generation stage. This type of solution mainly works for token-level acceleration in a single model and a single inference process, and cannot achieve the sharing of intermediate results across tasks, models, and inference stages.
[0005] Another type is the retrieval caching scheme based on retrieval enhancement, which caches embedding results, vector retrieval results, or knowledge fragments to reduce the overhead of repeated retrieval. This type of scheme only performs local optimizations for the knowledge retrieval stage and cannot cover the complete inference chain process, including inference planning, tool invocation, multi-stage agent collaboration, and complex task execution.
[0006] Furthermore, some multi-agent systems introduce short-term and long-term memory mechanisms to store historical dialogues, tool call records, and some inference information. However, existing agent memory mechanisms are mainly used to enhance contextual continuity and dialogue consistency, with their core goal being to improve the interactive experience rather than achieving computational reuse at the inference chain level. In multi-model workflow systems, some frameworks use DAG or Pipeline approaches to orchestrate task flows and optimize execution order through node dependency management. However, these solutions primarily address task scheduling issues rather than the reuse of intermediate results.
[0007] Therefore, a method for reusing intermediate results and accelerating inference chains for multi-model services is essential. Summary of the Invention
[0008] To address the shortcomings of existing methods, specifically in the reuse of intermediate results, the following deficiencies exist: a lack of mechanisms for reusing staged results within the inference chain, making it difficult to effectively avoid extensive repetitive computations; a lack of unified structured descriptions and semantic modeling capabilities for staged inference results; inability to achieve dynamic result sharing across tasks, models, and inference stages; difficulty in supporting complex agent collaboration, multi-stage workflows, and dynamic inference path scenarios; and inability to accelerate incremental inference and local jumps based on existing stage results. This paper proposes a method for reusing intermediate results and accelerating the inference chain for multi-model services.
[0009] The purpose of this invention is to provide a method for reusing intermediate results and accelerating the inference chain for multi-model services. This method can uniformly manage and dynamically reuse intermediate results in the inference chain, thereby reducing the overhead of repeated inference computation and improving inference efficiency and system response speed.
[0010] The present invention also aims to provide a method for reusing intermediate results and accelerating inference chains for multi-model services, which can also achieve the reuse of stage results between non-identical tasks through semantic matching and inference dependency analysis.
[0011] The present invention also aims to provide a method for reusing intermediate results and accelerating inference chains for multi-model services, which can also achieve local incremental inference by constructing an inference chain dependency graph, thereby avoiding repeated execution of the complete inference chain.
[0012] In a first aspect of the present invention, a method for reusing intermediate results and accelerating inference chains for multi-model services is provided, the method comprising the steps of:
[0013] Receive user requests, parse the user requests for tasks, and construct an inference chain containing multiple stage nodes; during the execution of the inference chain, generate stage feature description information for the current stage node;
[0014] Based on the stage feature description information, historical intermediate results are retrieved from the intermediate result repository;
[0015] Make reuse decisions based on the search results;
[0016] When the reuse is determined to be successful, the historical intermediate results are loaded and the inference execution of the current stage node is skipped;
[0017] When reuse fails, the inference of the current stage node is executed, and the current intermediate result generated by the execution is stored in the intermediate result repository.
[0018] In a second aspect of the invention, a system for reusing intermediate results and accelerating inference chains for multi-model services is provided, comprising:
[0019] The user request processing module is used to receive user requests and parse tasks.
[0020] The inference chain building module is used to build an inference chain containing multiple stage nodes based on the task parsing results;
[0021] The intermediate results management module is used to store and manage historical intermediate results;
[0022] The semantic matching module is used to perform retrieval in the intermediate result repository based on stage feature description information;
[0023] The inference chain scheduling module is used to make reuse decisions based on the retrieval results and schedule the execution of the inference chain.
[0024] In a third aspect of the invention, a computer device is provided, including a processor, a memory, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method described in the first aspect.
[0025] In a fourth aspect of the invention, a computer-readable storage medium is provided having a computer program / instructions stored thereon, characterized in that the computer program / instructions, when executed by a processor, implement the steps of the method described in the first aspect.
[0026] This invention enables the system to uniformly store and reuse stage-level intermediate results such as embedding results, retrieval results, inference planning results, tool call results, context summaries, and stage status information through a stage-level intermediate result reuse mechanism in the inference chain. This avoids repeated execution of the complete inference chain and significantly reduces system resource consumption.
[0027] In the intermediate result matching stage, a dynamic matching mechanism based on semantic vectors and multi-dimensional constraints is introduced. By jointly judging multiple dimensions such as semantic similarity, stage type, context consistency, dependency relationship, model compatibility and time validity, stage-level semantic reuse under different tasks can be realized, which can effectively improve cache hit rate and reuse flexibility.
[0028] By constructing an inference chain dependency graph and managing the data dependencies between nodes at each stage, the system can automatically identify affected nodes when the task changes, perform local incremental inference only on the changed parts, avoid recalculating the entire chain, and significantly reduce the overall inference latency in complex tasks and long-chain inference scenarios.
[0029] By using a unified intermediate result abstraction structure, cross-model stage results sharing is achieved, enabling different models (including general large language models, mathematical reasoning models, code generation models, multimodal models, and retrieval enhancement models) to share stage reasoning results, effectively improving the collaborative efficiency of heterogeneous models.
[0030] A dynamic lifecycle management system is established for historical intermediate results. Based on factors such as access frequency, reuse count, generation cost, and time validity, hot caching retention, cold data archiving, result compression, and invalidation are dynamically executed to improve system efficiency and storage resource utilization in long-term operation.
[0031] In high-concurrency scenarios, multiple similar tasks can share the results of ongoing or completed stages, avoiding multiple requests from repeatedly occupying GPU and retrieval resources, thereby significantly improving system throughput and resource utilization efficiency.
[0032] This invention focuses on stage-level reuse of large model inference chains, organically combining semantic matching, dependency graph management, and incremental inference to achieve dynamic acceleration of the inference chain in complex multi-model service scenarios, effectively improving inference efficiency, system throughput, and resource utilization. Detailed Implementation
[0033] The more detailed description of embodiments of the invention below is not intended to limit the scope of the claimed invention, but is merely illustrative and does not limit the description of the features and characteristics of the invention, in order to suggest the best mode for carrying out the invention and to enable those skilled in the art to practice the invention. However, it should be understood that various modifications and variations can be made without departing from the scope of the invention as defined by the appended claims. The detailed description should be considered illustrative only and not restrictive, and any such modifications and variations shall fall within the scope of the invention described herein. Furthermore, the background art is intended to illustrate the current state of research and development and significance of the technology, and is not intended to limit the invention or the scope of application of this application.
[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to limit the invention.
[0035] The present invention will be further described below with reference to specific embodiments.
[0036] A preferred embodiment of the method for reusing intermediate results and accelerating inference chains for multi-model services according to the present invention is as follows: The system receives user requests and performs task parsing. After receiving a user input request, the system performs semantic parsing, task identification, and contextual analysis on the request content, extracting task objectives, input parameters, task types, and contextual constraint information. The user request can be submitted through a web input box, a speech recognition interface, or an API platform integration entry. Further, the system can break down complex tasks into subtasks based on a large language model, a task classification model, or a workflow scheduling module, and generate corresponding task execution plans. For example, if the user request is "Please summarize the reasons for the abnormal sales of Company A's product in the past three months and provide optimization suggestions," the system performs semantic parsing on the request, identifies the task type as "abnormality analysis and suggestion generation," and extracts the analysis object "Company A's product," the time range "the past three months," and the expected output "summary of abnormality reasons and optimization suggestions." The system then constructs a multi-stage inference chain. The system dynamically constructs the inference chain structure according to the task execution plan. The inference chain includes multiple stage nodes, each stage node corresponding to at least one processing task. The processing tasks include, but are not limited to: large language model inference, knowledge retrieval, embedding generation, agent planning, tool invocation, code execution, data query, context compression, and result aggregation. Simultaneously, the system establishes dependencies between nodes at each stage, forming an inference chain dependency graph.
[0037] For the above user request, the system generates the following inference chain:
[0038] (1) Task understanding stage; (2) Embedding generation stage; (3) Vector knowledge retrieval stage; (4) Sales data analysis stage; (5) Reasoning and planning stage; (6) Optimization suggestion generation stage; (7) Result summary and output stage. There are clear dependencies between the stages. For example, the knowledge retrieval stage depends on the output of the Embedding generation stage, and the reasoning and planning stage depends on the output of the knowledge retrieval stage and the data analysis stage.
[0039] During the generation of stage feature description information, the system generates corresponding stage feature description information for each stage node during inference chain execution. This stage feature description information includes, but is not limited to: current stage input content, input semantic vector, stage task type, context summary, belonging inference path, identifiers of preceding dependent nodes, model type information, tool call status, timestamp information, and historical task association identifiers. The system further vectorizes the stage input and context information to generate a stage semantic index.
[0040] For example, in the embedding generation stage, the system generates semantic vectors for user input; in the knowledge retrieval stage, the system generates semantic features for retrieval queries; and in the reasoning planning stage, the system generates semantic descriptions of reasoning objectives.
[0041] The system retrieves historical intermediate results. Based on the feature description information of the current stage, it performs a historical result retrieval in the intermediate result repository. The retrieval process includes: semantic vector-based similarity retrieval, stage type-based filtering and matching, task dependency-based association matching, context consistency-based constraint verification, model compatibility-based adaptation verification, and time validity-based result filtering. The system can use a vector database, graph database, or hybrid index structure to store intermediate result information.
[0042] For example, in the current embedding generation stage, the system searches the intermediate result repository for semantically similar embedding results; in the current knowledge retrieval stage, the system searches for similar retrieval results; and in the current inference planning stage, the system searches for planning results with similar inference paths. The system then performs a reuse decision. Based on the matching degree between historical intermediate results and the current stage, the system calculates the reuse confidence score. When the reuse confidence score is higher than a preset threshold, the system determines that the current stage can directly reuse historical results and skips the corresponding inference execution process. When the reuse confidence score is lower than the preset threshold but some reusable information exists, the system enters incremental inference mode. For example, if the system detects that the historical task "analyze the reasons for the recent decline in sales of Company A's product" has a high semantic similarity to the current task, and the reuse confidence score of its embedding results and knowledge retrieval results is higher than the preset threshold, then it determines that it can be directly reused.
[0043] Execution phase result reuse. For phase nodes that meet the reuse conditions, the system directly loads the corresponding historical phase results and writes them into the current inference chain context. The reused results include, but are not limited to: retrieval results, inference summaries, tool call outputs, phase planning results, embedding results, context compression results, and phase status information. Simultaneously, the system updates the execution status of the current inference chain. Local incremental inference is then performed.
[0044] For stage nodes that cannot be fully reused, the system performs local incremental inference only on the differing parts based on the reused stage results. The system automatically identifies affected nodes based on the inference chain dependency graph and generates local inference sub-chains. Furthermore, the system can reduce redundant computations by using methods such as difference context splicing, local prompt reconstruction, subtask re-execution, local retrieval updates, or dynamic parameter adjustments.
[0045] For example, when the time range in a user's request changes from "the past three months" to "the past six months," the system determines that the results of the Embedding generation and knowledge retrieval stages remain valid. Only incremental inference needs to be performed for the newly added time range data in the sales data analysis and inference planning stages, thus avoiding the need to re-execute the complete inference chain. This dynamically optimizes the inference execution path.
[0046] During the execution of the inference chain, the system dynamically optimizes the execution path based on the stage reuse status, model load, and task complexity. This dynamic optimization includes, but is not limited to: skipping reused nodes, merging duplicate stages, pruning redundant inference paths, adjusting the stage execution order, performing parallel inference scheduling, and dynamically switching inference models, thereby reducing overall inference latency.
[0047] After completing the inference chain, the system aggregates, organizes, and post-processes the results from each stage to generate the final task output, which is then returned to the user.
[0048] The system updates the intermediate results repository. It writes the interim results generated during this inference process into the intermediate results repository. The stored content includes, but is not limited to: interim inputs, interim outputs, interim semantic vectors, context summaries, tool call results, inference state, dependency information, model execution parameters, and result validity information. Simultaneously, the system establishes an association index between interim results to support rapid matching and reuse in subsequent tasks.
[0049] Furthermore, the system also includes lifecycle management of intermediate results. Based on access frequency, reuse count, result timeliness, and storage resource availability, the system dynamically manages historical intermediate results, including result updates, result eviction, result compression, retention of hot results, and archiving of cold data, to ensure storage efficiency and retrieval performance during long-term system operation.
[0050] Example 1: Implementation of intermediate result reuse based on RAG and large language model collaboration. This example is applied to an enterprise knowledge question answering system. The system consists of a user request module, a task orchestration module, an embedding module, a vector retrieval module, a large language model inference module, an intermediate result management module, and an inference chain scheduling module.
[0051] When a user initiates a question request, the system first performs semantic parsing on the request content and then constructs an inference chain through the task orchestration module.
[0052] The user request is: "Please summarize the reasons for the abnormal sales of Company A's product in the past three months and provide optimization suggestions." The system generates the following reasoning chain: (1) Task understanding stage; (2) Embedding generation stage; (3) Vector knowledge retrieval stage; (4) Sales data analysis stage; (5) Reasoning planning stage; (6) Optimization suggestion generation stage; (7) Result summary and output stage.
[0053] During the embedding generation phase, the system vectorizes the user input and context, generating corresponding semantic vectors. The system then retrieves results from previous phases from the intermediate results repository to determine if semantically similar embedding or retrieval results exist. If the system detects that the historical task "Analyze the reasons for the recent decline in sales of Company A's product" has a high semantic similarity to the current task, the system directly reuses its corresponding embedding results and some knowledge retrieval results without re-executing the embedding generation and vector retrieval operations.
[0054] Furthermore, during the inference planning phase, the system identifies similar inference paths between the current task and historical analysis tasks through stage semantic matching. Therefore, it directly reuses some data analysis conclusions and anomaly classification results from historical stages, performing only local incremental analysis for newly added time ranges. Ultimately, the system only performs supplementary inference on discrepancies, thereby reducing the repeated execution of the complete inference chain and improving overall response speed.
[0055] Example 2: Implementation of Inference Chain Reuse Based on Multi-Agent Collaborative System:
[0056] This embodiment is applied to a multi-agent automated task processing system. The system includes: a master control agent, a retrieval agent, a planning agent, a tool invocation agent, a summarization agent, and an intermediate result reuse module.
[0057] When a user submits a complex task, the master agent breaks the task down into multiple sub-tasks, which are then completed collaboratively by different agents.
[0058] The user's request was: "Analyze recent market changes in a certain industry and generate an investment risk assessment report." The system first generates market data retrieval tasks, industry news analysis tasks, risk assessment tasks, and report generation tasks.
[0059] During execution, each agent generates corresponding intermediate results, including: industry news summaries, market data analysis results, risk scoring results, tool call outputs, and inference planning paths. The system writes these results into an intermediate results repository and establishes a stage dependency index.
[0060] When a user requests "analysis of the industry's risk trends for the next six months" again, the system finds through stage semantic matching that the news analysis results and market data analysis results in the historical task are still valid. Therefore, the corresponding stage results are directly reused, and only the stage related to future trend prediction is re-executed.
[0061] Furthermore, in high-concurrency scenarios, when multiple users simultaneously request similar industry analysis tasks, the system allows multiple agents to share the results of intermediate stages being executed, thereby avoiding multiple agents repeatedly executing the same retrieval and analysis process.
[0062] Example 3: Incremental Inference Implementation Based on Dynamic Workflow
[0063] This embodiment is applied to a complex workflow automation system. The system adopts a DAG inference chain structure, with each node corresponding to an inference stage.
[0064] During execution, the system maintains the inference chain dependency graph in real time and records: node inputs, node outputs, context states, dependencies, tool call information, and model execution parameters. When a local change occurs in the input at a certain stage, the system automatically identifies the affected nodes based on the dependency graph and re-executes only the corresponding sub-chain.
[0065] For example, the original task includes A: data retrieval; B: data cleaning; C: statistical analysis; D: risk prediction; and E: report generation. When the user only updates a portion of the data range, the system determines that the results of nodes A and B are still valid, and only re-executes nodes C, D, and E, without needing to re-execute the entire inference chain.
[0066] Example 4: Cross-Model Result Reuse Based on Heterogeneous Model Collaboration: This example is applied to a heterogeneous large-scale model collaborative service platform. The system simultaneously deploys a general-purpose large language model, a mathematical reasoning model, a code generation model, a multimodal model, and a retrieval enhancement model. Due to differences in input and output formats between different models, this invention achieves cross-model stage result sharing by unifying the intermediate result abstraction structure.
[0067] For example, after the general-purpose large language model completes task planning, it generates a structured reasoning plan. The mathematical reasoning model directly reuses the stage goals and context summaries in this reasoning plan without re-parsing the user's original request. The code generation model further reuses the structured computation results output by the mathematical model, executing only the code implementation stage. The system achieves the flow and sharing of stage results between different models through a unified semantic description and stage indexing mechanism.
[0068] Example 5: Intermediate Result Lifecycle Management Implementation: In this example, the system performs dynamic lifecycle management for historical intermediate results. The system calculates the result value score based on the following indicators: access frequency, most recent reuse time, reuse success rate, task relevance, result generation cost, and timeliness. For frequently reused results, the system retains their hot cache state; for low-frequency results, the system performs compressed storage or cold archiving; for invalid results, the system automatically performs eviction processing. Furthermore, the system can establish a hierarchical caching mechanism based on task domains to achieve dynamic resource optimization under different business scenarios.
[0069] Through the above implementation methods, the present invention can achieve the reuse of intermediate results in stages, local incremental reasoning, and acceleration of dynamic reasoning chains in scenarios such as multi-model collaboration, complex inference chains, agent systems, and dynamic workflows, thereby effectively reducing the overhead of repeated reasoning and improving system response efficiency and resource utilization.
[0070] In another embodiment, a computer device is provided, including a processor, a memory, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method described above.
[0071] In another embodiment, a computer-readable storage medium is provided that stores a computer program / instructions thereon, characterized in that the computer program / instructions, when executed by a processor, implement the steps of the method described above.
[0072] The various embodiments described herein, or their specific features, structures, or characteristics, may be suitably combined in one or more embodiments of the invention. Furthermore, in some cases, the order of steps described in the flowcharts and / or pipeline processes may be modified where appropriate, and they need not necessarily be performed in the exact order described. Additionally, various aspects of the invention may be implemented using software, hardware, firmware, or combinations thereof, and / or other computer-implemented modules or devices that perform the described functions. Software implementations of the invention may include executable code stored in a computer-readable medium and executed by one or more processors. Computer-readable media may include computer hard disk drives, ROM, RAM, flash memory, portable computer storage media such as CD-ROM, DVD-ROM, flash drives, and / or other devices having a Universal Serial Bus (USB) interface, and / or any other suitable tangible or non-transitory computer-readable medium or computer memory on which executable code can be stored and executed by a processor. The invention may be used in conjunction with any suitable operating system.
[0073] This invention achieves the following through a multi-stage evidence credibility assessment and citation constraint mechanism: reducing the probability of generating illusions in large models; improving the authenticity of generated results; enhancing the interpretability of generated content; improving the stability of complex reasoning scenarios; establishing a traceable relationship between generated content and evidence; and improving the security and engineering deployability of large model applications in high-credibility scenarios.
[0074] The present invention has been described in detail above with reference to specific exemplary embodiments. However, it should be understood that various modifications and variations can be made without departing from the scope of the invention as defined by the appended claims. The detailed description should be considered illustrative only and not restrictive, and any such modifications and variations shall fall within the scope of the invention described herein. Furthermore, the background art is intended to illustrate the current state of development and significance of the technology and is not intended to limit the present invention or the scope of application of the present application.
[0075] More specifically, although exemplary embodiments of the invention have been described herein, the invention is not limited to these embodiments, but includes any and all embodiments modified, omitted, such as combinations between various embodiments, adaptive changes, and / or substitutions, as would be apparent to those skilled in the art from the foregoing detailed description. The limitations in the claims are to be interpreted broadly as used in the language of the claims and are not limited to the examples described in the foregoing detailed description or during the implementation of this application, which should be considered non-exclusive. Any step listed in any method or process claim may be performed in any order and is not limited to the order set forth in the claims. Therefore, the scope of the invention should be determined solely by the appended claims and their legal equivalents, and not by the description and examples given above.
Claims
1. A method for reusing intermediate results and accelerating inference chains for multi-model services, characterized in that, The method includes the following steps: Receive user requests, parse the user requests for tasks, and construct an inference chain containing multiple stage nodes; During the execution of the inference chain, stage feature description information is generated for the current stage node; Based on the stage feature description information, historical intermediate results are retrieved from the intermediate result repository; Make reuse decisions based on the search results; When the reuse is determined to be successful, the historical intermediate results are loaded and the inference execution of the current stage node is skipped; When reuse fails, the inference of the current stage node is executed, and the current intermediate result generated by the execution is stored in the intermediate result repository.
2. The method for reusing intermediate results and accelerating inference chains for multi-model services according to claim 1, characterized in that, The stage feature description information includes one or more of the following: stage input content, stage semantic vector, stage task type, context summary, inference path, identifier of preceding dependent nodes, model type information, and tool call status.
3. The method for reusing intermediate results and accelerating inference chains for multi-model services according to claim 1, characterized in that, The retrieval of historical intermediate results in the intermediate results repository includes at least one of the following retrieval methods: Similarity retrieval based on semantic vectors; Stage-based filtering and matching; Association matching based on task dependencies; Constraint validation based on context consistency; Model compatibility-based adaptation verification; Results selection based on time validity.
4. The method for reusing intermediate results and accelerating inference chains for multi-model services according to claim 1, characterized in that, The process of making reuse decisions based on search results includes: Calculate the reuse confidence of the historical intermediate results and the current stage node; When the reuse confidence level is higher than a preset threshold, the reuse is determined to be successful; When the reuse confidence level is lower than a preset threshold but some reusable information exists, it is determined to enter the incremental inference mode.
5. The method for reusing intermediate results and accelerating inference chains for multi-model services according to claim 1, characterized in that, The construction of the inference chain, which includes multiple stage nodes, includes: The inference chain structure is dynamically generated based on the task execution plan; The processing tasks corresponding to the stage nodes include at least one of the following: large language model reasoning, knowledge retrieval, embedding generation, agent planning, tool invocation, code execution, data query, context compression, and result aggregation; Establish the dependencies between nodes at each stage to form a reasoning chain dependency graph.
6. The method for reusing intermediate results and accelerating inference chains for multi-model services according to claim 5, characterized in that, The affected nodes are automatically identified based on the inference chain dependency graph, and local incremental inference is performed only on the affected nodes.
7. The method for reusing intermediate results and accelerating inference chains for multi-model services according to claim 1, characterized in that, Also includes: The execution path of the inference chain is dynamically optimized based on the reuse status of intermediate results; The dynamic optimization includes at least one of the following: skipping reused nodes, merging duplicate stages, trimming redundant inference paths, adjusting the execution order of stages, performing parallel inference scheduling, and dynamically switching inference models.
8. The method for reusing intermediate results and accelerating inference chains for multi-model services according to claim 1, characterized in that, Also includes: Perform lifecycle management on historical intermediate results in the intermediate results repository; The lifecycle management includes: based on one or more factors such as access frequency, reuse count, result timeliness, and generation cost, performing result updates, result elimination, result compression, retention of hot results, or archiving of cold data.
9. The method for reusing intermediate results and accelerating inference chains for multi-model services according to claim 1, characterized in that, Also includes: Establish a unified intermediate result abstraction structure, which is used to share stage results among different models.
10. A system for reusing intermediate results and accelerating inference chains for multi-model services, characterized in that, The user request processing module is used to receive user requests and parse tasks. The inference chain building module is used to build an inference chain containing multiple stage nodes based on the task parsing results; The intermediate results management module is used to store and manage historical intermediate results; The semantic matching module is used to perform retrieval in the intermediate result repository based on stage feature description information; The inference chain scheduling module is used to make reuse decisions based on the search results and schedule the execution of the inference chain; The incremental reasoning module is used to perform local reasoning only on the affected nodes when the reuse decision is in incremental reasoning mode; The lifecycle management module is used to perform dynamic lifecycle management on historical intermediate results in the intermediate results repository.