A large language model inference dynamic scheduling method and system based on a heterogeneous model pool
Patent Information
- Application Number
- CN202611289874.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-25
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本发明提出了一种基于异构模型池的大语言模型推理动态调度方法及系统,解决现有大语言模型推理系统在面对提示注入、模型异常、输出不稳定和长期运行状态变化时,缺乏动态模型选择和运行时自适应降权机制的问题
通过基于置信度的动态调度机制,使大语言模型推理系统能够根据模型运行时表现自动调整模型调用概率,避免固定模型组合带来的长期暴露风险。将每轮参与推理的模型数量设为非固定数量,使系统能够根据任务风险、输入复杂度和系统资源状态动态调整冗余强度。通过一致性反馈更新置信度,使可靠模型在后续推理中被更多调用,使异常模型、低质量模型或被攻击模型逐步降权,从而提升系统鲁棒性。本发明不需要访问模型内部参数、梯度或训练数据,也不需要重新训练模型,可作为部署后的运行时调度层应用于已有大语言模型系统。
Smart Images

Figure CN122840261A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a dynamic scheduling method and system for large language model inference based on heterogeneous model pools. Background Technology
[0002] With the widespread application of large language models in scenarios such as intelligent question answering, text generation, code assistance, autonomous driving decision-making, security auditing, and intelligent agent control, the security and reliability issues of model inference systems are becoming increasingly prominent. Existing large language models are typically deployed as a single model or a fixed combination of models. When the model is affected by hint injection, backdoor attacks, data poisoning, or when hallucinations and logical deviations occur in distributed external inputs, fuzzy semantic inputs, or complex inference tasks, fixed inference processes struggle to identify and suppress abnormal model outputs in a timely manner, easily leading to erroneous, unstable, or insecure responses from the system.
[0003] Existing technologies commonly employ protection methods such as input filtering, output auditing, rule constraints, model fine-tuning, adversarial training, and human feedback enhancement. While these methods can improve model security to some extent, they typically rely on known attack patterns or predefined security rules, making them ill-suited to unknown attacks, migration attacks, or runtime state changes. Furthermore, protection schemes based on retraining or fine-tuning are costly and unsuitable for continuously evolving online inference environments after deployment.
[0004] The concept of dynamic heterogeneous redundancy emphasizes improving the reliability and security of a system under abnormal environments through heterogeneity, redundancy, and dynamism. Applying this concept to large language model inference systems requires further solutions to problems such as how different models are selected from the model pool, how to adjust the calling frequency based on historical performance, and how to mitigate the impact of some models malfunctioning or being attacked. If multiple models participate in inference in a fixed manner, or are simply randomly selected, it is impossible to dynamically adjust the model calling frequency based on the actual performance of the models during operation. When some models consistently perform poorly, have unstable output, or are controlled by attackers, these models may still continue to participate in inference, thus affecting the final output of the system. Summary of the Invention
[0005] This invention proposes a dynamic scheduling method and system for large language model inference based on heterogeneous model pools, which solves the problem that existing large language model inference systems lack dynamic model selection and runtime adaptive weight reduction mechanisms when facing prompt injection, model anomalies, unstable output, and long-term changes in operating status.
[0006] This invention proposes a dynamic scheduling method for large language model inference based on a heterogeneous model pool, comprising the following steps: selecting at least two heterogeneous large language models as working models from a preset model pool according to the sampling probability based on confidence, and performing inference on the current round of inference requests through the working models to obtain the inference result of each heterogeneous large language model; Based on the inference results, generate the final output data for the current round of inference request; Based on the difference between the final output data and the inference result, the confidence of the working model is updated, and the updated confidence is used to calculate the sampling probability of the working model in the next round of inference.
[0007] Optionally, before the step of selecting at least two heterogeneous large language models as working models from a preset model pool based on the sampling probability of confidence level, the method further includes: An initial confidence level is set for each heterogeneous large language model in the preset model pool based on its historical performance data, task adaptability data, response stability data, and / or historical security data. The initial confidence level is used to calculate the sampling probability of each heterogeneous large language model in the preset model pool during the first round of inference.
[0008] Optionally, before the step of selecting at least two heterogeneous large language models as working models from a preset model pool based on the sampling probability of confidence level, the method further includes: The sampling probability calculation formula is as follows:
[0009] Calculate the sampling probability of each heterogeneous large language model in the preset model pool, where, Representing heterogeneous large language models In the Confidence level of the wheel.
[0010] Optionally, the step of selecting at least two heterogeneous large language models as working models from a preset model pool based on the sampling probability of confidence includes: The number of preset working models is determined based on the task type, security risk level, input complexity, response latency requirements, and / or system resource status data. Based on the sampling probability based on confidence level, a preset number of heterogeneous large language models are selected as working models from a preset model pool, wherein the preset number is greater than or equal to 2.
[0011] Optionally, the step of selecting at least two heterogeneous large language models as working models from a preset model pool based on the sampling probability of confidence includes: Using any of the following sampling methods—sampling without replacement, sampling with replacement, weighted random sampling, stratified weighted sampling, or weighted sampling combined with task type constraints—at least two heterogeneous large language models are selected as working models from a pre-defined model pool based on the sampling probability of confidence level.
[0012] Optionally, the step of updating the confidence level of the working model based on the difference data between the final output data and the inference result includes: Based on the difference between the final output data and the inference result, a preset update formula is used:
[0013] The confidence level of the working model is updated, wherein, Indicates the result of reasoning. This indicates the final output data. , , This represents the upper confidence level.
[0014] Optionally, the method further includes: Get the number of times the inference results of each heterogeneous large language model in the preset model pool are inconsistent with the final output; Heterogeneous large language models whose inference results are inconsistent with the final output more than a preset threshold will be removed from the preset model pool.
[0015] Optionally, the method further includes: Update the confidence-based sampling probability based on at least one of the following: response time, resource consumption, deployment node status, or service load data of each heterogeneous large language model.
[0016] On the other hand, this application provides a dynamic scheduling system for large language model inference based on heterogeneous model pools, including: The sampling module is used to select at least two heterogeneous large language models as working models from a preset model pool according to the sampling probability based on confidence, and to perform inference on the current round of inference requests through the working models to obtain the inference result of each heterogeneous large language model. An integration module is used to generate final output data for the current round of inference requests based on the inference results; The update module is used to update the confidence of the working model based on the difference between the final output data and the inference result. The updated confidence is used to calculate the sampling probability of the working model in the next round of inference.
[0017] The beneficial effects and experimental verification data of this invention are as follows: By employing a confidence-based dynamic scheduling mechanism, the large language model inference system can automatically adjust the model invocation probability based on the model's runtime performance, avoiding the long-term exposure risks associated with fixed model combinations. The number of models participating in each round of inference is set to a variable number, allowing the system to dynamically adjust redundancy intensity based on task risk, input complexity, and system resource status. Confidence is updated through consistency feedback, ensuring that reliable models are invoked more frequently in subsequent inferences, while abnormal, low-quality, or attacked models are gradually de-weighted, thereby improving system robustness. This invention does not require access to model internal parameters, gradients, or training data, nor does it require model retraining. It can be applied as a runtime scheduling layer in existing large language model systems after deployment. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating a dynamic scheduling method for large language model inference based on a heterogeneous model pool, according to the present invention. Figure 2 This is a schematic diagram of the structure of a large language model inference dynamic scheduling system based on a heterogeneous model pool according to the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device according to the present invention; Figure 4 This is a schematic diagram of the structure of a storage medium according to the present invention. Detailed Implementation
[0019] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0020] like Figure 1 As shown, this invention provides a dynamic scheduling method for large language model inference based on heterogeneous model pools, including: S101. Select at least two heterogeneous large language models as working models from the preset model pool according to the sampling probability based on confidence, and use the working models to perform inference on the current round of inference requests to obtain the inference result of each heterogeneous large language model. S102. Generate the final output data for the current round of reasoning request based on the reasoning result; S103. Based on the difference between the final output data and the inference result, the confidence of the working model is updated, and the updated confidence is used to calculate the sampling probability of the working model in the next round of inference.
[0021] Optionally, before the step of selecting at least two heterogeneous large language models as working models from a preset model pool based on the sampling probability of confidence level, the method further includes: An initial confidence level is set for each heterogeneous large language model in the preset model pool based on its historical performance data, task adaptability data, response stability data, and / or historical security data. The initial confidence level is used to calculate the sampling probability of each heterogeneous large language model in the preset model pool during the first round of inference.
[0022] Optionally, before the step of selecting at least two heterogeneous large language models as working models from a preset model pool based on the sampling probability of confidence level, the method further includes: The sampling probability calculation formula is as follows:
[0023] Calculate the sampling probability of each heterogeneous large language model in the preset model pool, where, Representing heterogeneous large language models In the Confidence level of the wheel.
[0024] Optionally, the step of selecting at least two heterogeneous large language models as working models from a preset model pool based on the sampling probability of confidence includes: The number of preset working models is determined based on the task type, security risk level, input complexity, response latency requirements, and / or system resource status data. Based on the sampling probability based on confidence level, a preset number of heterogeneous large language models are selected as working models from a preset model pool, wherein the preset number is greater than or equal to 2.
[0025] Optionally, the step of selecting at least two heterogeneous large language models as working models from a preset model pool based on the sampling probability of confidence includes: Using any of the following sampling methods—sampling without replacement, sampling with replacement, weighted random sampling, stratified weighted sampling, or weighted sampling combined with task type constraints—at least two heterogeneous large language models are selected as working models from a pre-defined model pool based on the sampling probability of confidence level.
[0026] Optionally, the step of updating the confidence level of the working model based on the difference data between the final output data and the inference result includes: Based on the difference between the final output data and the inference result, a preset update formula is used:
[0027] The confidence level of the working model is updated, wherein, Indicates the result of reasoning. This indicates the final output data. , , This represents the upper confidence level.
[0028] Optionally, the method further includes: Get the number of times the inference results of each heterogeneous large language model in the preset model pool are inconsistent with the final output; Heterogeneous large language models whose inference results are inconsistent with the final output more than a preset threshold will be removed from the preset model pool.
[0029] Optionally, the method further includes: Update the confidence-based sampling probability based on at least one of the following: response time, resource consumption, deployment node status, or service load data of each heterogeneous large language model.
[0030] In one possible implementation, a dynamic scheduling method for large language model inference based on a heterogeneous model pool includes the following steps: Construct a heterogeneous model pool. Select... A large language model constitutes the model pool:
[0031] The models differ in at least one of the following aspects: model architecture, parameter size, training corpus, alignment, prompt word templates, deployment location, response latency, or inference capability. For each model... Set initial confidence level And set a confidence level cap. Reward coefficient Penalty coefficient and the number of models per round of sampling .
[0032] Obtain the current inference state. In the... During round-robin reasoning, the system receives the current input request. And read the current confidence vector of each model in the model pool:
[0033] Calculate the model sampling probability. Based on the current confidence level of each model, calculate the sampling probability. In the Sampling probability of the round Defined as:
[0034] in, Representation Model In the Confidence level of the wheel, This represents the total number of models in the model pool. This calculation method gives high-confidence models a higher probability of being scheduled, while allowing other models to still retain the opportunity to participate.
[0035] Perform dynamic model extraction. Based on the sampling probability. From the model pool Extraction Several models participate in the current round of inference, among which:
[0036] The The value can be preset to a fixed value or dynamically adjusted based on the current task type, security risk level, input complexity, response latency requirements, or system resource status. For example, for tasks with high security risks or complex input semantics, the system can increase the value. To improve the strength of redundant inference; for low-risk or latency-sensitive tasks, the system can reduce... This is to reduce reasoning costs.
[0037] Perform multi-model inference. Request the current input. Send to the selected Each model generates candidate outputs independently:
[0038] in, Representation Model In the Wheel to input The generated candidate outputs can be classification labels, true / false judgments, text answers, structured decision fields, action control instructions, security risk labels, or tool call results.
[0039] The system receives the final output and updates the confidence level. It obtains the final output based on the candidate outputs of the extracted model. And the candidate output of each extracted model With the final output Perform a consistency comparison. If... Then the model is considered If the current round of reasoning aligns with the system consensus, its confidence level is updated with a reward; if Then the model is considered If there is an output deviation in the current round of inference, its confidence is updated as a penalty.
[0040] In a preferred embodiment, the confidence update formula is:
[0041] in, , , This represents the upper confidence level.
[0042] Furthermore, the model The evolution of confidence after multiple rounds of reasoning can be expressed as:
[0043] in, This is an indicator function. It takes the value when the condition is true. Otherwise, the value is This formula shows that models that frequently match the system's final output will receive continuous rewards, while models that frequently deviate from the system's final output will be subject to continuous penalties.
[0044] Updated confidence level This serves as the basis for calculating the sampling probability in the next round of inference, enabling the system to form the following dynamic closed loop:
[0045] in, Indicates the first Round sampling probability vector, Indicates the first The set of models extracted from the round, and:
[0046] Furthermore, the method may also include a low-confidence model suppression mechanism. When a model is in continuous... The proportion of rounds of reasoning that do not match the final output exceeds a preset threshold. When this occurs, the model is marked as a low-confidence model. The inconsistency ratio can be expressed as:
[0047] When the following conditions are met:
[0048] Then reduce the model The subsequent sampling probability, or suspend its participation in scheduling within a preset time window.
[0049] Furthermore, the method may also include an exploration-and-hold mechanism. To avoid scheduling results being concentrated in a few models for a long period, the system can set a base sampling probability for each model. The corrected sampling probability is:
[0050] in:
[0051] This mechanism allows low-confidence models to still have a non-zero probability of being sampled, thus adapting to changes in model state, changes in task distribution, and recovery from temporary anomalies.
[0052] Furthermore, the method may also include a task-aware scheduling mechanism. For different types of inference requests, the system can adjust the number of models extracted in each round based on task risk or complexity. In one implementation, It can be defined as:
[0053] in, Indicates the risk level of the current input request. This represents the complexity score of the current input request. This represents the minimum number of extraction models. , , To adjust the parameters, This indicates rounding up. Therefore, the system can automatically increase redundancy in high-risk or high-complexity tasks and reduce inference costs in low-risk tasks.
[0054] Furthermore, the method may also include a latency-aware scheduling mechanism. The system can adjust the sampling probability by considering factors such as model response time, server load, deployment node status, and invocation cost. In one implementation, the sampling probability incorporating response latency can be expressed as:
[0055] in, Representation Model In the Response latency within a round or history window Let this be the time delay penalty function. Preferably, the time delay penalty function can be set as:
[0056] in, This represents the delay penalty coefficient.
[0057] For example, the construction contains A heterogeneous model pool for large language models:
[0058] in, The settings can be configured based on system deployment conditions and task requirements. For example, the model pool can contain large language models with different architectures, parameter sizes, training data sources, or alignment methods. During initialization, each model is given the same initial confidence level.
[0059] Simultaneously set the reward coefficient:
[0060] Set the penalty coefficient:
[0061] Set the confidence level cap:
[0062] The number of models participating in inference in each round is not fixed to a specific value, but is set as a configurable parameter. And satisfy:
[0063] When the current input request is received Next, the current confidence level of each model in the model pool is read. If the confidence levels of all models are the same in the initial stage, then each model has the same sampling probability:
[0064] The system is based on the sampling probability from Extracted from each model Each model participates in the current round of inference, and will The inputs are fed into the extracted models respectively. The set of extracted models is represented as follows:
[0065] And it satisfies:
[0066] For any extracted model Its candidate output is:
[0067] The final output is obtained by outputting the ruling. Then, the candidate outputs of each extracted model are... and Compare the results. If the output of a model matches the final output, multiply the confidence score of that model by [the factor]. and not exceeding If the output of a model is inconsistent with the final output, then the confidence score of that model is multiplied by . Models that are not sampled can maintain their original confidence level or be slightly adjusted according to the preset recovery strategy.
[0068] For example, in a certain round of inference, the set of models drawn from the model pool is as follows:
[0069] in:
[0070] The outputs of each model are:
[0071] If the final output is obtained through the output decision:
[0072] Then for any extracted model Its confidence level is updated as follows:
[0073] After multiple rounds of inference, models that frequently match the final output will accumulate higher confidence levels, thus increasing their probability of being selected in subsequent sampling. Conversely, models that frequently deviate from the final output will have their confidence levels reduced, thus decreasing their probability of being selected in subsequent sampling. Therefore, the system can achieve a dynamic scheduling effect online, prioritizing reliable models and de-weighting abnormal models, without retraining the models or modifying their internal parameters.
[0074] For example, if some models in the model pool are affected by malicious prompts and continuously output content specified by the attacker, while other models still output normal results, the attacked models are more likely to be inconsistent with the final consensus result in multiple rounds of inference. The scheduling module updates the model with penalties based on the inconsistent results, gradually reducing its confidence level. As the confidence level decreases, the probability of the attacked model being subsequently extracted decreases, thereby reducing its impact on the overall system output.
[0075] For example, multiple large language models or multimodal large models output driving decisions based on road images, traffic signals, navigation information, and vehicle status, respectively. The system dynamically extracts data based on historical consistency and current confidence. Several models participate in decision-making. For high-risk scenarios, such as intersections, pedestrian crossings, or traffic light recognition conflicts, the system can improve... To enhance redundancy detection; for low-risk scenarios, the system can reduce... This reduces inference costs. Through task-aware scheduling, the system can achieve a balance between security and computational efficiency.
[0076] For example, if a model has a high confidence level but its response time consistently exceeds a preset threshold, the system can incorporate a latency penalty term into the sampling probability calculation, allowing the scheduling process to consider both reliability and response efficiency. The corrected sampling probability can be calculated based on confidence level, average response latency, and current load, making it suitable for distributed deployment environments.
[0077] On the other hand, such as Figure 2As shown, this application provides a dynamic scheduling system for large language model inference based on heterogeneous model pools, including: The sampling module 201 is used to select at least two heterogeneous large language models as working models from a preset model pool according to the sampling probability based on confidence, and to perform inference on the current round of inference request through the working models to obtain the inference result of each heterogeneous large language model. Integration module 202 is used to generate final output data for the current round of inference request based on the inference results; The update module 203 is used to update the confidence of the working model based on the difference data between the final output data and the inference result. The updated confidence is used to calculate the sampling probability of the working model in the next round of inference.
[0078] In one possible implementation, such as Figure 3 As shown, this application embodiment provides an electronic device 300, including: a memory 310, a processor 320, and a first computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the first computer program 311, it performs the following: selects at least two heterogeneous large language models as working models from a preset model pool according to the sampling probability based on confidence, and performs inference on the current round of inference request through the working models to obtain the inference result of each heterogeneous large language model; Based on the inference results, generate the final output data for the current round of inference request; Based on the difference between the final output data and the inference result, the confidence of the working model is updated, and the updated confidence is used to calculate the sampling probability of the working model in the next round of inference.
[0079] In one possible implementation, such as Figure 4 As shown, this application embodiment provides a computer-readable storage medium 400, on which a second computer program 411 is stored. When the second computer program 411 is executed by a processor, it performs the following: selects at least two heterogeneous large language models as working models from a preset model pool according to the sampling probability based on confidence, and performs inference on the current round of inference request through the working models to obtain the inference result of each heterogeneous large language model. Based on the inference results, generate the final output data for the current round of inference request; Based on the difference between the final output data and the inference result, the confidence of the working model is updated, and the updated confidence is used to calculate the sampling probability of the working model in the next round of inference.
[0080] The present invention has been described above through a preferred embodiment. This solution can also be implemented by means of an apparatus or device to perform the methods described in the above embodiments.
[0081] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of this solution includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which the embodiments of this solution pertain. The processor performs the various methods and processes described above. For example, the method embodiments of this solution can be implemented as software programs tangibly contained in a machine-readable medium, such as memory. In some embodiments, part or all of the software program can be loaded and / or installed via memory and / or a communication interface. When the software program is loaded into memory and executed by the processor, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the processor can be configured to perform one of the methods described above by any other suitable means (e.g., by means of firmware).
[0082] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A dynamic scheduling method for large language model inference based on heterogeneous model pools, characterized in that, Includes the following steps: At least two heterogeneous large language models are selected as working models from a preset model pool based on the sampling probability of confidence level, and the current round of inference requests are inferred through the working models to obtain the inference result of each heterogeneous large language model. Based on the inference results, generate the final output data for the current round of inference request; Based on the difference between the final output data and the inference result, the confidence of the working model is updated, and the updated confidence is used to calculate the sampling probability of the working model in the next round of inference.
2. The method for dynamic scheduling of large language model inference based on heterogeneous model pools according to claim 1, characterized in that, Before the step of selecting at least two heterogeneous large language models as working models from a preset model pool based on the sampling probability of confidence level, the method further includes: An initial confidence level is set for each heterogeneous large language model in the preset model pool based on its historical performance data, task adaptability data, response stability data, and / or historical security data. The initial confidence level is used to calculate the sampling probability of each heterogeneous large language model in the preset model pool during the first round of inference.
3. The method for dynamic scheduling of large language model inference based on heterogeneous model pools according to claim 1, characterized in that, Before the step of selecting at least two heterogeneous large language models as working models from a preset model pool based on the sampling probability of confidence level, the method further includes: The sampling probability calculation formula is as follows: Calculate the sampling probability of each heterogeneous large language model in the preset model pool, where, Representing heterogeneous large language models In the Confidence level of the wheel.
4. The method for dynamic scheduling of large language model inference based on heterogeneous model pools according to claim 1, characterized in that, The step of selecting at least two heterogeneous large language models as working models from a preset model pool based on the sampling probability of confidence level includes: The number of preset working models is determined based on the task type, security risk level, input complexity, response latency requirements, and / or system resource status data. Based on the sampling probability based on confidence level, a preset number of heterogeneous large language models are selected as working models from a preset model pool, wherein the preset number is greater than or equal to 2.
5. The method for dynamic scheduling of large language model inference based on heterogeneous model pools according to claim 1, characterized in that, The step of selecting at least two heterogeneous large language models as working models from a preset model pool based on the sampling probability of confidence level includes: Using any of the following sampling methods—sampling without replacement, sampling with replacement, weighted random sampling, stratified weighted sampling, or weighted sampling combined with task type constraints—at least two heterogeneous large language models are selected as working models from a pre-defined model pool based on the sampling probability of confidence level.
6. The method for dynamic scheduling of large language model inference based on heterogeneous model pools according to claim 1, characterized in that, The step of updating the confidence level of the working model based on the difference between the final output data and the inference result includes: Based on the difference between the final output data and the inference result, a preset update formula is used: The confidence level of the working model is updated, wherein, Indicates the result of reasoning. This indicates the final output data. , , This represents the upper confidence level.
7. The method for dynamic scheduling of large language model inference based on heterogeneous model pools according to claim 1, characterized in that, The method further includes: Get the number of times the inference results of each heterogeneous large language model in the preset model pool are inconsistent with the final output; Heterogeneous large language models whose inference results are inconsistent with the final output more than a preset threshold will be removed from the preset model pool.
8. The method for dynamic scheduling of large language model inference based on heterogeneous model pools according to claim 1, characterized in that, The method further includes: Update the confidence-based sampling probability based on at least one of the following: response time, resource consumption, deployment node status, or service load data of each heterogeneous large language model.
9. A dynamic scheduling system for large language model inference based on heterogeneous model pools, characterized in that, include: The sampling module is used to select at least two heterogeneous large language models as working models from a preset model pool according to the sampling probability based on confidence, and to perform inference on the current round of inference requests through the working models to obtain the inference result of each heterogeneous large language model. An integration module is used to generate final output data for the current round of inference requests based on the inference results; The update module is used to update the confidence of the working model based on the difference between the final output data and the inference result. The updated confidence is used to calculate the sampling probability of the working model in the next round of inference.