Self-adaptive deep integration model deployment method based on dynamic load prediction

Through dynamic load prediction and greedy heuristic algorithms to adjust the model deployment solution, the problem of difficult performance requirements under resource constraints in multi-model integrated inference scenarios is solved, and efficient inference performance and resource utilization are achieved.

CN120104146APending Publication Date: 2025-06-06UNIV OF SCI & TECH OF CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510087039.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively respond to high-performance requirements in resource-constrained environments in multi-model integrated inference scenarios, resulting in low resource utilization or excessive task processing delay.

Method used

Adaptive deep integration model deployment method based on dynamic load prediction is adopted, and the model deployment scheme is dynamically adjusted to adapt to load fluctuations and optimize inference efficiency and model accuracy through mathematical modeling and greedy heuristic algorithms.

Benefits of technology

It significantly improves the accuracy and response speed of multi-model integrated inference, reduces the cut-off time miss rate by 20%, and is suitable for a variety of deep ensemble learning inference scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104146A_ABST
    Figure CN120104146A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic load prediction-based adaptive deep integration model deployment method, which comprises the following steps of: performing mathematical modeling on a problem according to computing resource limitation of a scene and an integration model, and determining an objective function for optimizing the problem; the future load is predicted, and a load prediction result is obtained; and based on the objective function and according to the load prediction result, generating a deployment scheme of a deep integration model by adopting a greedy heuristic algorithm, and adjusting the generated deployment scheme of the deep integration model according to the currently deployed maximum load throughput. According to the method, the condition of load fluctuation can be well adapted by dynamically adjusting a model deployment scheme, and the accuracy and response speed of multi-model integrated reasoning are remarkably improved; in addition, the method has the advantage of being high in universality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to an adaptive deep integration model deployment method based on dynamic load prediction. Background Art

[0002] With the widespread application of deep learning models in various scenarios, reasoning systems are facing increasing performance requirements. Currently, in multi-model integrated reasoning scenarios, a static deployment strategy is usually adopted, that is, a fixed number of model instances are loaded for each model. However, since the query load of each model in real scenarios varies and fluctuates frequently, the corresponding static deployment scheme is difficult to cope with the high performance requirements under resource-constrained environments, which leads to low resource utilization or excessive task processing delays.

[0003] In view of this, the present invention is proposed. Summary of the invention

[0004] The purpose of the present invention is to provide an adaptive deep integration model deployment method based on dynamic load prediction, so as to effectively cope with high performance requirements under resource-constrained environments while taking into account reasoning performance, and solve the technical problems existing in the prior art.

[0005] The objective of the present invention is achieved through the following technical solutions:

[0006] An adaptive deep integration model deployment method based on dynamic load prediction, comprising:

[0007] The problem is mathematically modeled according to the computing resource constraints and integrated model of the scenario, and the objective function of the optimization problem is determined; the future load is also predicted to obtain the load prediction result;

[0008] Based on the objective function and according to the load prediction result, a deployment plan of the deep integration model is generated by using a scheduling optimization method based on a greedy heuristic algorithm;

[0009] According to a predetermined period and / or time point, the maximum load throughput of the current deployment is calculated, and the generated deployment plan of the deep integration model is adjusted according to the maximum load throughput of the current deployment.

[0010] Preferably, the objective function of determining the optimization problem includes:

[0011] The objective function of the optimization problem is: max S f(D,S,λ);

[0012] And need to meet:

[0013] in, For the current deployment plan; is the deployment adjustment plan; λ is the load prediction result; c i is the memory usage of the model; M is the memory limit of the scene computing resources; i is the model number; m is the total number of models.

[0014] Preferably, the process of predicting future load includes:

[0015] A time series model based on the long short-term memory (LSTM) algorithm is used to predict future loads based on past loads to obtain load forecast results.

[0016] Preferably, the process of generating the deployment side of the deep integration model by scheduling optimization based on the greedy heuristic algorithm includes:

[0017] The optimized objective function is determined as:

[0018] g(D,S,λ)=g 1 +σg 2 ;

[0019] in:

[0020] In the formula, is the load throughput of the ith model; a is a preset threshold used to control the degree to which the load throughput of the model exceeds the predicted load;

[0021] Where n is the number of models in the ensemble model, i and j iterate from 1 to n to calculate the sum of the differences in load throughput between different models g 2 , σ is a weight parameter, which takes a negative value and is used to adjust the balance between load throughput and accuracy;

[0022] A deployment plan for the deep integration model is generated based on the optimized objective function g.

[0023] Preferably, the value of the weight parameter σ is adjusted and set through experiments.

[0024] Preferably, the weight parameter σ is -1 to minimize g 2 , and maximize g 1 .

[0025] Preferably, the process of adjusting the generated deployment scheme of the deep integration model includes:

[0026] Calculate the maximum load throughput of the current deployment according to the scheduled period or scheduled time point

[0027] When the maximum load throughput q meets the predetermined deployment scheme adjustment condition, the deployment adjustment operation is triggered to adjust the deployment scheme of the deep integration model.

[0028] Preferably, the predetermined deployment scheme adjustment condition includes:

[0029] q<λ, indicating that the current load throughput is insufficient;

[0030] and / or,

[0031] q>λ+2a: indicates that the current load throughput is excessive.

[0032] Preferably, in the method, the process of adjusting the deployment scheme of the deep integration model includes:

[0033] Calling a linear programming solver to further optimize the previously obtained optimized objective function g under the condition of satisfying resource constraints, and generating an adjustment plan S;

[0034] Update the deployment plan according to the adjustment plan S Complete the adjustment of the deployment plan of the deep integration model.

[0035] Compared with the prior art, the adaptive deep integration model deployment solution provided by the present invention can effectively improve system performance, that is, it can adapt well to load fluctuations by dynamically adjusting the model deployment solution, significantly improve the accuracy and response speed of multi-model integrated reasoning, and experiments have shown that it can reduce the deadline miss rate by 20%. Moreover, the technical solution provided by the present invention also has the advantage of strong versatility. It can be applied to a variety of deep integration learning and reasoning scenarios, and can be extended to complex systems such as cloud-edge collaboration. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0037] Figure 1 A schematic diagram of the implementation principle of the deep integration model deployment solution provided in an embodiment of the present invention;

[0038] Figure 2 A schematic diagram of the processing flow of the deep integration model deployment solution provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0039] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in combination with the specific content of the present invention; it is obvious that the described embodiments are only part of the embodiments of the present invention, not all of the embodiments, which does not constitute a limitation of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of the present invention.

[0040] First, the terms that may be used in this article are explained as follows:

[0041] The term “and / or” means that either or both of them can be realized at the same time. For example, X and / or Y means both “X” or “Y” and “X and Y”.

[0042] The terms "include", "comprises", "contains", "has" or other descriptions with similar semantics should be interpreted as non-exclusive inclusion. For example, including certain technical feature elements (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or products, etc.) should be interpreted as including not only certain technical feature elements explicitly listed, but also other technical feature elements known in the art that are not explicitly listed.

[0043] The term "consisting of..." means excluding any technical feature elements not explicitly listed. If this term is used in a claim, it will make the claim closed, so that it does not contain technical feature elements other than the technical feature elements explicitly listed, except for the conventional impurities related to them. If this term only appears in a clause of a claim, it only limits the elements explicitly listed in the clause, and the elements recorded in other clauses are not excluded from the overall claim.

[0044] Unless otherwise specified or limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example: it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in this article can be understood according to specific circumstances.

[0045] When concentration, temperature, pressure, size or other parameters are expressed in the form of a numerical range, the numerical range should be understood to specifically disclose all ranges formed by the pairing of any upper limit, lower limit, and preferred value in the numerical range, regardless of whether the range is explicitly stated; for example, if a numerical range of "2 to 8" is stated, the numerical range should be interpreted as including ranges such as "2 to 7", "2 to 6", "5 to 7", "3 to 4 and 6 to 7", "3 to 5 and 7", "2 and 5 to 7", etc. Unless otherwise specified, the numerical ranges stated herein include both their end values ​​and all integers and fractions within the numerical range.

[0046] The orientation or position relationship indicated by terms such as "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", etc. are based on the orientation or position relationship shown in the drawings and are only for the convenience and simplification of description, and do not explicitly or implicitly indicate that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation of this document.

[0047] During the specific implementation process of the present invention, it was found through research and analysis that in the field of cloud computing, dynamic load adjustment technology (such as automatic expansion and contraction) has been widely studied and applied. However, this technology cannot be applied to deep model integrated reasoning systems with multiple different models. This is because: unlike the "infinite" resource assumption in the cloud environment, the deep integrated model deployment process is limited by device memory and computing power; and it is also necessary to achieve a dynamic balance of resource allocation among multiple models, rather than adjustments to a single model. Based on this, the present invention specifically provides an adaptive method that can dynamically adjust the model deployment plan to take into account both reasoning performance and resource efficiency.

[0048] The technical solution provided by the present invention relates to the fields of deep learning reasoning, dynamic resource allocation and system optimization. It specifically provides a load adaptive deployment method based on a deep model integrated reasoning system (i.e., an adaptive deep integrated model deployment method based on dynamic load prediction), which aims to solve the performance bottleneck of the static deployment strategy in the multi-model integrated reasoning system. It specifically introduces a query load prediction mechanism and an optimization algorithm to dynamically adjust the deployment quantity of different models and optimize the reasoning efficiency and model accuracy.

[0049] Furthermore, the objectives to be achieved during the implementation of the embodiments of the present invention include at least:

[0050] Under resource-constrained conditions, the overall accuracy of the reasoning system can be improved by dynamically optimizing the model deployment plan. In the process of dynamically optimizing the model deployment plan, the deployment plan can be adjusted in advance by predicting load fluctuations to avoid system performance degradation under high or low load conditions. Furthermore, efficient heuristic algorithms can be used to ensure that the response time of dynamic deployment adjustments is short enough to meet real-time reasoning requirements.

[0051] Once the above objectives are achieved, the implementation scheme provided by the present invention can effectively optimize the performance of the multi-model integrated reasoning system.

[0052] To facilitate the understanding of the present invention, a specific implementation scheme of an adaptive deep integration model deployment method based on dynamic load prediction proposed by the present invention will be described in detail below with reference to the accompanying drawings.

[0053] Specifically, refer to Figure 1 and Figure 2 As shown, the present invention provides an adaptive deep integration model deployment method based on dynamic load prediction, which includes at least the following multiple processing steps:

[0054] Step 21, deployment problem modeling, that is, mathematically modeling the problem according to the computing resource constraints of the scenario and the integrated model, taking the resource constraints as the conditions of the optimization problem, and integrating the integrated model information into the mathematical model to obtain an ILP (integer linear programming problem);

[0055] In the deployment problem modeling process, we first need to formally analyze the problem's goals and constraints. The goal is to dynamically adjust the model deployment plan to adapt to real-time load changes and maximize the inference accuracy. The constraint is to ensure that resource consumption is within a certain range.

[0056] Specifically, the deployment problem can be modeled as an ILP problem, namely:

[0057] For each model i in the ensemble model, adjust the number of its instances s i (, where a positive value indicates loading, and a negative value indicates unloading;

[0058] The objective function of the optimization problem is: max S f(D,S,λ) represents the expected maximization of the performance that the system can achieve under the deployment scheme and resource constraints, and the constraint is that for any model i, the following formula must be satisfied:

[0059]

[0060] in, For the current deployment plan; is the adjustment plan, i.e., the deployment adjustment plan; λ is the predicted query load, i.e., the load prediction result; c i is the memory usage of the model; M is the memory limit of the scene computing resources; i is the model number; m is the total number of models.

[0061] Step 22, predicting the future load and obtaining a load prediction result;

[0062] Specifically, a time series model based on the LSTM (Long Short-Term Memory) algorithm can be used to predict future query loads in order to provide input basis for dynamic adjustments;

[0063] Furthermore, for the unknown load prediction results λ in the optimization problem, that is, the load of the model at the future moment, the load of each model in the future can be predicted before the optimization problem is solved; for example, the time series model based on the LSTM architecture can be used to predict the future load prediction results λ according to the existing work in the related field. The model can accurately predict the future system load based on the past load value and obtain the corresponding load prediction results λ, thereby guiding the model deployment and avoiding excessive or insufficient adjustments;

[0064] Step 23, generating a deployment plan for the deep integration model by using a heuristic scheduling optimization method;

[0065] Specifically, in the heuristic scheduling optimization process, in order to overcome the computational complexity of the ILP problem, a greedy heuristic algorithm can be used to generate a deployment adjustment plan by weighing the throughput (i.e., load throughput) and the inference accuracy. After that, the performance of the current deployment plan needs to be periodically evaluated. When it is determined according to the evaluation results that an adjustment condition needs to be triggered, that is, when the current deployment plan needs to be adjusted, a new deployment plan is generated through the optimization algorithm.

[0066] Furthermore, in the process of solving the optimization using the heuristic scheduling algorithm, since the above optimization target f (that is, the objective function of the optimization problem is: max S f(D, S, λ)) is difficult to directly obtain an accurate value, so solving the ILP problem will be too complicated. For this reason, a greedy heuristic algorithm is proposed in the implementation process of the present invention to optimize the following objective function, and the optimized objective function is:

[0067] g(D,S,λ)=g 1 +σg 2 ;

[0068] in:

[0069] Through this formula, Select a smaller value between and a; where, is the load throughput of the i-th model; a is a preset threshold value, which is used to control the degree to which the load throughput of the model exceeds the predicted load, that is, when the load throughput of the model exceeds the predicted load by an increment exceeding a, the optimized objective function will not continue to increase; in other words, the present invention expects the sum of the load throughputs of all models to be close to the future load measurement results, so as to ensure that the load demand is met;

[0070] Where n is the number of models in the ensemble model, i and j are iterated from 1 to n to calculate the sum of the throughput differences between different models g 2 If we keep g 2 The smaller the g is, the more different models the user can use. Otherwise, the user can only use the model with the largest load throughput. 2 It can promote the diversity of model deployment. Since the accuracy of the integrated model comes from the diversity of the model, it can also effectively improve the inference accuracy.

[0071] The σ in the formula is a weight parameter used to adjust the balance between load throughput and accuracy. The weight parameter can be adjusted and set through experiments. If you want to minimize g 2 , and maximize g1, then the weight parameter needs to be a negative value. For example, the weight parameter can be selected as -1 by default;

[0072] Based on the above optimized objective function, the corresponding current deployment plan can be generated.

[0073] Next, the deployment scheme of the current model generated above needs to be adjusted accordingly to adapt to the application scenario of dynamic load. The following algorithm can be used to implement the deployment adjustment process, namely:

[0074] (1) Periodically calculate the maximum load throughput of the current deployment Of course, the corresponding maximum load throughput may also be calculated at a predetermined time point, that is, this step may be started and executed at a predetermined period and / or time point to adjust the subsequent deployment plan;

[0075] (2) When the corresponding maximum load throughput q satisfies the following conditions according to the above calculation results, the corresponding deployment adjustment operation is triggered:

[0076] Condition 1: q<λ: indicates that the current load throughput is insufficient; and / or,

[0077] Condition 2: q>λ+2a: indicates that the current load throughput is excessive;

[0078] (3) Calling the linear programming solver to optimize the above optimized objective function g under the condition of satisfying resource constraints and generate an adjustment plan S;

[0079] (4) Update the deployment plan according to the adjustment plan S Deployment adjustment can be performed based on the updated deployment plan, completing the deployment plan update and adjustment process.

[0080] Through the above process, periodic adjustments to the deployment plan can be achieved to effectively adapt to possible load fluctuations.

[0081] Furthermore, the adaptive deep integration model deployment method provided by the present invention can effectively improve system performance, including the ability to dynamically adjust the model deployment scheme to adapt to load fluctuations, and can also significantly improve the accuracy and response speed of multi-model integrated reasoning. Experiments show that the deadline miss rate can be reduced by 20% during the implementation of the present invention. Moreover, the embodiment of the present invention also has the advantage of strong versatility, that is, it can be applied to a variety of deep integrated learning and reasoning scenarios, and can be extended to complex systems such as cloud-edge collaboration.

[0082] The present invention can be applied to various suitable scenarios during the specific implementation process. For example, it can be deployed in a bank's intelligent customer service system, but is not limited to it, in which thousands of bank-related information inquiries are usually generated per second. The customer service system deploys multiple deep learning models for matching reasoning, and the query load fluctuates greatly over time. After using the adaptive deployment scheme provided by the embodiment of the present invention, the customer service system can dynamically adjust the number of model instances according to the predicted traffic fluctuations. During peak hours, faster models can be adaptively deployed, while maintaining a certain recommendation accuracy, the response speed can be greatly improved, allowing users to receive feedback in a timely manner, and the user deadline miss rate is reduced by 20% compared to traditional static deployment schemes. In periods with fewer inquiries, more models with high accuracy can be deployed to further improve the accuracy.

[0083] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with the technical field within the technical scope disclosed in the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims. The information disclosed in the background technology section of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or in any form that the information constitutes prior art known to those skilled in the art.

Claims

1. An adaptive deep integration model deployment method based on dynamic load prediction, characterized in that: include: Mathematically model the problem based on the scenario's computing resource constraints and integrated model, and determine the objective function of the optimization problem; And also predict the future load to obtain the load prediction result; Based on the objective function and according to the load prediction result, a deployment plan of the deep integration model is generated by using a scheduling optimization method based on a greedy heuristic algorithm; According to a predetermined period and / or time point, the maximum load throughput of the current deployment is calculated, and the generated deployment plan of the deep integration model is adjusted according to the maximum load throughput of the current deployment.

2. The method according to claim 1, characterized in that The objective function of determining the optimization problem includes: The objective function of the optimization problem is: max S f(D,S,λ); And need to meet: in, For the current deployment plan; is the deployment adjustment plan; λ is the load prediction result; c i is the memory usage of the model; M is the memory limit of the scene computing resources; i is the model number; m is the total number of models.

3. The method according to claim 1, characterized in that The process of predicting future load includes: A time series model based on the long short-term memory (LSTM) algorithm is used to predict future loads based on past loads to obtain load forecast results.

4. The method according to claim 1, 2 or 3, characterized in that: The process of generating a deployment side of a deep integration model by scheduling optimization based on a greedy heuristic algorithm includes: The optimized objective function is determined as: g(D,S,λ)=g1+σg2; in: In the formula, is the load throughput of the ith model; a is a preset threshold used to control the degree to which the load throughput of the model exceeds the predicted load; Where n is the number of models in the ensemble model, i and j iterate from 1 to n to calculate the sum of the differences in load throughput between different models g2, σ is a weight parameter, which is a negative value and is used to adjust the balance between load throughput and accuracy; A deployment plan for the deep integration model is generated based on the optimized objective function g.

5. The method according to claim 4, characterized in that The value of the weight parameter σ is adjusted and set through experiments.

6. The method according to claim 5, characterized in that The weight parameter σ takes a value of -1 to minimize g2 and maximize g1.

7. The method according to claim 4, characterized in that The process of adjusting the generated deployment scheme of the deep integration model includes: Calculate the maximum load throughput of the current deployment according to the scheduled period or scheduled time point When the maximum load throughput q meets the predetermined deployment plan adjustment condition, the deployment adjustment operation is triggered to adjust the deployment plan of the deep integration model.

8. The method according to claim 7, characterized in that The predetermined deployment scheme adjustment conditions include: q<λ, indicating that the current load throughput is insufficient; and / or, q>λ+2a: indicates that the current load throughput is excessive.

9. The method according to claim 7, characterized in that: The process of adjusting the deployment scheme of the deep integration model includes: Calling a linear programming solver to further optimize the optimized objective function g obtained previously under the condition of satisfying resource constraints, and generating an adjustment plan S; Update the deployment plan according to the adjustment plan S Complete the adjustment of the deployment plan of the deep integration model.