Resource-efficient and service-quality-aware adaptive scheduling method for inference service systems

By introducing the deep learning model automatic selection module and the collaborative management module based on deep reinforcement learning in the deep learning inference service system, the optimal deep learning model is automatically selected and resource configuration is dynamically adjusted, which solves the problems of low GPU resource utilization and difficult to guarantee service quality in the existing technology, and achieves efficient resource utilization and excellent service quality.

CN115129477BActive Publication Date: 2025-05-23SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210918942.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-05-02
Filing Date
2022-08-01
Publication Date
2025-05-23
Estimated Expiration
2042-08-01

AI Technical Summary

Technical Problem

The GPU resource utilization rate of existing deep learning inference service systems is low, mainly because it is difficult for users to accurately select the optimal deep learning model, and due to the dynamic changes in workloads, it is difficult to effectively adjust the resource allocation strategy, resulting in increased inference delay and difficult to guarantee service quality.

Method used

An adaptive scheduling method for inference service system with efficient resource and service quality perception is designed. The deep learning model is used to automatically select the optimal deep learning model and dynamically adjust the GPU resource allocation and batch size to improve the utilization rate of GPU resource and ensure service quality.

Benefits of technology

On the premise of ensuring user service quality, the GPU resource utilization rate of the inference service system has been significantly improved, and the ease of use and resource utilization efficiency of the system has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115129477B_ABST
    Figure CN115129477B_ABST
Patent Text Reader

Abstract

The present invention discloses an adaptive scheduling method for an inference service system with high resource efficiency and service quality awareness, comprising: a deep learning model automatic selection module uses a collaborative filtering method to predict the inference performance of an inference workload running on different deep learning models; the deep learning model automatic selection module uses a greedy algorithm to select an optimal deep learning model that meets the user's service quality requirements, and deploys the optimal deep learning model to a container to serve the inference workload in the inference service system; a collaborative management module uses a deep reinforcement learning method to collaboratively adjust GPU resource allocation and batch size settings according to dynamic changes in the inference workload. The present invention can automatically select a deep learning model according to user needs, and can collaboratively adjust GPU resource allocation and batch size settings according to dynamic changes in the inference workload.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of deep learning and distributed computing, and in particular to an adaptive scheduling method for an inference service system with high resource efficiency and service quality awareness. Background Art

[0002] As the reasoning accuracy of deep learning models continues to improve, more and more applications rely on deep learning reasoning as decision support, such as face recognition, autonomous driving, and intelligent voice assistants. However, as the reasoning accuracy of deep learning models increases, the reasoning model becomes more complex, which greatly increases the latency of using only the CPU to process deep learning reasoning requests. In a production environment, deep learning reasoning services are usually interactive and have quality of service requirements, such as requiring 99% of reasoning requests to be completed within 100ms. Therefore, current deep learning reasoning service systems increasingly use GPUs to accelerate the processing of reasoning requests to reduce reasoning latency and thereby ensure user service quality.

[0003] However, the GPU resource utilization rate of the current deep learning inference service system is very low, mainly due to the following two reasons. First, the existing deep learning inference service system usually requires users to independently select deep learning models and resource allocation strategies according to service quality requirements (such as inference latency, inference accuracy, etc.). However, since deep learning models are jointly defined by model architecture, programming framework, optimization compiler, etc., there are a large number of deep learning models, and it is difficult for users to accurately select the optimal model manually. At the same time, due to the dynamic changes of workloads in production environments, users usually adopt conservative overprovisioning strategies to allocate resources to ensure service quality, which greatly reduces GPU resource utilization. Secondly, modern GPUs have a large amount of memory and computing resources, and require a high parallel mechanism to achieve peak throughput. However, the existing deep learning inference service system usually adopts a smaller batch size to ensure user service quality. The smaller batch size limits the parallelism of program running, so GPU resources cannot be fully utilized.

[0004] The existing technology usually allows users to independently select deep learning models and bind the models to the hardware platform, and improves the resource utilization of the reasoning service system through dynamic batch size adjustment and automatic virtual machine or container instance expansion. However, on the one hand, it is difficult for users to accurately select the optimal model manually, and the manual method reduces the usability of the system; on the other hand, the adjustment method of virtual machines or instances is coarse-grained, and due to the limitation of batch size, GPU resources cannot be fully utilized. For this reason, the existing technology cannot enable the reasoning service system to effectively utilize GPU resources while ensuring the quality of user services. Summary of the invention

[0005] In view of the above problems, the present invention designs a resource-efficient and service quality-aware reasoning service system adaptive scheduling method, which is applied to a container-based reasoning service system and can automatically select a deep learning model according to the user's service quality requirements; according to the change of the reasoning workload, GPU resources can be dynamically allocated and the batch size can be adjusted for the operation of the deep learning model, thereby improving the GPU resource utilization of the reasoning service system while ensuring the user service quality. The adaptive scheduling method includes:

[0006] Build a deep learning model automatic selection module that can analyze the inference workload in the inference service system, and build a deep reinforcement learning-based collaborative management module that can interact with the inference service system in real time;

[0007] The deep learning model automatic selection module predicts the inference performance of the inference workload running on different deep learning models using a collaborative filtering method, wherein the inference performance includes inference latency and inference accuracy;

[0008] The deep learning model automatic selection module selects the optimal deep learning model that meets the user's service quality requirements using a greedy algorithm, and deploys the optimal deep learning model into a container to serve the inference workload in the inference service system;

[0009] The collaborative management module uses deep reinforcement learning methods to collaboratively adjust GPU resource allocation and batch size settings according to the dynamic changes of inference workloads based on the optimal deep learning model deployed in the container.

[0010] Furthermore, the deep learning model automatic selection module for analyzing the reasoning workload in the reasoning service system includes: collecting deep learning models in the production environment to construct a model knowledge base V = {v 1 , v 2 , ..., v j , ..., v m}, where v jrepresents the j-th deep learning model.

[0011] Furthermore, the method of using the collaborative filtering method to predict the inference performance of the inference workload running on different deep learning models includes:

[0012] ① Select the inference workload U of n users from the historical data of the inference service system = {u 1 ,u 2 , ..., u i , ..., u m}, where u i represents the inference workload of the i-th user;

[0013] Each inference workload u i Perform analysis on all the deep learning models in the model knowledge base V to construct a utility matrix P n×m , where the utility matrix P n×m Each element p in ij represents the inference workload u i Model v in model knowledge base V j For the inference workload u newly submitted by the user n+1 In the model knowledge base V, any two deep learning models are selected for analysis to obtain u n+1 The inference performance on any two of the deep learning models is calculated and the utility matrix is ​​inserted into the new utility matrix P (n+1)×m , the above specific expressions are as follows:

[0014]

[0015] ②Use collaborative filtering method based on online matrix decomposition to predict the inference workload u n+1 The reasoning performance on the deep learning models other than the two deep learning models selected in ① in the model knowledge base V is as follows:

[0016] Online matrix factorization maps the inference workload and deep learning model into a low-dimensional joint latent factor space; then the inner product of the joint latent factors in the joint latent factor space is used to obtain the new utility matrix P (n+1)×m Each element p in ij The value of , and then predict the inference workload u n+1 The reasoning performance on the deep learning models other than the two deep learning models selected in ① in the model knowledge base V is specifically:

[0017] The potential inference workload factor mapped by the inference workload and the potential deep learning model factor mapped by the deep learning model are expressed as and Among them, d is much smaller than n+1 and m;

[0018] By derivation and Fit the new utility matrix P (n+1)×m ,Right now This can then accurately predict the inference workload u n+1 The reasoning performance on the deep learning models other than the two deep learning models selected in ① in the model knowledge base V; wherein, the derivation and The process is as follows:

[0019] Based on the observed inference workload u i In the deep learning model v j Reasoning performance on p ij , construct the loss function as follows,

[0020]

[0021] in, and is the regularization coefficient, and is u i and v j The eigenvector of ,||·|| 2 is the Euclidean norm, g ij express g(·) is a logical function that maps values ​​to [0, 1],

[0022] Then use the stochastic gradient descent method to minimize the loss function Derivation and The details are as follows:

[0023]

[0024] where g′ ij express g′(·) is the derivative of g(·);

[0025] Finally, by calculating and The inner product can predict the inference workload u n+1 The inference performance on other deep learning models in the model knowledge base V is expressed as follows:

[0026]

[0027] ③ Based on the above method ②, predict the inference workload u respectively n+1 The inference delay on the other deep learning models in the model knowledge base V except the two deep learning models selected in ① and inference accuracy And get the inference workload u n+1 The set of inference latencies on all deep learning models in the model repository V and the inference accuracy set

[0028] ④ Based on the inference delay set obtained in ③ above and the inference accuracy set Calculate the inference workload u n+1 The inference latency on each of the deep learning models and the inference accuracy Reasoning about latency requirements in user quality of service and reasoning accuracy requirements difference and Get the set S L ={L (n+1)j} and S A ={A (n+1)j Delete S after} L and S A The elements less than 0 in the set S′ L ={L (n+1)j′} and S′ A ={A (n+1)j″}.

[0029] Furthermore, the method of selecting the optimal deep learning model that meets the user's service quality requirements by using a greedy algorithm includes:

[0030] (i) For the set S′ A ={A (n+1)j″} are arranged in ascending order, and the set S′ L ={L (n+1)j′} to sort the elements in ascending order;

[0031] (ii) From S′ A ={A (n+1)j″} Take out the deep learning model v′ corresponding to the first element 1 , then to the set S′ L ={L (n+1)j′} to find out whether the deep learning model v′ exists 1 ;

[0032] If it exists, then the deep learning model v′1 is the optimal model;

[0033] If it cannot exist, then from S′ A ={A (n+1)j″} to extract the second element v′ 2 , repeat to set S′ L ={L (n+1)j′} until it finds a pair that exists in S′ A ={A (n+1)j″} and S′ L ={L (n+1)j′}, wherein the first deep learning model is the optimal model v that satisfies the user service quality. optimal .

[0034] Furthermore, the construction of a collaborative management module based on deep reinforcement learning that can interact with the reasoning service system in real time includes:

[0035] Assume that there are k containers in the inference service system that are providing inference services for k user inference workloads. The collaborative management module receives the feedback status of the inference service system. According to the feedback status, the collaborative management module maximizes the GPU resource utilization of the k containers and meets the inference latency requirements of the k user inference workloads. Specifically:

[0036]

[0037] in, represents the GPU resources actually used by the i-th inference workload, represents the GPU resources allocated to the container hosting the i-th inference workload, Indicates the GPU resource allocation for the i-th inference workload and batch size BZ i The inference latency under, D represents the total amount of GPU resources in the system, represents the inference latency requirement of the i-th inference workload.

[0038] Furthermore, the collaborative management module uses a deep reinforcement learning method to collaboratively adjust GPU resource allocation and batch size settings based on the optimal deep learning model deployed in the container according to the dynamic changes of the inference workload, including:

[0039] 1) Build an agent b for each container that carries the inference workload i , agent b i Responsible for GPU resource allocation and corresponding batch size setting for containers carrying inference workload i, for each agent b iAccording to the Markov decision process model, it is as follows:

[0040] State space; S represents the state space, let s = {s 1 ,s 2 ,...,s k} represents agent b 1 , b 2 , …, b k At the state at time t, in the Markov decision process model, s i Represents agent b i The arrival rate of inference requests at time t, s i ∈S;

[0041] Action space; C represents the action space, C = {C 1 , C 2 , ..., C k}, where C i = {R i , BZ i} represents agent b i The action space, R i Indicates b i Optional GPU allocation space, BZ i Indicates b i The optional batch size configuration space is c = {c 1 , c 2 , ..., c k} represents agent b 1 , b 2 , …, b k The joint action taken at time t, c i ∈C i ;

[0042] Reward function; Agent b i The reward function is:

[0043]

[0044] in, when When it is less than or equal to 1, the inference latency requirement of the inference workload i is met. Then give agent b i Positive reward; otherwise, give agent b i Negative reward, σ is a pre-specified threshold used to control the limit of the negative reward;

[0045] 2) Based on step 1), define the agent b i At time t, the state is s i When a single action branch The state-action value function is:

[0046]

[0047] in, Represents agent b i The d-th dimension of the action space of Represents agent b i The number of actions that can be selected in the d-dimensional action space, Represents agent b i At time t, the state is s i When , the action selected in the d-dimensional action space; the action space C i The two dimensions correspond to GPU resource allocation and batch size settings, that is, when d=1, the d-th dimension action space corresponds to the optional configuration space of GPU resource allocation, and when d=2, the d-th dimension action space corresponds to the optional configuration space of batch size settings;

[0048] The above state space-action space value function is solved by defining the loss function as follows:

[0049]

[0050] Among them, BF represents the cache of experience replay, h is the number of sub-action spaces of each agent, because each agent b i There are only two sub-action spaces including GPU resource allocation and batch size setting, so h = 2,

[0051]

[0052] Then use gradient descent to minimize the above loss function to obtain the approximately optimal parameter θ, as follows 2

[0053]

[0054] After obtaining the approximately optimal θ, the ε-greedy strategy is used to select actions. Specifically, an action is randomly selected with a probability of ε, and an action is selected with a probability of 1-ε according to the following formula:

[0055]

[0056] According to the above formula, the joint action of each agent at time t is obtained:

[0057]

[0058] Beneficial effects: The resource-efficient and service quality-aware adaptive scheduling method for the inference service system disclosed in the present invention can automatically select a deep learning model according to the user's service quality requirements; it can dynamically allocate GPU resources and adjust the batch size for the operation of the deep learning model according to changes in workload, thereby improving the GPU resource utilization of the inference service system while ensuring the user service quality.

[0059] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0061] Figure 1 A schematic diagram of the resource efficient and service quality aware reasoning service system adaptive scheduling method of the present invention is shown. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0063] The present invention provides a resource-efficient and service quality-aware adaptive scheduling method for an inference service system, which is applied to a container-based inference service system;

[0064] Reference Figure 1 , the adaptive scheduling method includes:

[0065] Build a deep learning model automatic selection module that can analyze the inference workload in the inference service system, and build a deep reinforcement learning-based collaborative management module that can interact with the inference service system in real time;

[0066] The deep learning model automatic selection module uses collaborative filtering methods to predict the inference performance of inference workloads running on different deep learning models. The inference performance includes inference latency and inference accuracy.

[0067] The deep learning model automatic selection module uses a greedy algorithm to select the optimal deep learning model that meets the user's service quality requirements, and deploys the optimal deep learning model to the container to serve the inference workload in the inference service system;

[0068] The collaborative management module uses deep reinforcement learning methods to collaboratively adjust GPU resource allocation and batch size settings according to the dynamic changes of inference workloads based on the optimal deep learning model deployed in the container.

[0069] In an embodiment of the present invention, a deep learning model automatic selection module is constructed that can analyze the inference workload in the inference service system; the deep learning model automatic selection module uses a collaborative filtering method to predict the inference performance of the inference workload running on different deep learning models; the deep learning model automatic selection module uses a greedy algorithm to select the optimal deep learning model that meets the user's service quality requirements; the details are as follows:

[0070] 1) Model knowledge base construction

[0071] The deep learning model in the inference service system is defined by the model architecture, programming framework, and optimization compiler. The model architecture includes ResNet50, VGG16, etc., the programming framework includes TensorFlow, PyTorch, etc., and the optimization compiler includes TensorRT, TVM, etc. For example, ResNet50-TensorFlow-TensorRT-GPU and VGG16-TensorFlow-TVM-GPU for image recognition inference tasks are two different inference models, which are different in inference latency, inference accuracy, etc. Therefore, by collecting deep learning models in the production environment, a model knowledge base V = {v 1 , v 2 , ..., v j , ..., v m}, where v j represents the jth deep learning model;

[0072] 2) Design of reasoning performance prediction method based on collaborative filtering

[0073] The reasoning performance of the new inference workload performed on the deep learning model in the model knowledge base using the collaborative filtering method based on online matrix decomposition is as follows2

[0074] ① Select n users’ inference workload U={u 1 ,u 2 , ..., u i , ..., u m}, where u i represents the inference workload of the i-th user, and then each inference workload u i Analyze all models in the model knowledge base and construct the utility matrix P n×m , where P n×m Each element p in ij represents the inference workload u i In the model knowledge base, model v j For the inference workload u newly submitted by the user, n+1 Select any two deep learning models in the model knowledge base V for a brief analysis and obtain u n+1 The inference performance on these two models is calculated and the utility matrix is ​​inserted to obtain the new utility matrix P (n+1)×m , the above details are as follows:

[0075]

[0076] ②Use collaborative filtering method based on online matrix decomposition to predict the inference workload u n+1 Reasoning performance on other deep learning models in the model knowledge base V;

[0077] Online matrix factorization can map the inference workload and deep learning model into a low-dimensional joint latent factor space. Then, the inner product of the latent factors in this space can be used to obtain the efficiency of each element p in the utility matrix. ij The value of is the inference performance of the inference workload on the deep learning model, which can then predict the inference workload u n+1 The reasoning performance on other deep learning models in the model knowledge base V is as follows:

[0078] The latent inference workload factor mapped by the inference workload and the latent deep learning model factor mapped by the deep learning model are expressed as and Among them, d is much smaller than n+1 and m, and by derivation and To fit the new utility matrix P (n+1)×m ,Right now This can then accurately predict the inference workload u n+1 Inference performance on other deep learning models in the model knowledge base V. In order to derive and First, based on the observed inference workload ui In the deep learning model v j Reasoning performance on p ij , construct the loss function as follows,

[0079]

[0080] in and is the regularization coefficient, and is u i and v j The eigenvector of ,||·|| 2 is the Euclidean norm, g ij express g(·) is a logical function that maps values ​​to [0, 1], Secondly, the stochastic gradient descent method is used to minimize the loss function To deduce and The details are as follows:

[0081]

[0082] where g′ ij express g′(·) is the derivative of g(·); finally, the inference workload u can be predicted by calculating the corresponding inner product n+1 Reasoning performance on other deep learning models in the model knowledge base V:

[0083] ③Based on the above method ②, the inference workload u can be predicted respectively n+1 Inference latency on other deep learning models in model knowledge base V and inference accuracy Then we get the inference workload u n+1 The set of inference latencies on all deep learning models in the model repository V and the inference accuracy set

[0084] ④Based on the inference delay set obtained in ③ above and the inference accuracy set We calculate the inference workload u n+1 Inference latency on each deep learning model and inference accuracy Reasoning about latency requirements in user quality of service and reasoning accuracy requirements difference and Get the set S L={L (n+1)j} and S A ={A (n+1)j}, then delete the elements less than 0 in the set to get the set S′ L ={L (n+1)j′} and S′ A ={A (n+1)j″}, then, a greedy algorithm is used to select the optimal model that meets the user's service quality requirements, as follows:

[0085] (i) For the set S′ A ={A (n+1)j″} in ascending order, and also sort the elements of the set S′ L ={L (n+1)j′} to sort the elements in ascending order;

[0086] (ii) From S′ A ={A (n+1)j″} Take out the model v′ corresponding to the first element 1 , then to the set S′ L ={L (n+1)j′} to find out whether the model exists. If so, the model is the optimal model. If not, the optimal model is selected from S′ A ={A (n+1)j″} to extract the second element v′ 2 , repeat the above process until you find the A ={A (n+1)j″} and S′ L ={L (n+1)j′}, this model is the optimal model v that satisfies the user service quality. optimal .

[0087] In an embodiment of the present invention, a collaborative management module based on deep reinforcement learning is constructed, which can interact with the inference service system in real time; the collaborative management module uses the deep reinforcement learning method, based on the optimal deep learning model deployed in the container, and according to the dynamic changes of the inference workload, collaboratively adjusts the GPU resource allocation and batch size settings, as follows:

[0088] After determining the optimal deep learning model according to the user's service quality requirements, the determined deep learning model is deployed in the container and provides services for the user's inference workload; however, the user's inference workload changes dynamically over time, so the inference service system needs to obtain the runtime information of the inference workload in real time and send the runtime information as a feedback state value to the collaborative management module, and then dynamically adjust the GPU resources allocated to the container carrying the deep learning model and the batch size for processing inference requests through the collaborative management module, so as to meet the user's service quality requirements while maximizing the GPU resource utilization; since the inference accuracy of the deep learning model is determined after it is determined, we only need to consider the inference delay when coordinating the GPU resources and batch size;

[0089] Consider that there are k containers in the inference service system providing inference services for k user inference workloads. The goal to be achieved is to maximize the GPU resource utilization of k containers while meeting the inference latency requirements of k user inference workloads. The details are as follows:

[0090]

[0091]

[0092] in represents the GPU resources actually used by the i-th inference workload, represents the GPU resources allocated to the container hosting the i-th inference workload, Indicates the GPU resource allocation for the i-th inference workload and batch size BZ i The inference latency under, D represents the total amount of GPU resources in the system, represents the inference latency requirement of the i-th inference workload;

[0093] The above problems are solved through deep reinforcement learning:

[0094] 1) First, build an agent b for each container that carries the inference workload i i , agent b i The GPU resource allocation and corresponding batch size setting of the container responsible for carrying the inference workload i, for each agent b i , according to the Markov decision process modeling, it can be expressed as follows:

[0095] State space; S represents the state space, let s = {s 1 ,s 2 ,...,s k} represents agent b 1 , b2 , …, b k At the state at time t, in this model, s i Represents agent b i The arrival rate of inference requests at time t, s i ∈S;

[0096] Action space; C represents the action space, C = {C 1 , C 2 , ..., C k}, where C i = {R i , BZ i} represents agent b i The action space, R i Indicates b i Optional GPU allocation space, BZ i Indicates b i The optional batch size configuration space is c = {c 1 , c 2 , ..., c k} represents agent b 1 , b 2 , …, b k The joint action taken at time t, c i ∈C i ;

[0097] Reward function; the goal is to maximize the utilization of GPU resources as much as possible while meeting the user's reasoning delay requirements. Therefore, agent b i The reward function is:

[0098]

[0099] in when When it is less than or equal to 1, the inference latency requirement of workload i is met. Now give agent b i Positive reward; otherwise, give agent b i Negative reward, σ is a pre-specified threshold used to control the boundary of negative reward;

[0100] 2) Based on step 1), define agent b i At time t, the state is s i When a single action branch The state-action value function is:

[0101]

[0102] in Represents agent bi The d-th dimension of the action space of Represents agent b i The number of actions that can be selected in the d-dimensional action space, Represents agent b i At time t, the state is s i When the action selected in the d-dimensional action space is C i There are two dimensions: GPU resource allocation and batch size setting, i.e., d = 1 or 2. The first dimension corresponds to the optional configuration space of GPU resource allocation, and the second dimension is the optional configuration space of batch size setting.

[0103] In order to solve the above state-action value function, the loss function is defined as follows:

[0104]

[0105] Where BF represents the experience replay buffer, h is the number of sub-action spaces for each agent, because in our scenario each agent only contains two sub-action spaces: GPU allocation and batch size setting, so h = 2,

[0106]

[0107] Then, gradient descent is used to minimize the above loss function to obtain the approximately optimal parameter θ, as follows:

[0108]

[0109] After obtaining the approximately optimal θ, the ε-greedy strategy is used to select actions. Specifically, an action is randomly selected with a probability of ε, and an action is selected with a probability of 1-ε according to the following formula:

[0110]

[0111] In the actual calculation process, at time t: with probability ε, each agent randomly selects action c; finally, the joint action of each agent at time t is obtained That is, the adjustment actions of the collaborative management module on GPU resource allocation and batch size can be obtained.

[0112] In summary, the present invention discloses a resource-efficient and service quality-aware adaptive scheduling method for an inference service system. Users only need to provide service quality requirements, and the inference service system can automatically select the optimal deep learning model for them according to the user's service quality requirements, thereby improving the usability of the inference service system.

[0113] During the operation of user inference workloads, the inference service system can adaptively and collaboratively adjust the allocation and batch size of GPU resources to maximize GPU resource utilization while meeting user service quality.

[0114] Finally, it should be noted that, in this article, relational terms such as one and another are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include one..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0115] Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent substitutions for some of the technical features therein; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. Resource-efficient and service quality-aware adaptive scheduling method for inference service systems, applied to container-based inference service systems, It is characterized in that The adaptive scheduling method comprises: Build a deep learning model automatic selection module that can analyze the inference workload in the inference service system, and build a deep reinforcement learning-based collaborative management module that can interact with the inference service system in real time; The deep learning model automatic selection module predicts the inference performance of the inference workload running on different deep learning models using a collaborative filtering method, wherein the inference performance includes inference latency and inference accuracy; The deep learning model automatic selection module selects the optimal deep learning model that meets the user's service quality requirements using a greedy algorithm, and deploys the optimal deep learning model into a container to serve the inference workload in the inference service system; The collaborative management module uses a deep reinforcement learning method to collaboratively adjust GPU resource allocation and batch size settings according to the dynamic changes of the inference workload based on the optimal deep learning model deployed in the container; The deep learning model automatic selection module for reasoning workload in the analyzable reasoning service system includes: collecting deep learning models in the production environment to build a model knowledge base V = {v 1 , v 2 , ..., v j , ..., v m }, where v j represents the jth deep learning model; The method of using a collaborative filtering method to predict the inference performance of the inference workload running on different deep learning models includes: ① Select the inference workload U of n users from the historical data of the inference service system = {u 1 ,u 2 , ..., u i , ..., u m }, where u i represents the inference workload of the i-th user; Each inference workload u i Perform analysis on all the deep learning models in the model knowledge base V to construct a utility matrix P n×m , where the utility matrix P n×m Each element p in ij represents the inference workload u i Model v in model knowledge base V j For the inference workload u newly submitted by the user n+1 In the model knowledge base V, any two deep learning models are selected for analysis to obtain u n+1 The inference performance on any two of the deep learning models is calculated and the utility matrix is ​​inserted into the new utility matrix P (n+1)×m , the above specific expressions are as follows: ②Use collaborative filtering method based on online matrix decomposition to predict the inference workload u n+1 The reasoning performance on the deep learning models other than the two deep learning models selected in ① in the model knowledge base V.

2. According to the resource efficient and service quality aware reasoning service system adaptive scheduling method of claim 1, It is characterized in that The specific step ② of using the collaborative filtering method to predict the inference performance of the inference workload running on different deep learning models is as follows: Online matrix factorization maps the inference workload and deep learning model into a low-dimensional joint latent factor space; then the inner product of the joint latent factors in the joint latent factor space is used to obtain the new utility matrix P (n+1)×m Each element p in ij The value of , and then predict the inference workload u n+1 The reasoning performance on the deep learning models other than the two deep learning models selected in ① in the model knowledge base V is specifically: The potential inference workload factor mapped by the inference workload and the potential deep learning model factor mapped by the deep learning model are expressed as and Among them, d is much smaller than n+1 and m; By derivation and Fit the new utility matrix P (n+1)×m ,Right now This can then accurately predict the inference workload u n+1 The reasoning performance on the deep learning models other than the two deep learning models selected in ① in the model knowledge base V; wherein, the derivation and The process is as follows: Based on the observed inference workload u i In the deep learning model v j Reasoning performance on p ij , construct the loss function as follows, in, and is the regularization coefficient, and is u i and v j The eigenvector of ,||·|| 2 is the Euclidean norm, g ij express g(·) is a logical function that maps values ​​to [0, 1], Minimize the loss function using the stochastic gradient descent method again Derivation and The details are as follows: where g′ ij express g′(·) is the derivative of g(·); Finally, by calculating and the inner product, the inference workload u can be predicted n+1 for the inference performance on other deep learning models in the model knowledge base V. The expression is as follows:

3. According to the resource efficient and service quality aware reasoning service system adaptive scheduling method of claim 2, It is characterized in that The method of using a collaborative filtering method to predict the inference performance of the inference workload running on different deep learning models also includes: ③ Based on the above method ②, predict the inference workload u respectively n+1 The inference delay on the other deep learning models in the model knowledge base V except the two deep learning models selected in ① and inference accuracy And get the inference workload u n+1 The set of inference latencies on all deep learning models in the model repository V and the inference accuracy set ④ Based on the inference delay set obtained in ③ above and the inference accuracy set Calculate the inference workload u n+1 The inference latency on each of the deep learning models and the inference accuracy Reasoning about latency requirements in user quality of service and reasoning accuracy requirements difference and Get the set S L ={L (n+1)j } and S A ={A (n+1)j Delete S after} L and S A The elements less than 0 in the set S′ L ={L (n+1)j′ } and S′ A ={A (n+1)j″ }.

4. According to the resource efficient and service quality aware reasoning service system adaptive scheduling method of claim 3, It is characterized in that The method of using a greedy algorithm to select an optimal deep learning model that meets the user's service quality requirements includes: (i) For the set S′ A ={A (n+1)j″ } are arranged in ascending order, and the set S′ L ={L (n+1)j′ } to sort the elements in ascending order; (ii) From S′ A ={A (n+1)j″ } Take out the deep learning model v′ corresponding to the first element 1 , then to the set S′ L ={L (n+1)j′ } to find out whether the deep learning model v′ exists 1 ; If it exists, then the deep learning model v′ 1 is the optimal model; If it cannot exist, then from S′ A ={A (n+1)j″ } to extract the second element v′ 2 , repeat to set S′ L ={L (n+1)j′ } until it finds a pair that exists in S′ A ={A (n+1)j″ } and S′ L ={L (n+1)j′ }, wherein the first deep learning model is the optimal model v that satisfies the user service quality. optimal .

5. According to the resource efficient and service quality aware reasoning service system adaptive scheduling method of claim 1, It is characterized in that The construction of a collaborative management module based on deep reinforcement learning that can interact with the reasoning service system in real time includes: Assume that there are k containers in the inference service system providing inference services for k user inference workloads. The collaborative management module receives the feedback status of the inference service system. According to the feedback status, the collaborative management module maximizes the GPU resource utilization of the k containers and meets the inference latency requirements of the k user inference workloads. Specifically: in, represents the GPU resources actually used by the i-th inference workload, represents the GPU resources allocated to the container hosting the i-th inference workload, Indicates the GPU resource allocation for the i-th inference workload and batch size BZ i The inference latency under, D represents the total amount of GPU resources in the system, represents the inference latency requirement of the i-th inference workload.

Citation Information

Patent Citations

  • Game strategy optimization method and system and storage medium

    CN111291890A

  • Mobile edge computing unloading method based on multi-agent reinforcement learning

    CN112367353A