K8s cluster large model reasoning GPU sharing adaptive dynamic scheduling method

Through multivariate linear regression model and middleware proxy services, GPU resources are adaptively configured, which solves the problem of resource waste and system instability of large-scale model inference in Kubernetes clusters, and realizes efficient sharing and dynamic management of GPU resources.

CN120429102APending Publication Date: 2025-08-05NANJING FIBERHOME STARRYSKY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510480678.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The GPUs used by large-model inference in existing Kubernetes clusters are not efficient and cannot adaptively configure GPU resources, resulting in waste of resources and instability of the system.

Method used

The multivariate linear regression model is used to predict the time-consuming request response, and GPU resources are configured adaptively, inferred service instances are dynamically expanded, and business logic is simplified by middleware proxy services, and GPU resources are shared and efficiently managed.

Benefits of technology

It improves the efficiency of GPU resources and the stability of the system, reduces resource waste, simplifies the complexity of the business model, and realizes efficient management of dynamic scaling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429102A_ABST
    Figure CN120429102A_ABST
Patent Text Reader

Abstract

The invention discloses a K8s cluster large model reasoning GPU sharing adaptive dynamic scheduling method, belongs to the field of artificial intelligence, can adaptively configure GPU resources required by large model reasoning, and improves the correctness and use efficiency of GPU configuration in a cluster. According to the adaptive method, GPU resources in the cluster can be correctly configured according to the model type, the reasoning or accelerated reasoning mode, the model parameter quantity and the precision; in addition, the to-be-reasoned task quantity and the reasoning service state are detected in real time, a regression model is designed to predict request processing time according to the request length, the graphics card computing power, the graphics card utilization rate and historical request response time, and according to the regression model, when the request flow is too large or hardware fails, reasoning examples are additionally deployed, and the reasoning efficiency is improved. According to the method, the task processing throughput and the stability and reliability of the system are improved, and when the request quantity is sharply reduced, the deployment reasoning instances are reduced, so that GPU resources are saved, and the use efficiency of the GPU in the system is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a K8s cluster large model inference GPU shared adaptive dynamic scheduling method. Background Art

[0002] The application of large models is becoming more and more widespread, and most small and medium-sized companies only have a small number of old graphics cards. This has also given rise to the development of large model inference methods, such as quantization methods such as GPTQ (Generative Pretrained Transformer Quantization) and AWQ (Aware Weight Quantization), accelerated inference methods such as PagedAttention and FlashAttention, and inference engines such as vLLM and LMDeploy.

[0003] Deploying large-model inference services in Kubernetes (K8s) clusters is a common approach for large-scale inference. However, existing deployment methods inefficiently utilize GPUs in clusters. This is reflected in the adaptive relationship between large-model inference and graphics cards, including model type, parameter count, precision, inference method, graphics card type, supported precision, Tensor Core acceleration type, driver, CUDA, and the dynamic relationship between dynamic changes in user request volume and inference instance scaling. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a K8s cluster large model inference GPU sharing adaptive dynamic scheduling method to address the shortcomings of the background technology, which can adaptively configure the GPU resources required for large model inference and improve the correctness and utilization efficiency of GPU configuration in the cluster.

[0005] The present invention adopts the following technical solutions to solve the above technical problems:

[0006] A method for adaptive dynamic scheduling of GPU sharing for large-scale model inference in a K8s cluster, specifically comprising the following steps:

[0007] Step 1: Obtain business deployment requirements, including model information: large model inference service, model type, parameter quantity, storage accuracy, and inference method;

[0008] Step 2: Obtain GPU information in the cluster, including the number and type of nodes, graphics card driver and CUDA version, and GPU usage;

[0009] Step 3: Based on the obtained model information and GPU information, adaptively complete the GPU configuration of different models and deploy a single inference instance;

[0010] Step 4: Predict the request response time using a multivariate linear regression model based on the request length, graphics card computing power, and graphics card utilization rate.

[0011] Step 5: Dynamically scale the inference service instance based on the multivariate linear regression model to balance user waiting time and GPU utilization efficiency.

[0012] Step 6: When dynamically scaling the multivariate linear regression model, implement statistical prediction of the regression model through the middleware proxy service to avoid increasing the difficulty of business implementation.

[0013] As a further preferred solution of the GPU shared adaptive dynamic scheduling method for large-scale model inference in a K8s cluster of the present invention, in step 4, the form of the multivariate linear regression model is as follows:

[0014] y=β0+β1x1+β2x2+...+β m x m +ε

[0015] Among them, β0 is a constant term, β1, β2, ..., β m is the regression coefficient; the dependent variable y can be approximately expressed as the independent variables x1, x2, ..., x m The linear function of , ε is the residual term after removing the influence of the independent variable on y;

[0016] The matrix form is expressed as:

[0017] Y=Xβ+ε

[0018] Among them, X is called the design matrix, which is x1, x2, ...x m ; β is β0, β1, ..., β m The goal is to minimize the gap between the model predicted value Xβ and the true value Y, that is, to minimize the residual sum of squares to estimate the regression coefficient β; the formula is:

[0019]

[0020] Where S(β) is the residual sum of squares, is the representation of the sum of squares of the residuals, (Y-Xβ) T is the transpose of (Y-Xβ);

[0021] The process of minimizing the residual sum of squares is to solve the partial derivatives and make them equal to 0. The process is as follows:

[0022]

[0023] X T βX-X T Y=0

[0024] XT βX=X T Y

[0025] Get the solution of the equation:

[0026] β'=(X T X) -1 X T Y

[0027]

[0028] in, Indicates that S(β) is the derivative of β, βX is the product of the two, β T represents the transpose of β, X T

[0029] represents the transpose of X, X T Y represents the product of the transpose of X and Y;

[0030] Where β' is the estimated regression coefficient vector, X T is the transpose of the design matrix, (X T X) -1 It's X T The inverse matrix of X;

[0031] During prediction, the word embedding model is introduced to estimate the missing values of the corresponding result length of the large model inference.

[0032] As a further preferred solution of the GPU shared adaptive dynamic scheduling method for large model inference in a K8s cluster of the present invention, in step 1, business deployment requirements are obtained, including model information: large model inference service, model type, parameter quantity, storage precision, and inference method; including the Llama-3.1-8BInstruct model, using vLLM accelerated inference, referring to a Llama-3.1 type model, 8 billion parameters, float16 precision by default, and vLLM inference method.

[0033] As a further preferred solution of the GPU sharing adaptive dynamic scheduling method for K8s cluster large model inference of the present invention, in step 2, GPU sharing is used to enable multiple services to run on the same graphics card.

[0034] As a further preferred solution of the GPU shared adaptive dynamic scheduling method for large-model inference in a K8s cluster of the present invention, in step 3, a REDIS database service is deployed to store tasks to be inferred, task IDs, and inference task response results, to accept user inference requests, and the inference service instance processes the request and returns the inference result.

[0035] As a further preferred solution of the GPU shared adaptive dynamic scheduling method for large-model inference in a K8s cluster of the present invention, in step 4, a MYSQL database service is deployed to store request data, graphics card computing power, model accuracy, graphics card usage rate, current number of inference instances, number of request queue tasks, response result length, and response time, and a multivariate linear regression model is designed to predict the time required to process requests.

[0036] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects:

[0037] 1. The present invention provides a shared adaptive dynamic scheduling method for GPUs in large-model inference in a K8s cluster. This method can adaptively configure the GPU resources required for large-model inference, improving the accuracy and efficiency of GPU configuration in the cluster. The adaptive method can correctly configure GPU resources in the cluster based on the model type, inference or accelerated inference method, model parameter quantity, and accuracy. Given that GPU utilization is generally low in large-model inference tasks, the method implements shared management of GPU resources, allowing different inference instances to be deployed on the same GPU to improve GPU utilization.

[0038] 2. The present invention also detects the amount of tasks to be inferred and the status of the inference service in real time. Based on the request length, graphics card computing power, graphics card utilization, and historical request response time, a regression model is designed to predict the time required to process the request. Based on the regression model, when the request flow is too large or there is a hardware failure, more inference instances are deployed to improve task processing throughput and system stability and reliability. When the request volume drops sharply, fewer inference instances are deployed to save GPU resources and further improve the efficiency of GPU usage in the system.

[0039] 3. When using HPA (Horizontal Pod Autoscaler) for automatic scaling, the present invention adopts a middleware proxy service method to implement more complex statistical prediction models such as regression, avoiding increasing the complexity and implementation difficulty of the business model. At the same time, the middleware service can act as a proxy for multiple business models, and can also include request traffic, regression prediction, and other more complex statistical prediction models, making the dynamic scaling of business services in K8s both diverse and simple. The business model only needs to implement business functions, and there is no need to pay attention to complex dynamic scaling logic. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is a flow chart of a K8s cluster large model inference GPU shared adaptive dynamic scheduling method of the present invention;

[0041] Figure 2 This is a schematic diagram of the effect of the inference response time of the multivariate linear regression prediction large model of the present invention;

[0042] Figure 3 This is a flow chart of the middleware proxy service of the present invention. DETAILED DESCRIPTION

[0043] The technical solution of the present invention is further described in detail below with reference to the accompanying drawings:

[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. The present invention is described in detail below based on the drawings and preferred embodiments. The purpose and effect of the present invention will become more clear. It should be understood that the specific embodiments described here are only used to explain the present invention and are not used to limit the present invention.

[0045] A K8s cluster large model inference GPU shared adaptive dynamic scheduling method, such as Figure 1 As shown, the specific steps include:

[0046] Step 1: Obtain business deployment requirements, including model information: large model inference service, model type, parameter quantity, storage accuracy, and inference method;

[0047] Step 2: Obtain GPU information in the cluster, including the number and type of nodes, graphics card driver and CUDA version, and GPU usage;

[0048] Step 3: Based on the obtained model information and GPU information, adaptively complete the GPU configuration of different models and deploy a single inference instance;

[0049] Step 4: Predict the request response time using a multivariate linear regression model based on the request length, graphics card computing power, and graphics card utilization rate.

[0050] Step 5: Dynamically scale the inference service instance based on the multivariate linear regression model to balance user waiting time and GPU utilization efficiency.

[0051] Step 6: When dynamically scaling the multivariate linear regression model, implement statistical prediction of the regression model through the middleware proxy service to avoid increasing the difficulty of business implementation.

[0052] The detailed steps are as follows:

[0053] Prepare a Kubernetes cluster environment that includes GPU devices, as well as components such as the GPU Operator, Training Operator, Docker, and Container RD, as well as MySQL and Redis images.

[0054] Obtain deployment business requirements, including large model inference or accelerated inference deployment requirements, build a service dependency environment, and encapsulate it into an independent container. This includes: the large model name and type structure, number of parameters, inference deployment precision, and inference method. For example, Llama-3.1-8BInstruct uses vLLM accelerated inference, which refers to a Llama-3.1 model with 8 billion parameters, float16 precision by default, and vLLM inference method.

[0055] Building a DOCKER image requires including the independent container dependency environment for inference, such as PYTORCH, CUDA, PYTHON, TRANSFORMERS, vLLM, etc., as well as large model files, inference service code, and inference service startup scripts.

[0056] Based on the model structure, number of parameters, inference accuracy, and inference method, we can estimate the amount of video memory required for a single instance of a large model. For example, the model weights are stored in various precisions, such as float32, float16, int8, and int4. Different GPU models support different precisions, so we estimate that the video memory occupied by the Llama-3.1-8B-Instruct model mentioned above is approximately 16GB when stored in float16.

[0057] For training, the video memory usage includes model weights, optimizers, gradients, and activation states, while for inference, the video memory usage includes model weights and the overhead of forward inference calculations. This can be used to estimate the video memory usage of the Llama-3.1-8BInstruct model during inference.

[0058] Obtain GPU node information within the cluster, including the number of nodes, the number of GPUs, their models, driver versions, CUDA versions, and GPU usage, and use a GPU sharing method to enable multiple services to run on the same graphics card.

[0059] According to the above steps, calculate the required video memory size, number of GPUs, number of nodes, etc. for a single model inference instance. For the video memory required by the Llama-3.1-8B-Instruct model, take the NVIDIA RTX3090 graphics card as an example. Its video memory is 24GB. Therefore, a single inference instance requires at least one GPU and one node, or two such graphics cards with a usage rate below 50%. Generate the required graphics card configuration and start the inference service in the K8s cluster.

[0060] Deploy the REDIS database service to store tasks to be inferred, task IDs, and inference task response results, and to receive user inference requests, process requests with the inference service instance, and return inference results.

[0061] Deploy MYSQL database services to store request data, graphics card computing power, model accuracy, graphics card usage, current number of inference instances, number of request queue tasks, response result length, response time, etc. Design a multivariate linear regression model to predict request processing time;

[0062] In step 4, the multiple linear regression model is of the following form:

[0063] y=β0+β1x1+β2x2+…+β m x m +ε

[0064] Among them, β0 is a constant term, β1, β2, ..., β m is the regression coefficient; the dependent variable y can be approximately expressed as the independent variables x1, x2, ..., x m The linear function of , ε is the residual term after removing the influence of the independent variable on y;

[0065] The matrix form is expressed as:

[0066] Y=Xβ+ε

[0067] Among them, X is called the design matrix, which is x1, x2, ...x m ; β is β0, β1, ..., β m The goal is to minimize the gap between the model predicted value Xβ and the true value Y, that is, to minimize the residual sum of squares to estimate the regression coefficient β; the formula is:

[0068]

[0069] Where S(β) is the residual sum of squares, is the representation of the sum of squares of the residuals, (Y-Xβ) T is the transpose of (Y-Xβ);

[0070] The process of minimizing the residual sum of squares is to solve the partial derivatives and make them equal to 0. The process is as follows:

[0071]

[0072] X T βX-X T Y=0

[0073] X T βX=X T Y

[0074] Get the solution of the equation:

[0075] β'=(X T X) -1 X T Y

[0076]

[0077] in, Indicates that S(β) is the derivative of β, βX is the product of the two, β T represents the transpose of β, X T

[0078] represents the transpose of X, X T Y represents the product of the transpose of X and Y;

[0079] Where β' is the estimated regression coefficient vector, X T is the transpose of the design matrix, (X T X) -1 It's X T The inverse matrix of X;

[0080] During prediction, the word embedding model is introduced to estimate the missing values of the corresponding result length of the large model inference.

[0081] During prediction, this solution introduces a word embedding model to estimate the missing value of the corresponding result length of the large model inference. The effect of multivariate linear regression on predicting the response time of large model inference is as follows: Figure 2 As shown:

[0082] Monitor the request pool size in the REDIS service in real time and use the above regression model to predict response time. When the number of requests accumulates and the response time is long, HPA will be started to automatically increase the inference instance.

[0083] Monitor the inference service status and request pool in real time. If the request pool is empty or has few requests, which is lower than the number of inference instances, delete the inference service in the cluster to release GPU resources. This dynamic deployment method can improve the utilization efficiency of the cluster's GPU resources.

[0084] The middleware proxy service implements the statistical model of HPA expansion and contraction mentioned above, decoupling it from the business model. The flow chart is as follows Figure 3 As shown:

[0085] The middleware proxy service is responsible for implementing indicator calculation statistics and prediction, throwing statistical indicators to business services and forwarding business interfaces.

[0086] Business services implement business functions and receive statistical prediction results from the middleware proxy service. Prometheus monitors the statistical indicators returned by / metrics, thereby enabling automatic scaling of HPA.

[0087] The middleware proxy service avoids coupling with business services and can provide proxies for multiple businesses. The proxy service itself can also provide multiple statistical models, making service deployment in K8s more reasonable and diverse, and the deployment process is relatively simple.

[0088] The present invention can adaptively configure the GPU resources required for large-model reasoning, improving the accuracy and efficiency of GPU configuration in the cluster. The adaptive method can correctly configure the GPU resources in the cluster according to the model type, reasoning or accelerated reasoning method, model parameter quantity, and accuracy. Given that the GPU utilization rate of large-model reasoning tasks is generally low, the method shares and manages GPU resources, allowing different reasoning instances to be deployed on the same GPU to improve GPU utilization.

[0089] The present invention also detects the amount of tasks to be inferred and the status of the inference service in real time. Based on the request length, graphics card computing power, graphics card utilization rate, and historical request response time, a regression model is designed to predict the time required to process the request. According to the regression model, when the request flow is too large or there is a hardware failure, more inference instances are deployed to improve the task processing throughput and the stability and reliability of the system. When the request volume drops sharply, fewer inference instances are deployed to save GPU resources and further improve the efficiency of GPU utilization in the system.

[0090] When using HPA (Horizontal Pod Autoscaler) for automatic scaling, the present invention adopts a middleware proxy service method to implement more complex statistical prediction models such as regression, avoiding increasing the complexity and implementation difficulty of the business model. At the same time, the middleware service can act as a proxy for multiple business models, and can also include request traffic, regression prediction, and other more complex statistical prediction models, making the dynamic scaling of business services in K8s both diverse and simple. The business model only needs to implement business functions, and there is no need to pay attention to complex dynamic scaling logic.

[0091] Those skilled in the art will understand that the above descriptions are merely preferred embodiments of the invention and are not intended to limit the invention. Although the invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will still be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, etc. made within the spirit and principles of the invention shall be included within the scope of protection of the invention. All technical features in this embodiment may be freely combined according to actual needs.

[0092] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for adaptive dynamic scheduling of GPU sharing for large-scale model inference in a K8s cluster, characterized by: The specific steps include: Step 1: Obtain business deployment requirements, including model information: large model inference service, model type, parameter quantity, storage accuracy, and inference method; Step 2: Obtain GPU information in the cluster, including the number and type of nodes, graphics card driver and CUDA version, and GPU usage; Step 3: Based on the obtained model information and GPU information, adaptively complete the GPU configuration of different models and deploy a single inference instance; Step 4: Predict the request response time using a multivariate linear regression model based on the request length, graphics card computing power, and graphics card utilization rate. Step 5: Dynamically scale the inference service instance based on the multivariate linear regression model to balance user waiting time and GPU utilization efficiency. Step 6: When dynamically scaling the multivariate linear regression model, implement statistical prediction of the regression model through the middleware proxy service to avoid increasing the difficulty of business implementation.

2. A K8s cluster large model inference GPU shared adaptive dynamic scheduling method according to claim 1, characterized in that: In step 4, the multiple linear regression model is of the following form: y=β0+β1x1+β2x2+…+β m x m +e Among them, β0 is a constant term, β1, β2, ..., β m is the regression coefficient; the dependent variable y can be approximately expressed as the independent variables x1, x2, ..., x m The linear function of , ε is the residual term after removing the influence of the independent variable on y; The matrix form is expressed as: Y=Xβ+ε Among them, X is called the design matrix, which is x1, x2, ...x m ; β is β0, β1, ..., β m The goal is to minimize the gap between the model predicted value Xβ and the true value Y, that is, to minimize the residual sum of squares to estimate the regression coefficient β; the formula is: Where S(β) is the residual sum of squares, is the representation of the sum of squares of the residuals, (Y-Xβ) T is the transpose of (Y-Xβ); The process of minimizing the residual sum of squares is to solve the partial derivatives and make them equal to 0. The process is as follows: X T βX-X T Y=0 X T βX=X T Y Get the solution of the equation: β'=(X T X) -1 X T Y βX,β T ,X T ,X,X T Y in, Indicates that S(β) is the derivative of β, βX is the product of the two, β T represents the transpose of β, X T represents the transpose of X, X T Y represents the product of the transpose of X and Y; Where β' is the estimated regression coefficient vector, X T is the transpose of the design matrix, (X T X) -1 It's X T The inverse matrix of X; During prediction, the word embedding model is introduced to estimate the missing values of the corresponding result length of the large model inference.

3. The method for adaptive dynamic scheduling of GPU sharing for large-scale model inference in a K8s cluster according to claim 1 is characterized by: In step 1, obtain the business deployment requirements, including model information: large model inference service, model type, parameter number, storage precision, and inference method; including the Llama-3.1-8BInstruct model, using vLLM accelerated inference, referring to the Llama-3.1 type model, 8 billion parameters, default float16 precision, and vLLM inference method.

4. A K8s cluster large model inference GPU shared adaptive dynamic scheduling method according to claim 1, characterized in that: In step 2, GPU sharing is used to enable multiple services to run on the same graphics card.

5. The method for adaptive dynamic scheduling of GPU sharing for large-scale model inference in a K8s cluster according to claim 1 is characterized by: In step 3, deploy the REDIS database service to store the tasks to be inferred, task IDs, and inference task response results. It is used to accept user inference requests, and the inference service instance processes the requests and returns the inference results.

6. A K8s cluster large model inference GPU shared adaptive dynamic scheduling method according to claim 1, characterized in that: In step 4, deploy the MYSQL database service to store request data, graphics card computing power, model accuracy, graphics card usage, the current number of inference instances, the number of request queue tasks, the length of the response result, and the response time. Design a multivariate linear regression model to predict the time it takes to process requests.

Citation Information

Cited By

  • Graphic processor resource management system, method and server

    CN120765447A