The invention discloses a K8s cluster
large model reasoning GPU sharing adaptive dynamic scheduling method, belongs to the field of
artificial intelligence, can adaptively configure GPU resources required by
large model reasoning, and improves the
correctness and use efficiency of GPU configuration in a cluster. According to the
adaptive method, GPU resources in the cluster can be correctly configured according to the model type, the reasoning or accelerated reasoning mode, the
model parameter quantity and the precision; in addition, the to-be-reasoned task quantity and the reasoning service state are detected in real time, a regression model is designed to predict request
processing time according to the request length, the
graphics card computing power, the
graphics card
utilization rate and historical
request response time, and according to the regression model, when the request flow is too large or hardware fails, reasoning examples are additionally deployed, and the reasoning efficiency is improved. According to the method, the task
processing throughput and the stability and reliability of the
system are improved, and when the request quantity is sharply reduced, the deployment reasoning instances are reduced, so that GPU resources are saved, and the use efficiency of the GPU in the
system is further improved.