Heterogeneous GPU cluster-oriented deep neural network model parallel reasoning method

By adopting a deep reinforcement learning scheduling mechanism in heterogeneous GPU clusters, intelligently matches the DNN model and GPU resources, solving the problem of low inference efficiency in heterogeneous environments, and achieving efficient parallel inference and resource optimization.

CN120297426AActive Publication Date: 2025-07-11SHANDONG UNIV

Patent Information

Application Number
CN202510786813.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-11
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

In a heterogeneous GPU cluster environment, how to intelligently match and deploy GPU resources based on the operation characteristics of different models to improve the parallel inference efficiency and resource utilization of deep neural network models.

Method used

A scheduling mechanism based on deep reinforcement learning is adopted to intelligently match different DNN models and heterogeneous GPU servers through deep reinforcing Q networks and greedy strategies to achieve optimal deployment and parallel inference.

Benefits of technology

It significantly improves the overall parallel efficiency and resource utilization of inference tasks of heterogeneous GPU clusters, reduces the GPU idle rate, improves the throughput and energy efficiency ratio of inference computing, and adapts to the needs of diversified task scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297426A_ABST
    Figure CN120297426A_ABST
Patent Text Reader

Abstract

The invention discloses a deep neural network model parallel reasoning method for a heterogeneous GPU cluster, and relates to the field of distributed machine learning, and the method comprises the steps: obtaining current information, remaining selectable DNN models, deployed DNN models on each GPU server, and GPU servers which do not meet the number constraint of the DNN models; the scheduler selects a DNN model and deploys the DNN model on the selected GPU server, and calculates throughput for executing parallel reasoning at the moment; the combination of the DNN model and the GPU server with the maximum throughput is found, and related information is updated; judging whether the DNN models deployed on the GPU meet the number constraint or not, and updating GPU cluster information until all GPUs meet the specific DNN model number constraint; and repeating the steps until the algorithm converges. According to the method, limited heterogeneous GPU resources are fully utilized, and the DNN model with high compatibility is selected to deploy and execute parallel reasoning, so that the throughput is maximized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of distributed machine learning, and in particular to a method for parallel inference of deep neural network models for heterogeneous GPU clusters. Background Art

[0002] With the rapid development of deep learning technology, the representational ability of deep neural network (DNN) models has been continuously enhanced, greatly promoting the innovation and application expansion of the artificial intelligence industry. DNN models have made breakthroughs in many key fields such as natural language understanding, image recognition, and autonomous driving. A typical example is the successful application of large-scale language models such as ChatGPT. The realization of these applications depends on an efficient model inference calculation process, that is, in a large-scale server cluster, quickly responding to external task requests and making decisions based on pre-trained models. Currently, DNN model inference usually relies on a data center pre-deployed on the server. After receiving a task request, it dynamically loads and executes the corresponding model for real-time inference calculation. With the continuous improvement of task complexity and user concurrency, how to improve the real-time performance and throughput of DNN model inference has become one of the core bottlenecks restricting the service capabilities of deep learning systems.

[0003] To improve the inference efficiency, the industry widely uses GPU clusters to execute parallel inference tasks of DNN models. However, GPU resources in the actual deployment environment are usually heterogeneous, including differences in hardware parameters such as computing power and video memory size, as well as differences in driver compatibility and model execution efficiency. In a heterogeneous GPU cluster environment, different types of DNN models have different degrees of dependence on GPU resources. For example, convolutional neural networks (CNNs) are more dependent on tensor parallel computing capabilities in object detection tasks for autonomous driving, while Transformer models are more focused on memory bandwidth and matrix operation efficiency in machine translation and text generation tasks. Therefore, in a heterogeneous GPU cluster, how to intelligently match and deploy according to the running characteristics of different models and the heterogeneity of GPU resources has become a key challenge in improving the parallel inference efficiency of models.

[0004] Research shows that in the DNN inference stage, reasonably scheduling and combining different DNN models to efficiently utilize heterogeneous GPU resources can significantly improve the parallel execution efficiency of models and the overall throughput of the cluster. Therefore, for the problem of parallel inference of DNN models in a heterogeneous GPU cluster environment, there is an urgent need for a scheduling and deployment method with resource awareness and inference performance optimization capabilities to achieve efficient and scalable DNN inference services. Summary of the Invention

[0005] To overcome the above problems existing in the prior art, the present invention proposes a method for parallel inference of a deep neural network model for a heterogeneous GPU cluster.

[0006] The technical solution adopted by the present invention to solve its technical problems is: a method for parallel inference of a deep neural network model for a heterogeneous GPU cluster, including the following steps: Step 1, obtain the current information, the remaining selectable DNN models, the DNN models already deployed on each GPU server, and the GPU servers that do not meet the DNN model quantity constraint. Step 2, the scheduler selects a DNN model to be deployed on the selected GPU server and calculates the throughput of parallel inference when the currently selected DNN model is deployed on the selected GPU. Step 3, repeat Step 2 until a combination of a DNN model and a GPU server that maximizes the throughput is found, deploy the DNN model to the GPU, and update the relevant information. Step 4, determine whether the DNN models already deployed on the GPU meet the quantity constraint, and update the GPU cluster information until all GPUs meet the specific DNN model quantity constraint. Step 5, repeat Steps 1-4 until the algorithm converges.

[0007] For the above method for parallel inference of a deep neural network model for a heterogeneous GPU cluster, Step 1 is specifically as follows: Denote as the number of heterogeneous GPU servers, Denote as the set of heterogeneous DNN models, Denote as the set of DNN models deployed on GPU server ; the scheduler obtains, through observation, the set of the remaining selectable DNN models, the set of the DNN models already deployed on each GPU server and the set

[0008] For the above method for parallel inference of a deep neural network model for a heterogeneous GPU cluster, Step 2 is specifically as follows: Consider each GPU server as an agent, which has a local deep recurrent Q-network , and through the deep recurrent Q-network , For each of the remaining selectable DNN models in calculate the corresponding action value function ; The DNN model selection decision adopts -greedy strategy.

[0009] The above-mentioned method for parallel inference of a deep neural network model for a heterogeneous GPU cluster, the -greedy strategy is specifically as follows: The scheduler selects an action with a probability of using the corresponding to select an action , that is, deploying the DNN model on the GPU server for optimal parallel inference execution; with a probability of 1 - , the scheduler selects an optional DNN model from and deploys it to a slave selects a GPU server for deployment.

[0010] The above-mentioned method for parallel inference of a deep neural network model for a heterogeneous GPU cluster, the specific update of relevant information in step 3 includes: the set of remaining optional DNN models , the set of DNN models already deployed on each GPU server .

[0011] The above-mentioned method for parallel inference of a deep neural network model for a heterogeneous GPU cluster, step 4 is specifically as follows: Each GPU server has a specific number constraint on the deployed DNN model. After each DNN model is deployed to the GPU server, it is judged whether the DNN models already deployed on the GPU meet the number constraint; if meets the number constraint for the deployed DNN model, the information of the GPU servers in the GPU cluster that do not meet the constraint is updated.

[0012] The beneficial effects of the present invention are: (1) The present invention fully considers the heterogeneity of resources among GPU servers, and on this basis, proposes a multi-DNN model scheduling mechanism based on deep reinforcement learning to achieve the optimal matching between different DNN models and heterogeneous GPUs, effectively improving the overall parallel efficiency and resource utilization rate of the inference task.

[0013] (2) The model parallel inference method proposed by the present invention can automatically combine multiple DNN models and execute them efficiently and concurrently in a heterogeneous GPU cluster, so as to adapt to the requirements of diverse task scenarios, effectively reduce the GPU idle rate, improve the throughput rate and energy efficiency ratio of the inference calculation, and meet the actual deployment requirements of large-scale real-time inference systems. Brief Description of the Drawings

[0014] ​Figure 1 It is the schematic diagram of the process of the present invention; Figure 2 It is the comparison chart of throughput under different numbers of DNN models in the embodiment of the present invention; Figure 3 It is the comparison chart of convergence performance under different numbers of DNN models in the embodiment of the present invention. Detailed implementation manners

[0015] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific implementation manners.

[0016] The present invention provides a method for parallel inference of a deep neural network model for a heterogeneous GPU cluster, as Figure 1 shown, which includes the following steps: Step 1: Obtain the current information, the remaining selectable DNN models, the DNN models already deployed on each GPU server, and the GPU servers that do not meet the DNN model quantity constraint.

[0017] Let represent heterogeneous GPU servers, represent a set of heterogeneous DNN models, represent the set of DNN models deployed on the GPU server . The scheduler obtains the set of the remaining selectable DNN models by observation, the set of the DNN models already deployed on each GPU server and the set of the GPU servers that do not meet the DNN model quantity constraint.

[0018] Step 2: The scheduler selects a DNN model to be deployed on the selected GPU server, and determines whether the parallel inference performed by deploying the currently selected DNN model on the selected GPU maximizes the throughput. Repeat this step until a combination of the DNN model and the GPU server that maximizes the throughput is found, deploy the DNN model to the GPU, and update the relevant information.

[0019] (1) Regard each GPU server as an agent, which has a local deep recurrent Q-network . Through the deep recurrent Q-network , for each remaining selectable DNN model in calculate the corresponding action value function To achieve a trade-off between exploration and exploitation, the DNN model selection decision adopts -greedy strategy. Specifically, the scheduler uses probability to exploit the corresponding selected action , that is, deploy the DNN model on the GPU server to perform parallel inference optimally; with probability 1 - , the scheduler selects an optional DNN model from and deploys it to a slave select a GPU server .

[0020] (2) Update relevant information, that is, the set of remaining optional DNN models , the set of DNN models already deployed on each GPU server . .

[0021] Step 3, after each DNN model is deployed to a GPU server, determine whether the number of DNN models deployed on the GPU meets the quantity constraint, and update the GPU cluster information until all GPUs meet the specific DNN model quantity constraint.

[0022] Each GPU server has a specific quantity constraint for deploying DNN models. After each DNN model is deployed to a GPU server, it is necessary to determine whether the number of DNN models deployed on the GPU meets the quantity constraint. If meets the quantity constraint for deploying DNN models, then it is necessary to update the information of the GPU servers in the GPU cluster that do not meet the constraint.

[0023] Step 4, repeat Steps 1, 2, and 3 until the algorithm converges.

[0024] To comprehensively evaluate the performance of the method proposed in the present invention, systematic experiments were designed and carried out. Five types of typical DNN models were selected for the experiments, covering convolutional neural networks (CNNs), recurrent neural networks (RNNs), Transformer networks, generative adversarial networks (GANs), and graph neural networks (GNNs). The specific selected models include: ResNet50, ResNet101, ResNet152, VGG16, VGG19, DenseNet20, LSTM, GRU, ConvLSTM, ViT, Visformer, VOLO, BEiT, DaViT, DeiT, Vanilla-GAN, CGAN, DCGAN, W-GAN, cGAT, GCN, SageGCN, and StackGCN.

[0025] To verify the adaptability and performance of the method in a heterogeneous computing environment, the experiment was deployed and tested on five different types of NVIDIA GPUs, including: RTX 2080 (8GB), RTX 4090 (24GB), Tesla V100 (32GB), Tesla T4 (16GB), and A100 (40GB). The dataset used was ImageNet, which is widely representative.

[0026] During the experiment, based on the fixed GPU resource settings, DNN models with an increasing number were gradually introduced to evaluate the adaptability and inference efficiency of the scheduling method under different model scales. The specific model configurations are as follows: (1) 10 DNN models: ResNet50, VGG16, LSTM, GRU, ViT, VOLO, Vanilla - GAN, CGAN, cGAT, GCN; (2) 15 DNN models: ResNet50, VGG16, DenseNet201, LSTM, GRU, Con - vLSTM, ViT, VOLO, BEiT, Vanilla - GAN, CGAN, DCGAN, cGAT, GCN, SageGCN; (3) 17 DNN models: ResNet50, ResNet152, VGG16, DenseNet201, LSTM, GRU, Con - vLSTM, ViT, VOLO, BEiT, DEiT, Vanilla - GAN, CGAN, DCGAN, cGAT, GCN, SageGCN; (4) 20 DNN models: ResNet50, ResNet152, VGG16, DenseNet201, LSTM, LSTM, GRU, Con - vLSTM, ViT, VOLO, BEiT, DEiT, Vanilla - GAN, CGAN, DCGAN, W - GAN, cGAT, GCN, SageGCN, StackGCN.

[0027] The experimental results are as Figure 2 and Figure 3 shown. MoD - based represents the method proposed in the present invention, and the remaining methods are used as comparison baselines, including non - learning methods (Random and Greedy) and reinforcement learning methods (VDN - based and QMIX - based). Figure 2It shows the impact of different scheduling algorithms on the concurrent inference efficiency and overall system throughput when deploying multiple DNN models in a heterogeneous GPU cluster. The experiments were conducted in various settings where the number of models was gradually increased with a fixed number of GPUs. The results indicate that the MoD-based method significantly outperforms the comparative methods QMIX-based, VDN-based, Greedy, and Random in all tested scenarios, with the average increase in system throughput being 25%, 31%, 35%, and 51% respectively, fully verifying the significant advantages of the present invention in improving inference efficiency in heterogeneous environments. Figure 3 It further shows the training convergence processes of the MoD-based method and the QMIX-based and VDN-based methods under 15 model configurations, and analyzes their training dynamics through the changing trends of the cumulative reward (i.e., concurrent inference throughput) within 1000 training cycles. The experimental results show that the MoD-based method achieves a faster and more stable convergence process in various settings. Its convergence speed is approximately 1.60 times and 2.00 times higher than that of the QMIX-based and VDN-based methods respectively, and it also outperforms other methods in terms of the final throughput, fully demonstrating the high efficiency and convergence stability of the present invention in complex scheduling tasks.

[0028] The above embodiments are only exemplary embodiments of the present invention and are not used to limit the present invention. Those skilled in the art can make various modifications or equivalent substitutions to the present invention within the essence and protection scope of the present invention, and such modifications or equivalent substitutions should also be regarded as falling within the protection scope of the present invention.

Claims

1. A method for parallel inference of deep neural network models for heterogeneous GPU clusters, characterized in that, It includes the following steps: Step 1: Obtain the current information, the remaining selectable DNN models, the DNN models already deployed on each GPU server, and the GPU servers that do not meet the DNN model quantity constraint; Step 2: The scheduler selects a DNN model to be deployed on the selected GPU server and calculates the throughput of the currently selected DNN model for parallel inference when deployed on the selected GPU; Step 3: Repeat Step 2 until a combination of DNN model and GPU server that maximizes the throughput is found, deploy the DNN model to the GPU, and update the relevant information; Step 4: Determine whether the quantity constraint of the DNN models already deployed on the GPU is met, and update the GPU cluster information until all GPUs meet the specific DNN model quantity constraint; Step 5: Repeat Steps 1-4 until the algorithm converges.

2. The parallel inference method of a deep neural network model for a heterogeneous GPU cluster according to claim 1, wherein Step 1 specifically includes the following: denotes the number of heterogeneous GPU servers, denotes the set of heterogeneous DNN models, denotes the set of DNN models deployed on GPU server ; the scheduler obtains, by observation, the set of remaining optional DNN models , the set of DNN models already deployed on each GPU server , and the set of GPU servers that do not meet the DNN model quantity constraint . .​ 3. A parallel inference method for a deep neural network model oriented to heterogeneous GPU clusters according to claim 2, characterized in that, The specific content of step 2 is as follows: Each GPU server is regarded as an agent, and there is a local deep recurrent Q-network . Through the deep recurrent Q-network , For each remaining optional DNN model calculate the corresponding action value function ; The DNN model selection decision adopts -greedy policy.

4. A parallel inference method for a deep neural network model for a heterogeneous GPU cluster according to claim 3, characterized in that The -greedy policy is specifically as follows: with probability utilize corresponding to select an action , that is, deploy the DNN model on the GPU server for optimal parallel inference execution; with probability 1 - , the scheduler then selects an optional DNN model from and deploys it to the slave select a GPU server for execution.

5. The parallel inference method for a deep neural network model facing a heterogeneous GPU cluster according to claim 2, wherein The specific update of relevant information in step 3 includes: the set of remaining optional DNN models , the set of DNN models deployed on each GPU server . .

6. A parallel inference method for a deep neural network model oriented to a heterogeneous GPU cluster according to claim 1, wherein The specific content of step 4 is as follows: Each GPU server has a quantity constraint for specifically deploying the DNN model. After each DNN model is deployed to the GPU server, it is judged whether the DNN models already deployed on this GPU meet the quantity constraint; if the quantity constraint for deploying the DNN model is met, the information of the GPU servers in the GPU cluster that do not meet the constraint is updated.

Citation Information

Patent Citations

  • Method for accelerating multi-outlet DNN reasoning by heterogeneous processor under edge computing

    CN114662661A

  • GPU (Graphic Processing Unit) resource allocation method for deep learning reasoning performance interference perception

    CN115237586A

  • Server deployment method based on deep reinforcement learning in edge computing

    CN116321189A

  • Two-dimensional intelligent anti-interference decision-making method and system based on Q learning algorithm

    CN116390259A

  • Sequential decision-making method, device and equipment based on dual recurrent neural network

    CN116957053A

Cited By

  • Generative model heterogeneous collaborative reasoning method and system based on intelligent scheduling

    CN121560534A

  • Intelligent scheduling-based generative model heterogeneous collaborative reasoning method and system

    CN121560534B

  • Quantization-aware multi-model online inference scheduling method for heterogeneous GPU cluster

    CN122547552A