Heterogeneous GPU resource pooling method and system combining Ray and Volcano

By building the Ray+Volcano architecture on the Kubernetes cluster, combining Ray's task scheduling and Volcano's resource management, the resource pooling and dynamic scheduling problems in the heterogeneous GPU environment are solved, efficient and reliable computing resource management is achieved, and training efficiency and resource utilization are improved.

CN120276872AInactive Publication Date: 2025-07-08SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510766256.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing distributed computing frameworks such as Ray and Volcano are difficult to achieve resource pooling and dynamic scheduling in heterogeneous GPU environments, resulting in waste of computing resources, low training efficiency, ineffective load balancing, and high concurrent scheduling difficulties.

Method used

Combining the heterogeneous GPU resource pooling methods of Ray and Volcano, by building the Ray+Volcano architecture on the Kubernetes cluster, using Ray's task scheduling framework and Volcano's resource management mechanism, dynamic pooling and intelligent scheduling of GPU resources are achieved to ensure task load balancing and resource optimization allocation.

Benefits of technology

It realizes the rapid execution of large-scale tasks and efficient utilization of resources, has automatic failure recovery and data persistence functions, improves the high reliability and resource utilization of the system, and reduces hardware and operation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276872A_ABST
    Figure CN120276872A_ABST
Patent Text Reader

Abstract

The invention provides a heterogeneous GPU resource pooling method and system combining Ray and Volcano, and relates to the technical field of GPU resource scheduling. According to the method and the system, firstly, a Ray + Volcano framework is constructed on a Kubernetes cluster; then deep learning model training is carried out based on a Ray + Volcano framework, distributed management is carried out on deep learning training tasks through a task scheduling framework of a Ray cluster, and task load balance is ensured; and a resource management mechanism of Volcano job scheduling is utilized to realize control and optimal distribution of GPU resources. According to the method and the system, rapid execution of large-scale tasks and efficient utilization of resources are realized through the efficient distributed computing capability of Ray and the dynamic resource scheduling and isolation mechanism of Volcano.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of GPU resource scheduling, and in particular, to a heterogeneous GPU resource pooling method and system combining Ray and Volcano. Background Art

[0002] Currently, the demand for GPU resources in deep learning training tasks is increasing day by day. However, in a heterogeneous GPU environment, due to differences in different GPU models and performance, traditional resource scheduling methods often struggle to fully utilize the hardware performance, resulting in wasted computing resources or uneven task distribution. These problems lead to low training efficiency. Especially in deep learning tasks that require large-scale distributed computing, how to efficiently manage and schedule heterogeneous GPU resources has become an urgent technical problem.

[0003] Existing distributed computing frameworks such as Ray and Volcano already provide task scheduling and resource management functions. However, in a heterogeneous GPU environment, they still face challenges in resource pooling and dynamic scheduling. Existing solutions are difficult to meet the requirements of intelligent task allocation under multiple GPU models, and cannot effectively balance the load and improve training efficiency. Therefore, how to combine the resource scheduling capabilities provided by Ray and Volcano to optimize the management of computing resources in a heterogeneous GPU environment has become an important direction for improving deep learning training efficiency.

[0004] With the wide application of deep learning and machine learning, training complex neural network models requires a large amount of computing resources, especially GPUs. However, in a heterogeneous GPU environment, the resource requirements of computing tasks and the performance differences of GPUs are relatively large, which brings the following challenges: Difficult scheduling of heterogeneous GPU resources: Different models and different performance GPUs may perform differently in computing tasks. Traditional GPU resource scheduling methods often assume the same GPU type, making it difficult to fully utilize the computing power of heterogeneous GPUs.

[0005] Waste of GPU resources: Due to the performance differences between different GPUs, some high-performance GPUs may be inefficiently used, while low-performance GPUs may be over-allocated resources, resulting in low computing efficiency.

[0006] Low training efficiency: In a multi-node, heterogeneous GPU environment, it is difficult to reasonably balance task allocation and computing load, resulting in some nodes being overloaded and other nodes being idle, thus prolonging the model training time.

[0007] Difficult high-concurrency scheduling: In large-scale training tasks, traditional resource scheduling mechanisms cannot efficiently handle task scheduling and resource allocation in a heterogeneous GPU environment, resulting in wasted resources and computing bottlenecks.

[0008] With the rapid development of artificial intelligence, big data processing, and high-performance computing technologies, the Graphics Processing Unit (GPU) has become an essential computing resource in heterogeneous computing environments. To improve the utilization efficiency of GPU resources and reduce computing resource costs, current mainstream technical solutions usually adopt distributed computing frameworks or batch processing scheduling systems to manage and schedule GPU resources.

[0009] Existing distributed computing frameworks, such as Ray, possess basic resource scheduling and management capabilities. Ray mainly manages GPU resources through resource declarations and label management, and can dynamically allocate GPU resources on different nodes according to task requirements, enabling automatic scheduling and expansion of distributed computing tasks. In scenarios such as distributed inference and deep learning training, Ray can achieve task-level parallel scheduling and resource isolation, improving the utilization rate of computing resources. However, Ray focuses more on task scheduling in real-time tasks and microservice architectures, and has limited support for large-scale batch processing tasks and complex resource preemption mechanisms.

[0010] On the other hand, existing batch processing scheduling systems, such as Volcano, as a batch processing job scheduling framework based on Kubernetes, can achieve functions such as task priority, resource preemption, and fair scheduling through the scheduling plugin mechanism. Volcano is more suitable for processing large-scale and complex-dependent batch processing jobs and is widely used in high-load scenarios such as AI training and scientific computing. It realizes batch scheduling and job queuing of GPU resources by integrating the node scheduling and resource management mechanisms of Kubernetes. However, Volcano's scheduling mainly focuses on offline job processing and batch resource allocation, and has insufficient support for task scheduling with high real-time requirements, making it difficult to meet the dynamic scheduling requirements of GPU resources in scenarios such as online inference and stream processing.

[0011] The existing Ray and Volcano frameworks act independently in terms of scheduling architecture and resource management model, lacking unified scheduling control and resource pooling capabilities, resulting in the following defects and deficiencies in GPU resource scheduling and management: Low resource scheduling efficiency in heterogeneous GPU environments: Most existing technologies assume that GPU resources are homogeneous computing units, failing to fully consider the differences in computing power, video memory capacity, bandwidth, etc. among different models of GPUs (such as A100, V100, RTX3090, etc.). There is a lack of resource abstraction and scheduling optimization models based on heterogeneous GPU capability metrics, leading to a low matching degree between tasks and GPUs, low GPU utilization rate, and serious resource fragmentation problems.

[0012] Separation of batch task and real-time task scheduling: Ray and Volcano schedule different types of tasks independently, and cannot implement a unified scheduling strategy and resource sharing. In a hybrid load scenario, there is no dynamic coordination of GPU resources between batch tasks and real-time inference services. The resource allocation is inflexible, which is likely to cause high-priority tasks to wait and delay, affecting the overall throughput rate and service quality of the system.

[0013] Insufficient GPU resource preemption and dynamic elasticity: Although Volcano has a basic task preemption mechanism, it has limited ability to dynamically preempt GPU resources and release them to high-priority tasks, and it is difficult to meet the requirements of multi-task collaborative preemption and priority scheduling in complex business scenarios. Ray's resource scheduling strategy is biased towards static configuration, lacking the elastic scaling ability of GPU resources in the face of burst load scenarios, and unable to dynamically reclaim the resources occupied by low-priority tasks, affecting task execution efficiency and system response ability.

[0014] Lack of unified GPU resource pooling and sharing mechanism: Existing scheduling systems independently maintain resource information based on their respective frameworks, lacking a unified heterogeneous GPU resource pooling management mechanism, and unable to achieve a global unified view and scheduling control of GPU resources. The resource island effect is obvious among GPU nodes, and it is difficult to further improve resource utilization.

[0015] Insufficient scalability and intelligence of the scheduling system: Existing scheduling systems have limited intelligent support for resource scheduling and cannot dynamically adjust scheduling strategies according to real-time business requirements, GPU performance metrics, energy consumption models, etc. There are obvious deficiencies in intelligent scheduling algorithms, scalability of scheduling strategies, and cross-frame scheduling coordination in existing systems, and they cannot meet the requirements of flexible and intelligent resource scheduling in complex business scenarios in a large-scale heterogeneous GPU environment. Summary of the Invention

[0016] The technical problem to be solved by the present invention is to provide a heterogeneous GPU resource pooling method and system combining Ray and Volcano to achieve dynamic pooling and intelligent scheduling of GPU resources in view of the above-mentioned deficiencies of the prior art.

[0017] To solve the above technical problems, the technical solutions adopted by the present invention are as follows: On the one hand, the present invention provides a heterogeneous GPU resource pooling method combining Ray and Volcano, including: Construct a Ray+Volcano architecture on the Kubernetes cluster; Based on the Ray+Volcano architecture, perform deep learning model training, and use the task scheduling framework of the Ray cluster to manage deep learning training tasks in a distributed manner to ensure task load balancing; Utilize the resource management mechanism of Volcano job scheduling to achieve the control and optimized allocation of GPU resources.

[0018] Furthermore, the method specifically includes: Step 1: Collect data through on-site devices to generate an input data set for the deep learning model or directly use an existing data set as the input data set for the deep learning model; Step 2: Deploy a Kubernetes cluster; Step 3: Install GPU-related components on the Kubernetes cluster, integrate the GPU into the Kubernetes cluster, and make it available as a schedulable resource for containerized workloads; Step 4: Deploy the Ray+Volcano architecture on the Kubernetes cluster; Step 5: Package the deep learning model and the environment required for model training into an image and deploy containers on the Kubernetes cluster; Step 6: Use the obtained data set to train the deep learning model on the Kubenetes cluster using the Ray+Volcano architecture to achieve the control and optimized allocation of GPU resources.

[0019] Furthermore, the Kubernetes cluster deploying the Ray+Volcano architecture includes a Ray cluster, Volcano, and GPU; the Ray cluster runs in the Kubernetes cluster; Ray nodes are deployed on Kubernetes nodes; the Ray cluster performs distributed training on the deep learning model; Volcano collaborates with the Ray cluster to uniformly manage the resources of the Kubernetes cluster; Volcano is responsible for scheduling jobs to Kubernetes nodes with corresponding GPU resources; the Ray nodes deployed on Kubernetes nodes use the GPU to complete distributed computing tasks.

[0020] Furthermore, the Ray cluster consists of a Ray Driver, a Ray Head, and multiple Ray Worker nodes; the Ray Driver is the entry point for users to submit tasks, responsible for task submission, scheduling, and coordination; the Ray Head is the head node of the Ray cluster, undertaking functions such as global scheduling, resource management, task distribution, cluster initialization, resource management, node management, cluster status monitoring and recording, task reception and parsing, task scheduling strategy formulation, and task distribution and execution coordination, and is responsible for collecting and managing the status information of all Worker nodes; the Ray Worker is the working node that actually executes tasks and executes the computing tasks distributed by the Ray Head; the Ray cluster runs in the Kubernetes cluster, supports automatic scaling through the Autoscaler, dynamically creates and destroys the Pods of the Ray Head and Ray Worker according to task requirements, and realizes the utilization and scheduling of resources.

[0021] Furthermore, the Volcano consists of a Controller Manager, a Volcano Scheduler, and a resource queue Queue; the Controller Manager consists of a Job CM, a PodGroup CM, and a Queue CM, and is responsible for the lifecycle management of Jobs, PodGroups, and Queues. Among them, the Job CM manages the status of batch processing tasks, the PodGroup CM manages the Pod groups that need to be scheduled simultaneously, and the Queue CM is responsible for queue resource allocation to ensure fair resource allocation according to priorities; the Volcano Scheduler is a custom scheduler, responsible for making scheduling decisions based on the information of the Controller Manager, scheduling tasks to Kubernetes nodes, and supporting preemptive scheduling, quota management, and priority scheduling strategies; the resource queue Queue is a resource pool defined by Volcano, including GPU resources of different models, and tasks queue up waiting for scheduling according to requirements to realize the management and utilization of heterogeneous GPU resources.

[0022] Furthermore, the process of the Ray cluster performing distributed training on the deep learning model is as follows: Cluster startup: Initialize the Ray cluster and start it in the Kubernetes cluster in the form of containers. The Ray cluster interacts with the Kubernetes cluster, and the Volcano allocates node resources as needed to uniformly manage the cluster resources; Task Definition: Define the Ray distributed training task and specify the required resources. Ray schedules the task to Ray Worker nodes, and Volcano allocates and isolates resources in the Kubernetes cluster according to the task specifications; Task Scheduling and Execution: Ray distributes tasks to available Ray Worker nodes according to the task resource requests to start training. After the training is completed, the Worker nodes return the results to the Ray head node; Parallel Model Training: Ray distributes different parts of the model to different Ray Worker nodes for parallel training. Volcano ensures the full utilization of node resources and allocates computing resources according to the task scheduling strategy; Result Aggregation and Synchronization: The Ray Driver obtains the calculation results of all Ray Worker nodes and performs synchronization and merging.

[0023] Furthermore, the process of the Ray cluster for distributed training of deep learning models also includes: Fault Tolerance and Failure Recovery: If a certain Ray Worker node fails or crashes, the Ray head automatically detects the node failure and reschedules the task; while Volcano continues to provide resource management and scheduling services for Ray to ensure that the task is rescheduled to other healthy nodes for execution; Model Evaluation and Saving: The Ray Driver performs model evaluation by obtaining the training results on each Worker node.

[0024] On the other hand, the present invention also provides a heterogeneous GPU resource pooling system combining Ray and Volcano, including: Architecture Construction Module: Construct the Ray+Volcano architecture on the Kubernetes cluster; Model Training Module: Perform deep learning model training based on the Ray+Volcano architecture, and perform distributed management of deep learning training tasks through the task scheduling framework of the Ray cluster to ensure task load balancing; Resource Management and Optimization Module: Utilize the resource management mechanism of Volcano job scheduling to achieve control and optimal allocation of GPU resources.

[0025] The beneficial effects of adopting the above technical solution are as follows: An heterogeneous GPU resource pooling method and system combining Ray and Volcano provided by the present invention realizes the rapid execution of large-scale tasks and the efficient utilization of resources through the efficient distributed computing ability of Ray and the dynamic resource scheduling and isolation mechanism of Volcano. It has automatic fault recovery and data persistence functions to ensure the high reliability and fault tolerance of the system; at the same time, the simple programming model and good compatibility reduce the development difficulty and improve the development efficiency. In addition, the flexible resource expansion and significantly improved resource utilization effectively reduce the hardware and operation costs. This efficient, reliable, flexible and economical solution is particularly suitable for scenarios such as machine learning and data analysis that require large-scale distributed computing, and has important practical application value and broad development prospects. Brief Description of the Drawings

[0026] Figure 1 Schematic diagram of the Ray + Volcano architecture deployed in the Kubernetes cluster provided by the embodiment of the present invention; Figure 2 Schematic diagram of the deep learning model training process under the Ray+Volcano architecture provided by the embodiment of the present invention. Detailed Embodiment

[0027] The following combines the drawings and embodiments to further describe in detail the specific implementation manners of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0028] In this embodiment, taking a mechanical processing production environment as an example, the heterogeneous GPU resource pooling method combining Ray and Volcano of the present invention is used to train the deep learning model adopted by the production environment, and then the action signal is obtained through inference by the trained deep learning model, and the inferred action signal is output to the production environment; the production equipment in the production environment performs production operations through the obtained action signal.

[0029] In this embodiment, an heterogeneous GPU resource pooling method combining Ray and Volcano includes: Construct a Ray+Volcano architecture on the Kubernetes cluster; Based on the Ray+Volcano architecture, perform deep learning model training, and efficiently manage the deep learning training tasks in a distributed manner through the task scheduling framework of the Ray cluster to ensure task load balancing; Utilize the resource management mechanism of Volcano job scheduling to achieve precise control and optimal allocation of GPU resources.

[0030] Specifically, it includes the following steps: Step 1: Collect data through on-site devices to generate an input dataset for the deep learning model or directly use an existing dataset as the input dataset for the deep learning model; Step 2: Deploy a Kubernetes cluster; Kubernetes (abbreviated as k8s) is an open-source container orchestration platform that realizes the automated management of containerized applications through cluster deployment. Its core goal is to ensure the high availability, scalability, and elasticity of applications.

[0031] A Kubernetes cluster includes a control plane and worker nodes; the control plane is responsible for cluster decision-making and scheduling, and its core components include kube-apiserver, etcd, kube-scheduler, and kube-controller-manager; kube-apiserver provides API interfaces to handle all cluster operation requests; etcd is used for distributed key-value storage to save cluster configuration and status data; kube-scheduler schedules Pods to appropriate nodes according to resource requirements. Kube-controller-manager monitors the cluster status and performs automated repairs (such as node failure recovery).

[0032] Worker nodes run actual containerized applications, and the components include kubelet, kube-proxy, and a container runtime (such as Docker, containerd); kubelet communicates with the API Server to manage the Pod lifecycle on the node. Kube-proxy realizes service load balancing and network proxy. The container runtime (such as Docker, containerd) executes container operations.

[0033] The core resource objects of a Kubernetes cluster include Pod, Service, Deployment, and StatefulSet; a Pod is the smallest deployment unit and can contain one or more containers; a Service provides a stable network access entry and supports load balancing; a Deployment is a deployment controller for stateless applications and supports rolling updates and rollbacks; a StatefulSet manages stateful applications (such as databases) to ensure stable network identities and storage.

[0034] Step 3: Install GPU-related components on the Kubernetes cluster, integrate the GPU into the Kubernetes cluster, and make it available as a schedulable resource for containerized workloads; Step 4: Deploy the Ray+Volcano architecture on the Kubernetes cluster, as Figure 1 shown; Ray is an open-source distributed computing framework designed to simplify parallel and distributed computing in Python. It provides high-performance distributed computing capabilities while maintaining a user-friendly interface, suitable for scenarios that require rapid scaling and parallelization.

[0035] The core features of Ray include distributed object storage, elastic resource scheduling, and distributed tracing and debugging; distributed object storage supports sharing data between different nodes, facilitating the transfer of large-scale data between tasks. Elastic resource scheduling dynamically allocates and releases resources, improving resource utilization. Distributed tracing and debugging provide real-time monitoring and debugging tools to help developers locate performance bottlenecks.

[0036] The application scenarios of Ray include machine learning training, large-scale data processing, and asynchronous task processing; machine learning training accelerates model training and supports distributed training tasks. Large-scale data processing provides rich APIs for handling data loading, transformation, and aggregation. Asynchronous task processing efficiently handles a large number of concurrent tasks, improving execution efficiency.

[0037] Volcano is a cloud-native batch computing platform based on Kubernetes, focusing on high-performance computing task scheduling. It is the first batch computing project of the CNCF, aiming to optimize Kubernetes' support for scenarios such as AI and big data.

[0038] The core features of Volcano include multi-queue management, advanced scheduling policies, and heterogeneous device management; multi-queue management supports multi-tenant resource sharing and resource planning. Advanced scheduling policies such as Gang Scheduling, DRF (DominantResource Fairness), etc., improve cluster resource utilization. Heterogeneous device management supports the efficient utilization of heterogeneous computing resources such as GPUs and NPUs.

[0039] The application scenarios of Volcano include AI model training, big data processing, and high-performance computing; AI model training optimizes the resource allocation of distributed training tasks. Big data processing supports the resource scheduling of frameworks such as Spark and Flink. High-performance computing scenarios such as gene analysis, scientific simulation, and rendering.

[0040] In this embodiment, the Kubernetes cluster after deploying the Ray+Volcano architecture includes a Ray cluster, Volcano, and GPUs; the Ray cluster runs in the Kubernetes cluster; the Ray cluster performs distributed training on deep learning models; Volcano collaborates with the Ray cluster to uniformly manage the resources of the Kubernetes cluster; Volcano is responsible for scheduling jobs to Kubernetes nodes with corresponding GPU resources; Ray uses GPUs to complete distributed computing tasks.

[0041] The Ray cluster consists of a Ray Driver, a Ray Head, and multiple Ray Worker nodes; the Ray Driver is the entry point for users to submit tasks, responsible for task submission, scheduling, and coordination; the Ray Head is the head node of the Ray cluster, undertaking core functions such as global scheduling, resource management, task distribution, cluster initialization, resource management, node management, cluster status monitoring and recording, task reception and parsing, task scheduling strategy formulation, and task distribution and execution coordination, and is responsible for collecting and managing the status information of all Worker nodes; the Ray Worker is the working node that actually executes tasks and executes the computing tasks distributed by the Ray Head; the Ray cluster runs in the Kubernetes cluster, supports automatic scaling through the Autoscaler, dynamically creates and destroys the Pods of the Ray Head and Ray Worker according to task requirements, and realizes elastic utilization and efficient scheduling of resources.

[0042] The Volcano is a high-performance batch computing (HPC) and AI task scheduler on Kubernetes, which makes up for the deficiencies of the native Kubernetes scheduling in aspects such as large-scale parallel jobs and resource scheduling queues. Volcano consists of a Controller Manager, a Volcano Scheduler, and a resource queue Queue; the Controller Manager consists of a Job CM, a PodGroup CM, and a Queue CM, which is responsible for the lifecycle management of Jobs, PodGroups, and Queues. Among them, the Job CM manages the status of batch tasks, the PodGroup CM manages the group of Pods that need to be scheduled simultaneously, and the Queue CM is responsible for queue resource allocation to ensure fair resource allocation according to priorities; the Volcano Scheduler is a custom scheduler that is responsible for making scheduling decisions based on the information of the Controller Manager, efficiently scheduling tasks to appropriate Kubernetes nodes, and supporting preemptive scheduling, quota management, and priority scheduling strategies. The resource queue Queue is a resource pool defined by Volcano. In this example, it includes GPU resources of different models such as A100, H100, and t4. Tasks queue up according to requirements and wait for scheduling to achieve refined management and utilization of heterogeneous GPU resources.

[0043] The GPU (vGPU) is an actual hardware resource pool that includes physical GPUs or virtual GPUs (vGPUs). Volcano is responsible for scheduling jobs to Kubernetes nodes with corresponding GPU resources; the computing tasks of Ray are allocated resources through Kubernetes nodes and finally run on these Kubernetes nodes with GPUs to complete distributed computing tasks.

[0044] Step 5: Package the deep learning model and the environment required for model training into an image and deploy a container in the Kubernetes cluster; Step 6: Use the obtained dataset to train the deep learning model in the Kubenetes cluster using the Ray+Volcano architecture, as Figure 2 shown, to achieve control and optimized allocation of GPU resources; Step 7: Obtain an action signal through inference using the trained deep learning model and output the inferred action signal to the production environment; Step 8: The devices in the production environment operate based on the obtained action signal.

[0045] In this embodiment, the process of the Ray cluster for distributed training of the deep learning model is as follows: Cluster startup: Initialize the Ray cluster through ray.init(); the Ray cluster starts in the form of containers in the Kubernetes cluster environment; when Ray starts, it interacts with the Kubernetes cluster and begins task scheduling according to the Ray cluster configuration (such as resources like GPUs, CPUs, etc.); during this process, Volcano is used as the resource scheduler to ensure that the resources of each Kubernetes node (such as GPUs) can be correctly allocated as needed. The Ray cluster obtains unified management of the cluster resources through cooperation with Volcano; Task definition: Define the Ray distributed training tasks through the @ray.remote decorator, and specify the required resources (such as GPUs or CPUs) for each task; these tasks will be scheduled by Ray to the appropriate Ray Worker nodes in the Ray cluster for calculation; while Volcano is responsible for identifying the resource requirements of each task in the Kubernetes cluster and making appropriate resource allocation and isolation according to the task specifications submitted by Ray (such as the number of GPUs); in this way, Ray can efficiently schedule tasks in the Kubernetes cluster and ensure that each task has sufficient resources.

[0046] Task scheduling and execution: After the Ray tasks are defined and submitted, Ray distributes these tasks to the available Ray Worker nodes according to the resource requests of the tasks (such as GPUs), and starts the deep learning model training process. After the training ends, each Worker node in Ray will return the training results to the head node.

[0047] Parallel model training: Ray assigns different parts of the deep learning model to different worker nodes for training. Volcano continues to play its role during this process to ensure that the resources of each node are fully utilized and calculate resources are correctly allocated according to the task scheduling strategy.

[0048] Result Aggregation and Synchronization: During the distributed training process, each Worker node calculates its own results (such as gradients or model weights). Ray provides a mechanism for aggregating task results, which is jointly completed by the Ray Driver and Ray Workers, allowing the output of tasks to be obtained from all Worker nodes; through ray.get(), the calculation results of all Worker nodes are obtained, and the results are synchronized and merged; Volcano does not directly participate in the aggregation of results, but it ensures that the resources allocated to each node can operate normally according to requirements, so that Ray can efficiently aggregate and synchronize results.

[0049] Fault Tolerance and Failure Recovery: If a Worker node fails or crashes, the Ray head will automatically detect the node failure and reschedule the task. And Volcano continues to provide resource management and scheduling services for Ray, ensuring that tasks can be rescheduled to other healthy nodes for execution. The fault tolerance of Volcano enables the Ray cluster to seamlessly recover resource allocation when node failures occur, ensuring that the training tasks are not affected.

[0050] Model Evaluation and Saving: After the deep learning model training is completed, the Ray Driver evaluates the model by obtaining the training results on each Worker node. Ray collects the training data of all nodes, validates and saves the final model. This process does not involve direct operations of Volcano, but during the model training, Volcano has been ensuring the resource allocation and isolation, so that the training process can be completed smoothly.

[0051] In this embodiment, a heterogeneous GPU resource pooling system combining Ray and Volcano includes: Architecture Construction Module: Build the Ray+Volcano architecture on the Kubernetes cluster; Model Training Module: Based on the Ray+Volcano architecture, perform deep learning model training, and use the task scheduling framework of the Ray cluster to manage the deep learning training tasks distributively to ensure task load balancing; Resource Management and Optimization Module: Utilize the resource management mechanism of Volcano job scheduling to achieve control and optimal allocation of GPU resources.

[0052] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A heterogeneous GPU resource pooling method combining Ray and Volcano, characterized in that: Including: Build the Ray+Volcano architecture on the Kubernetes cluster; Perform deep learning model training based on the Ray+Volcano architecture, and perform distributed management of deep learning training tasks through the task scheduling framework of the Ray cluster to ensure task load balancing; Utilize the resource management mechanism of Volcano job scheduling to achieve control and optimal allocation of GPU resources.

2. A heterogeneous GPU resource pooling method combining Ray and Volcano according to claim 1, characterized in that: The method specifically includes: Step 1: Collect data through on-site devices to generate an input dataset for the deep learning model or directly use an existing dataset as the input dataset for the deep learning model; Step 2: Deploy the Kubernetes cluster; Step 3: Install GPU-related components on the Kubernetes cluster, integrate the GPU into the Kubernetes cluster, and make it available as a schedulable resource for containerized workloads; Step 4: Deploy the Ray+Volcano architecture on the Kubernetes cluster; Step 5: Package the deep learning model and the required environment for model training into an image and deploy a container on the Kubernetes cluster; Step 6: Train the deep learning model on the Kubenetes cluster using the Ray+Volcano architecture with the obtained dataset to achieve control and optimal allocation of GPU resources.

3. The heterogeneous GPU resource pooling method combining Ray and Volcano according to claim 2, wherein: The Kubernetes cluster deploying the Ray+Volcano architecture includes a Ray cluster, Volcano, and GPUs; the Ray cluster runs in the Kubernetes cluster; Ray nodes are deployed on Kubernetes nodes; the Ray cluster performs distributed training on the deep learning model; Volcano collaborates with the Ray cluster to uniformly manage the resources of the Kubernetes cluster; Volcano is responsible for scheduling jobs to Kubernetes nodes with corresponding GPU resources; the Ray nodes deployed on Kubernetes nodes use GPUs to complete distributed computing tasks.

4. An heterogeneous GPU resource pooling method combining Ray and Volcano according to claim 3, characterized in that: The Ray cluster consists of a Ray Driver, a Ray Head, and multiple Ray Worker nodes; the Ray Driver is the entry point for users to submit tasks, responsible for task submission, scheduling, and coordination; the Ray Head is the head node of the Ray cluster, undertaking functions such as global scheduling, resource management, task distribution, cluster initialization, resource management, node management, cluster status monitoring and recording, task reception and parsing, task scheduling strategy formulation, and task distribution and execution coordination, and is responsible for collecting and managing the status information of all Worker nodes; the Ray Worker is the working node that actually executes tasks, executing the computing tasks distributed by the Ray Head; the Ray cluster runs in the Kubernetes cluster, supports automatic scaling through the Autoscaler, dynamically creates and destroys the Pods of the Ray Head and Ray Worker according to task requirements, and realizes the utilization and scheduling of resources.

5. A heterogeneous GPU resource pooling method combining Ray and Volcano according to claim 4, characterized in that: The Volcano consists of a Controller Manager, a Volcano Scheduler, and a resource queue Queue; the Controller Manager consists of a Job CM, a PodGroup CM, and a Queue CM, responsible for the lifecycle management of Jobs, PodGroups, and Queues. Among them, the Job CM manages the status of batch tasks, the PodGroup CM manages the Pod groups that need to be scheduled simultaneously, and the Queue CM is responsible for queue resource allocation to ensure fair resource allocation according to priorities; the Volcano Scheduler is a custom scheduler, responsible for making scheduling decisions based on the information of the Controller Manager, scheduling tasks to Kubernetes nodes, and supporting preemptive scheduling, quota management, and priority scheduling strategies; the resource queue Queue is the resource pool defined by Volcano, including GPU resources of different models, and tasks queue up according to requirements waiting for scheduling to realize the management and utilization of heterogeneous GPU resources.

6. The heterogeneous GPU resource pooling method combining Ray and Volcano according to claim 5, characterized in that: The process of the Ray cluster performing distributed training on a deep learning model is as follows: Cluster startup: Initialize the Ray cluster and start it in the Kubernetes cluster in the form of a container. The Ray cluster interacts with the Kubernetes cluster, and the Volcano allocates node resources as needed to uniformly manage the cluster resources; Task definition: Define the Ray distributed training task and specify the required resources. The Ray schedules the task to the Ray Worker node, and the Volcano allocates and isolates resources in the Kubernetes cluster according to the task specifications; Task scheduling and execution: Ray distributes tasks to available Ray Worker nodes according to task resource requests to start training. After training is completed, the Worker nodes return the results to the Ray head node; Parallel model training: Ray distributes different parts of the model to different Ray Worker nodes for parallel training. Volcano ensures the full utilization of node resources and allocates computing resources according to the task scheduling strategy; Result aggregation and synchronization: The Ray Driver obtains the calculation results of all Ray Worker nodes and synchronizes and merges them.

7. An heterogeneous GPU resource pooling method combining Ray and Volcano according to claim 6, characterized in that: The process of the Ray cluster for distributed training of deep learning models also includes: Fault tolerance and recovery: If a certain Ray Worker node fails or crashes, the Ray head automatically detects the node failure and reschedules the task; while Volcano continues to provide resource management and scheduling services for Ray to ensure that the task is rescheduled to other healthy nodes for execution; Model evaluation and saving: The Ray Driver evaluates the model by obtaining the training results on each Worker node.

8. A heterogeneous GPU resource pooling system that combines Ray and Volcano, implemented based on the heterogeneous GPU resource pooling method described in claim 1, characterized in that: Including: Architecture construction module: Build the Ray+Volcano architecture on the Kubernetes cluster; Model training module: Conduct deep learning model training based on the Ray+Volcano architecture, and perform distributed management of deep learning training tasks through the task scheduling framework of the Ray cluster to ensure task load balancing; Resource management and optimization module: Utilize the resource management mechanism of Volcano job scheduling to achieve control and optimal allocation of GPU resources.

Citation Information

Patent Citations

  • A parallel deep learning scheduling training method and system based on a container

    CN109885389A

  • Deep learning task hybrid deployment method and system

    CN119987974A

Cited By

  • Preemptive distribution method and system for deep learning model training tasks

    CN121579156A