Multi-target distributed deep learning container scheduling method and system and storage medium

By combining GPU topology awareness, network awareness and load awareness technologies in a distributed deep learning environment, using graph neural networks and deep Q networks to optimize container scheduling, the problems of communication bottlenecks and load imbalance in the existing technology are solved, efficient multi-objective optimization scheduling is achieved, and cluster performance is significantly improved.

CN120029720AActive Publication Date: 2025-05-23ARTIFICIAL INTELLIGENCE RES INST OF HEFEI COMPREHENSIVE NAT SCI CENT (ANHUI ARTIFICIAL INTELLIGENCE LAB) +1

Patent Information

Application Number
CN202510499669.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-23
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

In distributed deep learning operations, existing container orchestrators cannot sense communication performance differences between nodes in fine-grained manner, resulting in communication bottlenecks and load imbalance, and lack of global resource coordination and multi-objective optimization capabilities, making it difficult to meet real-time scheduling needs.

Method used

GPU topology awareness, network awareness and load awareness technology are adopted, combined with graph neural networks and deep Q networks, embedded vectors representing node interaction characteristics are generated, and container scheduling scheme is optimized through reward functions to minimize communication overhead and load balancing.

Benefits of technology

It effectively alleviates the communication bottleneck in distributed deep learning jobs, realizes load balancing between nodes, significantly improves the overall performance of the cluster, and significantly reduces the solution time in real-time scheduling scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029720A_ABST
    Figure CN120029720A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of distributed deep learning, and discloses a multi-target distributed deep learning container scheduling method and system, and a storage medium, and the method comprises the steps: carrying out the modeling of a communication structure between GPUs in nodes, the communication performance in the nodes, and the communication performance between the nodes through employing a graph neural network, generating an embedded vector representing node interaction characteristics, and obtaining a cluster state; inputting the cluster state into the deep Q network, defining a mapping relation from a container to a node through an action space, and optimizing a deep Q network decision based on a reward function so as to minimize communication overhead, balance node loads and train the deep Q network; and deploying the deep Q network model after offline training to a scheduler, and generating a container scheduling scheme according to a real-time cluster state. According to the scheduling method provided by the invention, the communication bottleneck in distributed deep learning job training is effectively relieved while the scheduling real-time performance is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of distributed deep learning technology, and in particular to a multi-objective distributed deep learning container scheduling method, system and storage medium. Background Art

[0002] With the rapid development of artificial intelligence, the parameter scale of deep learning models has continued to expand, from the initial million level to the current trillion level. Distributed deep learning has become an inevitable choice for training complex models. The traditional single-machine training method is limited by the computing power of a single device and cannot meet the training requirements of these complex models. At the same time, cloud computing, with its flexible and efficient resource management capabilities and good scalability, provides an ideal deployment platform for distributed deep learning jobs. However, mainstream container orchestrators have obvious shortcomings when facing distributed deep learning jobs. First, the current container orchestrators cannot perceive the communication performance differences between nodes in a fine-grained manner, which easily leads to communication bottlenecks and "slow node effects", thereby reducing the overall training efficiency. Secondly, they lack accurate perception of the actual load of the nodes, which easily leads to overloading of some nodes, thereby increasing the risk of system failure and affecting the overall utilization of resources. In addition, traditional container scheduling methods mostly use heuristic rules, lack the ability to coordinate the global cluster resources and multi-objective optimization, and easily lead to scheduling decisions falling into local optimality. The traditional multi-objective optimization algorithm takes a long time to solve, is not suitable for real-time scheduling scenarios, and is difficult to meet the needs of deadline (DDL) jobs.

[0003] Therefore, how to achieve efficient multi-objective optimization scheduling in a cloud environment to maximize communication efficiency and improve cluster load balancing has become a key issue that needs to be urgently solved in the process of deploying distributed deep learning jobs on the cloud. Summary of the invention

[0004] In order to solve the above technical problems, the present invention provides a multi-objective distributed deep learning container scheduling method, system and storage medium; the present invention aims at the container scheduling problem of distributed learning jobs on cloud platforms, and designs a real-time scheduling method to alleviate the communication bottleneck in the job training process and improve the overall load balancing of the system.

[0005] In order to solve the above technical problems, the present invention adopts the following technical solutions: A multi-objective distributed deep learning container scheduling method, comprising: The communication structure between GPUs in each node in the distributed cluster and the communication performance within the node are collected through GPU topology perception technology, the communication performance between nodes is obtained through network perception technology, and the actual resource load of the node is monitored through load perception technology. Use graph neural networks to model the communication structure between GPUs within nodes, the communication performance within nodes, and the communication performance between nodes, generate an embedding vector that characterizes the node interaction characteristics, and then splice the embedding vector with the actual resource load of the node and the resource requirements of the container to be scheduled to obtain the cluster status; Input the cluster state into the deep Q network, define the mapping relationship between containers and nodes through the action space, and optimize the container scheduling scheme decided by the deep Q network based on the reward function to minimize the communication overhead and balance the node load, so as to realize the training of the deep Q network; The deep Q network model completed offline training is deployed to the scheduler, and a container scheduling plan is generated based on the real-time cluster status.

[0006] In one embodiment, the collecting of the communication structure between GPUs within each node in the distributed cluster and the communication performance within the node by using the GPU topology awareness technology specifically includes: According to the connection mode between GPUs in the node, the internal communication performance weight is quantified, and the internal communication performance of the node is measured based on the communication structure between GPUs.

[0007] In one embodiment, the obtaining of communication performance between nodes by using network perception technology specifically includes: The communication performance between nodes is measured by the following calculation : ; In the formula, and represent bandwidth weight and delay weight respectively, , , They represent bandwidth value, delay value and packet loss rate respectively.

[0008] In one embodiment, monitoring the actual resource load of the node by using the load sensing technology specifically includes: ; is the actual resource load of the node, where , , CPU resource utilization Weight coefficient, memory resource utilization Weight coefficient and GPU resource utilization The weight coefficient of .

[0009] In one embodiment, the use of a graph neural network to model the communication structure between GPUs within a node, the communication performance within a node, and the communication performance between nodes to generate an embedding vector representing the node interaction characteristics specifically includes: The cluster nodes are regarded as vertices in the graph structure, the communication performance within the nodes is regarded as the vertex attribute, and the communication performance between nodes is regarded as the edge attribute. The interaction features between nodes are extracted through graph convolution operations to generate embedding vectors.

[0010] In one embodiment, the step of concatenating the embedding vector with the actual resource load of the node and the resource demand of the container to be scheduled to obtain the cluster state specifically includes: Cluster Status ; in, Indicates the node resource utilization rate. Indicates the remaining resources of the node. Indicates the container resource requirements, Indicates the container deployment status. They are respectively the embedding vector representing the communication performance within the node and the embedding vector representing the communication performance between nodes.

[0011] In one embodiment, the step of inputting the cluster state into the deep Q network, defining the mapping relationship between containers and nodes through the action space, and optimizing the container scheduling scheme decided by the deep Q network based on the reward function specifically includes: Define each action as scheduling a container onto a node; Based on two optimization objectives, the reward function The scheduling scheme is comprehensively measured by linearly weighting the two optimization objectives: ; In the formula, Indicates the degree of similarity in the communication performance of the nodes deployed by the job. Indicates the similarity of the actual load of each node in the cluster; and are the weight coefficients of the two optimization objectives respectively, ; ; ; in, Indicates the minimum communication performance of the nodes where the job has been deployed. Indicates the maximum communication performance of the nodes deployed by the job; Indicates the minimum load of the cluster node. Indicates the maximum load of the cluster node.

[0012] In one of the embodiments, a continuous optimization process of the deep Q network is also included: the scheduling effect corresponding to the container scheduling solution generated for the real-time cluster status is fed back to the deep Q network model, and the container scheduling solution is continuously optimized through online incremental training.

[0013] A multi-objective distributed deep learning container scheduling system, comprising: Collection module: collects the communication structure between GPUs in each node in the distributed cluster and the communication performance within the node through GPU topology perception technology, obtains the communication performance between nodes through network perception technology, and monitors the actual resource load of the node through load perception technology; Encoding module: Use graph neural networks to model the communication structure between GPUs within nodes, the communication performance within nodes, and the communication performance between nodes, generate an embedding vector that characterizes the node interaction characteristics, and then splice the embedding vector with the actual resource load of the node and the resource requirements of the container to be scheduled to obtain the cluster status; Offline training module: Input the cluster state into the deep Q network, define the mapping relationship between containers and nodes through the action space, and optimize the container scheduling scheme decided by the deep Q network based on the reward function to minimize the communication overhead and balance the node load, so as to realize the training of the deep Q network; Inference deployment module: deploys the offline trained deep Q network model to the scheduler and generates a container scheduling plan based on the real-time cluster status.

[0014] A computer-readable storage medium stores a computer program, which implements the steps of the method of any embodiment when executed by a processor.

[0015] Compared with the prior art, the beneficial technical effects of the present invention are: The present invention makes up for the deficiencies of existing container orchestrators in communication and load perception. It models the communication topology inside and outside the cluster through graph neural networks, and combines load-sensing technology to monitor the utilization of node resources in real time, so that the scheduler can accurately perceive the communication performance and node load status in the cluster, thereby optimizing the scheduling strategy. Furthermore, the present invention uses reinforcement learning algorithms to achieve the coordination and multi-objective optimization of global resources. By combining offline training with online reasoning, the trained model can quickly generate the optimal scheduling strategy in the actual environment, significantly reducing the solution time. In summary, the scheduling method proposed in the present invention effectively alleviates the communication bottleneck in the training of distributed deep learning jobs while ensuring the real-time scheduling, and realizes load balancing between nodes, thereby significantly improving the overall performance of the cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 4 is a flow chart of a method in an embodiment of the present invention.

[0017] Figure 2 It is a schematic diagram of the scheduling method framework in an embodiment of the present invention. DETAILED DESCRIPTION

[0018] A preferred embodiment of the present invention is described in detail below with reference to the accompanying drawings.

[0019] In view of the shortcomings of mainstream container orchestrators in communication perception and load perception, the present invention uses GPU topology perception, network perception and load perception technologies to perform fine-grained perception and modeling of node status in distributed clusters. Specifically, GPU topology perception technology is used to model the communication structure between GPUs within the node, clarify the connection mode between GPUs in each node, and thus improve the perception accuracy of internal communication performance; network perception technology is used to capture the communication characteristics between nodes in the cluster, including performance indicators such as communication bandwidth, latency and packet loss rate, and improve the scheduler's perception of inter-node communication performance; load perception technology monitors the resource utilization of each node in the cluster in real time through monitoring components such as Metrics Server to ensure that the scheduler can perceive the actual load of the node during task scheduling. Secondly, in order to solve the problem that traditional container scheduling methods lack globality and traditional multi-objective optimization algorithms take a long time to solve, the present invention proposes a multi-objective container scheduling method based on deep reinforcement learning (DRL). First, the communication topology in the cluster is modeled using Graph Neural Networks (GNN), and the complex relationships between nodes are extracted and encoded into state vectors, thereby providing an accurate representation of communication performance for the scheduling strategy; then, the optimal container scheduling strategy is learned through the Deep Q Network (DQN) algorithm to achieve the goal of reducing communication overhead and improving cluster load balancing. This method can not only coordinate cluster resources from a global perspective, but also quickly generate scheduling decisions in a dynamic environment, significantly improving scheduling efficiency and system performance.

[0020] The present invention mainly collects the communication topology, node load and container information between cluster nodes, and uses graph neural network to encode the cluster state. Then the encoded state information is used as input, and the reinforcement learning agent makes decisions to generate a feasible scheduling plan to achieve efficient scheduling of containers.

[0021] like Figure 1 As shown, a multi-objective distributed deep learning container scheduling method in the present invention includes the following steps: S1: GPU topology awareness technology is used to collect the communication structure between GPUs in each node in the distributed cluster and the communication performance within the node. Network awareness technology is used to obtain the communication performance between nodes. Load awareness technology is used to monitor the actual resource load of the node.

[0022] S2: Use graph neural networks to model the communication structure between GPUs within nodes, the communication performance within nodes, and the communication performance between nodes, generate an embedding vector that characterizes the node interaction characteristics, and then concatenate the embedding vector with the actual resource load of the node and the resource requirements of the container to be scheduled to obtain the cluster status.

[0023] S3: Input the cluster state into the deep Q network, define the mapping relationship from container to node through the action space, and optimize the container scheduling scheme decided by the deep Q network based on the reward function to minimize the communication overhead and balance the node load, thereby realizing the training of the deep Q network.

[0024] S4: Deploy the deep Q network model that has completed offline training to the scheduler and generate a container scheduling plan based on the real-time cluster status.

[0025] In the technical solution of the present invention, a node is an independent computing unit in a distributed cluster, representing a physical server or virtual machine in the cluster, and has independent computing resources (CPU, GPU, memory) and network interfaces.

[0026] In one embodiment, the step S1 of collecting the communication structure between GPUs in each node in the distributed cluster and the communication performance within the node by using the GPU topology awareness technology specifically includes: According to whether the connection mode between GPUs in the node is NVLink or PCIe, the internal communication performance weight is quantified, and the internal communication performance of the node is measured based on the communication structure between GPUs.

[0027] In one embodiment, the obtaining of communication performance between nodes by using network perception technology specifically includes: The communication performance between nodes is measured by the following calculation : ; In the formula, and represent bandwidth weight and delay weight respectively, , , They represent bandwidth value, delay value and packet loss rate respectively.

[0028] In one embodiment, monitoring the actual resource load of the node by using the load sensing technology specifically includes: ; is the actual resource load of the node, where , , CPU resource utilization , memory resource utilization and the weight coefficient of GPU resource utilization ,and .

[0029] Traditional container orchestrators lack the perception of the communication performance and actual load of nodes in the cluster, resulting in the inability to fully optimize communication efficiency and load balancing when making scheduling decisions. Therefore, the present invention first collects data from the nodes in the cluster through a variety of technical means, and collects these useful indicators to lay the foundation for subsequent algorithm design. Specifically, the GPU device plug-in is used to collect the topological information between the GPUs in each node, and the network performance testing tool is used to collect network performance indicators between cluster nodes. The actual load of each node (such as CPU, GPU and memory utilization) is monitored in real time through monitoring components such as Kubernetes' Metrics Server. These indicators provide comprehensive data support for cluster state encoding and subsequent scheduling scheme design.

[0030] Specifically, Figure 2 As shown in the figure, the improved GPU device plug-in can be deployed on the cluster nodes. The device plug-in can be used to collect the topological structure and communication performance between the GPUs in each node in the cluster, such as the connection mode (NVLink or PCIe) and bandwidth, and use this to measure the communication performance within the node; then, the network measurement tool is used to regularly measure the network performance indicators between nodes, including communication bandwidth, latency, and packet loss rate, to measure the communication performance between nodes. In addition, the Kubernetes Metrics Server is used to monitor the actual resource load of each node in real time, including the actual utilization of CPU, memory, and GPU resources. After collecting these performance indicators, they need to be preprocessed and the data are quantified according to a certain scoring method. For example, the communication performance of the GPUs can be weighted according to their different topological connection methods; the communication cost between nodes can be scored, and the smaller the value of the communication cost, the lower the communication cost.

[0031] In one embodiment, the graph neural network in step S2 is used to model the communication structure between GPUs within the node, the communication performance within the node, and the communication performance between nodes to generate an embedding vector representing the node interaction characteristics, specifically including: The cluster nodes are regarded as vertices in the graph structure, the communication performance within the nodes is regarded as the vertex attribute, and the communication performance between nodes is regarded as the edge attribute. The interaction features between nodes are extracted through graph convolution operations to generate embedding vectors.

[0032] In order to effectively deal with variable-length and high-dimensional state data, it is necessary to encode the cluster state. The present invention uses a graph neural network to process the collected communication performance indicators and encode the communication topology structure inside and outside the cluster nodes. The graph neural network can directly model the irregular communication topology in the cluster, capture the complex interaction characteristics between nodes (such as bandwidth, delay, etc.), and encode them into a state vector. This method can not only retain the global topology information, but also has strong scalability and can adapt to the dynamic changes of cluster scale and communication topology. In addition, the embedded vector output by the graph neural network, the real resource load of each node, and the resource demand vector of each container are spliced ​​to form a complete deep Q network state input. In this way, the state input not only contains the global characteristics of the communication topology inside and outside the cluster, but also combines the resource load information of each node, thereby providing a comprehensive and accurate state description for the deep Q network model. This combination can better support the learning of scheduling strategies, so that the deep reinforcement learning model can simultaneously consider the communication performance and the resource load of the node when performing task scheduling, so as to achieve multi-objective optimization of alleviating communication bottlenecks and load balancing.

[0033] Graph neural networks can effectively process irregular and variable-length cluster structures, extract interactive features in complex networks, and generate embedding vectors. In addition, the actual load information of the node and the information of the container to be scheduled also need to be part of the state. Finally, the embedding vector generated by the graph neural network is spliced ​​with the resource load and container information vector to form a complete state input for the subsequent decision-making process of the reinforcement learning model.

[0034] In one embodiment, the step of concatenating the embedding vector with the actual resource load of the node and the resource demand of the container to be scheduled to obtain the cluster state specifically includes: Cluster Status ; in, Indicates the node resource utilization rate. Indicates the remaining resources of the node. Indicates the container resource requirements, Indicates the container deployment status. They are respectively the embedding vector representing the communication performance within the node and the embedding vector representing the communication performance between nodes.

[0035] Specifically, the cluster status It is necessary to fully express the information of the current cluster and the tasks to be scheduled so that the agent can judge the system status. In order to ensure that the agent fully considers the communication cost in the decision-making process, the design of the state of the present invention not only includes the utilization rate of resources on each node and remaining resources , and also incorporates key communication topology information, which is encoded through graph neural networks to generate structured embedding vectors , which represent the communication performance inside and outside the node, and are used as part of the state of the deep reinforcement learning algorithm. In addition, the cluster state also integrates the resource requirements of the container , related job information, and their current deployment status .

[0036] In one embodiment, in step S3, the cluster state is input into the deep Q network, a mapping relationship between containers and nodes is defined through the action space, and a container scheduling scheme determined by the deep Q network is optimized based on a reward function, specifically including: Define each action as scheduling a container onto a node; Based on two optimization objectives, the reward function The scheduling scheme is comprehensively measured by linearly weighting the two optimization objectives: ; In the formula, Indicates the degree of similarity in the communication performance of the nodes deployed by the job. Indicates the similarity of the actual load of each node in the cluster; and are the weight coefficients of the two optimization objectives respectively, ; ; ; in, Indicates the minimum communication performance of the nodes where the job has been deployed. Indicates the maximum communication performance of the nodes deployed by the job; Indicates the minimum load of the cluster node. Indicates the maximum load of the cluster node.

[0037] "Job deployed node" refers to the cluster node that has been assigned and is running the current job container during the scheduling process.

[0038] In the learning phase of the scheduling strategy, the present invention uses a deep Q network model to design the intelligent agent. In the design of the model action, each action Considered as scheduling a container to a node, the action set represents a scheduling scheme for a group of containers. When the number of containers to be scheduled is m and the total number of nodes is N, the action can be expressed as: ; .

[0039] In addition, in the design of model rewards, since the multi-container scheduling problem is a long-sequence decision problem, the scheduling between containers has a significant correlation. It is difficult to accurately evaluate the overall quality of the agent's decision before all containers are fully scheduled. Therefore, we adopt the reward sharing method and use the final reward as a shared reward for all steps in an epoch. Based on the two optimization objectives, we linearly weight multiple objective function values ​​to comprehensively measure the pros and cons of the scheduling scheme.

[0040] In the process of generating the scheduling strategy, the key step of the present invention is to use the encoded cluster state as input, and then make inference decisions through the reinforcement learning agent. The state vector contains the communication topology characteristics, resource load conditions and container resource requirements of the nodes in the cluster. This information is encoded and embedded by the graph neural network and spliced ​​with the resource vector to form a complete input representation. Based on these inputs, the agent continuously learns the optimal scheduling strategy through the deep Q network model in offline training to achieve the goal of minimizing communication overhead and ensuring node load balancing. In the online reasoning stage, the agent uses these state vectors in real time to generate the optimal scheduling plan for the container. When a new container needs to be scheduled, the agent will select the most suitable node to deploy the container based on the overall state of the current cluster to ensure the communication efficiency and load balancing of the system. Through such a design, the present invention can efficiently perform intelligent scheduling in a dynamic environment and realize real-time optimization of the scheduling plan.

[0041] In one of the embodiments, a continuous optimization process of the deep Q network is also included: a scheduling effect corresponding to the container scheduling plan is generated for the real-time cluster status, and is fed back to the deep Q network model, and the container scheduling plan is continuously optimized through online incremental training.

[0042] Once the agent completes the reasoning decision, the generated container scheduling plan will be actually executed by the scheduler. The scheduler assigns the containers to be scheduled to the appropriate nodes to run according to the scheduling decision output by the agent to optimize the overall performance of the cluster. The scheduling plan output by the reasoning is a container-node mapping set. For example, {0:3,1:2,2:1} means that container 0 should be scheduled to node 3, container 1 should be scheduled to node 2, and container 2 should be scheduled to node 1. During the execution process, the system will provide real-time feedback on the scheduling effect, and use this feedback information to continuously optimize the reinforcement learning model, thereby continuously improving the scheduling strategy and ensuring the efficiency and stability of the scheduling.

[0043] Offline training can fully learn scheduling strategies in complex simulation environments, thereby providing efficient decision support for online reasoning. Compared with traditional multi-objective optimization algorithms, the combination of offline training and online reasoning enables the scheduler to quickly infer the optimal container scheduling strategy based on the real-time cluster status, greatly improving scheduling efficiency and ensuring the stable operation of the system in a dynamic environment.

[0044] It should be understood that, although the steps in the flowcharts of the accompanying drawings of the specification are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts of the accompanying drawings of the specification may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0045] The present invention also provides a multi-objective distributed deep learning container scheduling system. Since the implementation scheme and method of the system to solve the problem are similar, the implementation of the specific system in the embodiment of this specification can refer to the implementation of the aforementioned method, and the repetitions will not be repeated. As used below, the term "module" or "module" can be a combination of software and / or hardware that implements a predetermined function. Although the system described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.

[0046] A multi-objective distributed deep learning container scheduling system, comprising: Collection module: collects the communication structure between GPUs in each node in the distributed cluster and the communication performance within the node through GPU topology perception technology, obtains the communication performance between nodes through network perception technology, and monitors the actual resource load of the node through load perception technology; Encoding module: Use graph neural networks to model the communication structure between GPUs within nodes, the communication performance within nodes, and the communication performance between nodes, generate an embedding vector that characterizes the node interaction characteristics, and then splice the embedding vector with the actual resource load of the node and the resource requirements of the container to be scheduled to obtain the cluster status; Offline training module: Input the cluster state into the deep Q network, define the mapping relationship between containers and nodes through the action space, and optimize the container scheduling scheme decided by the deep Q network based on the reward function to minimize the communication overhead and balance the node load, so as to realize the training of the deep Q network; Inference deployment module: deploys the offline trained deep Q network model to the scheduler and generates a container scheduling plan based on the real-time cluster status.

[0047] The present invention also provides a computer-readable storage medium including instructions, such as a memory including instructions, and the above instructions can be executed by a processor to perform the above method. The storage medium can be a computer-readable storage medium, for example, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0048] It is obvious to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention, and any reference numerals in the claims should not be regarded as limiting the claims involved.

[0049] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. A multi-objective distributed deep learning container scheduling method, characterized in that: include: The communication structure between GPUs in each node in the distributed cluster and the communication performance within the node are collected through GPU topology perception technology, the communication performance between nodes is obtained through network perception technology, and the actual resource load of the node is monitored through load perception technology. Use graph neural networks to model the communication structure between GPUs within nodes, the communication performance within nodes, and the communication performance between nodes, generate an embedding vector that characterizes the node interaction characteristics, and then splice the embedding vector with the actual resource load of the node and the resource requirements of the container to be scheduled to obtain the cluster status; Input the cluster state into the deep Q network, define the mapping relationship between containers and nodes through the action space, and optimize the container scheduling scheme decided by the deep Q network based on the reward function to minimize the communication overhead and balance the node load, so as to realize the training of the deep Q network; The deep Q network model completed offline training is deployed to the scheduler, and a container scheduling plan is generated based on the real-time cluster status.

2. A multi-objective distributed deep learning container scheduling method according to claim 1, characterized in that: The GPU topology perception technology is used to collect the communication structure between GPUs in each node in the distributed cluster and the communication performance within the node, specifically including: According to the connection mode between GPUs in the node, the internal communication performance weight is quantified, and the internal communication performance of the node is measured based on the communication structure between GPUs.

3. According to claim 1, a multi-objective distributed deep learning container scheduling method is characterized in that: The obtaining of the communication performance between nodes by using the network perception technology specifically includes: The communication performance between nodes is measured by the following calculation : ; In the formula, and represent bandwidth weight and delay weight respectively, , , They represent bandwidth value, delay value and packet loss rate respectively.

4. The multi-objective distributed deep learning container scheduling method according to claim 1, characterized in that: The monitoring of the actual resource load of the node by load sensing technology specifically includes: ; is the actual resource load of the node, where , , CPU resource utilization Weight coefficient, memory resource utilization Weight coefficient and GPU resource utilization The weight coefficient of .

5. The multi-objective distributed deep learning container scheduling method according to claim 1, characterized in that: The use of graph neural networks to model the communication structure between GPUs within nodes, the communication performance within nodes, and the communication performance between nodes to generate an embedding vector that characterizes the node interaction characteristics specifically includes: The cluster nodes are regarded as vertices in the graph structure, the communication performance within the nodes is regarded as the vertex attribute, and the communication performance between nodes is regarded as the edge attribute. The interaction features between nodes are extracted through graph convolution operations to generate embedding vectors.

6. A multi-objective distributed deep learning container scheduling method according to claim 1, characterized in that: The step of combining the embedding vector with the actual resource load of the node and the resource requirements of the container to be scheduled to obtain the cluster state specifically includes: Cluster Status ; in, Indicates the node resource utilization rate. Indicates the remaining resources of the node. Indicates the container resource requirements, Indicates the container deployment status. They are respectively the embedding vector representing the communication performance within the node and the embedding vector representing the communication performance between nodes.

7. The multi-objective distributed deep learning container scheduling method according to claim 1, characterized in that: The cluster state is input into the deep Q network, the mapping relationship between containers and nodes is defined through the action space, and the container scheduling scheme decided by the deep Q network is optimized based on the reward function, which specifically includes: Define each action as scheduling a container onto a node; Based on two optimization objectives, the reward function The scheduling scheme is comprehensively measured by linearly weighting the two optimization objectives: ; In the formula, Indicates the degree of similarity in the communication performance of the nodes deployed by the job. Indicates the similarity of the actual load of each node in the cluster; and are the weight coefficients of the two optimization objectives respectively, ; ; ; in, Indicates the minimum communication performance of the nodes where the job has been deployed. Indicates the maximum communication performance of the nodes deployed by the job; Indicates the minimum load of the cluster node. Indicates the maximum load of the cluster node.

8. The multi-objective distributed deep learning container scheduling method according to claim 1, characterized in that: It also includes the continuous optimization process of the deep Q network: the scheduling effect corresponding to the container scheduling plan generated for the real-time cluster status is fed back to the deep Q network model, and the container scheduling plan is continuously optimized through online incremental training.

9. A multi-objective distributed deep learning container scheduling system, characterized in that: include: Collection module: collects the communication structure between GPUs in each node in the distributed cluster and the communication performance within the node through GPU topology perception technology, obtains the communication performance between nodes through network perception technology, and monitors the actual resource load of the node through load perception technology; Encoding module: Use graph neural networks to model the communication structure between GPUs within nodes, the communication performance within nodes, and the communication performance between nodes, generate an embedding vector that characterizes the node interaction characteristics, and then splice the embedding vector with the actual resource load of the node and the resource requirements of the container to be scheduled to obtain the cluster status; Offline training module: Input the cluster state into the deep Q network, define the mapping relationship between containers and nodes through the action space, and optimize the container scheduling scheme decided by the deep Q network based on the reward function to minimize the communication overhead and balance the node load, so as to realize the training of the deep Q network; Inference deployment module: deploys the offline trained deep Q network model to the scheduler and generates a container scheduling plan based on the real-time cluster status.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Computing power routing method and system based on deep reinforcement learning and graph neural network

    CN117896306A

  • Multi-target container scheduling method and system for distributed learning operation of cloud platform

    CN118796470A

  • Graph neural network and reinforcement learning techniques for connection management

    US20220124543A1

  • Data path circuit design using reinforcement learning

    US20230139623A1

  • Computing power network system

    US20240064175A1

Cited By

  • SDN controller load balancing method based on graph neural network and evolutionary algorithm

    CN120434126A

  • A load balancing method for SDN controller based on graph neural network and evolutionary algorithm

    CN120434126B