Multi-objective Distributed Deep Learning Container Scheduling Method, System and Storage Medium

Through GPU topology awareness, network awareness and load awareness technology combined with graph neural network and deep Q network, the shortcomings of container orchestrator in communication and load awareness are solved, efficient scheduling of distributed deep learning jobs is achieved, communication bottlenecks are alleviated and load balancing is improved.

CN120029720BActive Publication Date: 2025-07-11ARTIFICIAL INTELLIGENCE RES INST OF HEFEI COMPREHENSIVE NAT SCI CENT (ANHUI ARTIFICIAL INTELLIGENCE LAB) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510499669.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-11
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing container orchestrator cannot perceive the communication performance differences between nodes in a fine-grained manner, resulting in communication bottlenecks and node overloads. The traditional scheduling methods lack global optimization capabilities, making it difficult to meet the efficient scheduling needs of distributed deep learning jobs.

Method used

GPU topology awareness, network awareness and load awareness technology are adopted, combined with graph neural networks and deep Q networks, embedded vectors representing node interaction characteristics are generated, and container scheduling strategies are optimized through reward functions to maximize communication efficiency and load balancing.

Benefits of technology

It significantly alleviates the communication bottleneck, improves the load balancing and overall performance of the cluster, meets the real-time scheduling needs, and improves the scheduling efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029720B_ABST
    Figure CN120029720B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of distributed deep learning, and discloses a multi-objective distributed deep learning container scheduling method, system and storage medium. The method includes: using a graph neural network to model the communication structure between GPUs inside a node, the internal communication performance of the node, and the communication performance between nodes, generating an embedding vector representing the interaction characteristics of the nodes, and obtaining the cluster state; inputting the cluster state into a deep Q network, defining the mapping relationship from the container to the node through the action space, optimizing the decision of the deep Q network based on the reward function to minimize the communication overhead and balance the node load, and realizing the training of the deep Q network; deploying the deep Q network model completed by offline training to the scheduler, and generating a container scheduling scheme according to the real-time cluster state. The scheduling method proposed by the present invention effectively alleviates the communication bottleneck in the training of distributed deep learning jobs while ensuring the scheduling real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of distributed deep learning technology, and in particular to a multi-objective distributed deep learning container scheduling method, system and storage medium. Background Art

[0002] With the rapid development of artificial intelligence, the parameter scale of deep learning models has continued to expand, from the initial million level to the current trillion level. Distributed deep learning has become an inevitable choice for training complex models. The traditional single-machine training method is limited by the computing power of a single device and cannot meet the training requirements of these complex models. At the same time, cloud computing, with its flexible and efficient resource management capabilities and good scalability, provides an ideal deployment platform for distributed deep learning jobs. However, mainstream container orchestrators have obvious shortcomings when facing distributed deep learning jobs. First, the current container orchestrators cannot perceive the communication performance differences between nodes in a fine-grained manner, which easily leads to communication bottlenecks and "slow node effects", thereby reducing the overall training efficiency. Secondly, they lack accurate perception of the actual load of the nodes, which easily leads to overloading of some nodes, thereby increasing the risk of system failure and affecting the overall utilization of resources. In addition, traditional container scheduling methods mostly use heuristic rules, lack the ability to coordinate the global cluster resources and multi-objective optimization, and easily lead to scheduling decisions falling into local optimality. The traditional multi-objective optimization algorithm takes a long time to solve, is not suitable for real-time scheduling scenarios, and is difficult to meet the needs of deadline (DDL) jobs.

[0003] Therefore, how to achieve efficient multi-objective optimization scheduling in a cloud environment to maximize communication efficiency and improve cluster load balancing has become a key issue that needs to be urgently solved in the process of deploying distributed deep learning jobs on the cloud. Summary of the invention

[0004] In order to solve the above technical problems, the present invention provides a multi-objective distributed deep learning container scheduling method, system and storage medium; the present invention aims at the container scheduling problem of distributed learning jobs on cloud platforms, and designs a real-time scheduling method to alleviate the communication bottleneck in the job training process and improve the overall load balancing of the system.

[0005] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0006] A multi-objective distributed deep learning container scheduling method, comprising:

[0007] The communication structure between GPUs in each node in the distributed cluster and the communication performance within the node are collected through GPU topology perception technology, the communication performance between nodes is obtained through network perception technology, and the actual resource load of the node is monitored through load perception technology.

[0008] Model the communication structure between GPUs inside a node, the internal communication performance of the node, and the communication performance between nodes using a graph neural network, generate an embedding vector representing the node interaction characteristics, and concatenate the embedding vector with the actual resource load of the node and the resource requirements of the containers to be scheduled to obtain the cluster state;

[0009] Input the cluster state into a deep Q - network, define the mapping relationship from containers to nodes through the action space, optimize the container scheduling scheme of the deep Q - network decision based on the reward function to minimize the communication overhead and balance the node load, and achieve the training of the deep Q - network;

[0010] Deploy the deep Q - network model completed by offline training to the scheduler, and generate a container scheduling scheme according to the real - time cluster state.

[0011] In one embodiment, the method for collecting the communication structure between GPUs inside each node and the internal communication performance of the node in the distributed cluster through GPU topology - aware technology specifically includes:

[0012] Quantify the internal communication performance weight according to the connection mode between GPUs inside the node, and measure the internal communication performance of the node based on the communication structure between GPUs.

[0013] In one embodiment, the method for obtaining the communication performance between nodes through network - aware technology specifically includes:

[0014] Measure the communication performance between nodes by calculating the following formula :

[0015] ;

[0016] In the formula, and respectively represent the bandwidth weight and the latency weight, , , respectively represent the bandwidth value, the latency value, and the packet loss rate.

[0017] In one embodiment, the method for monitoring the actual resource load of a node through load - aware technology specifically includes:

[0018] ;

[0019] is the actual resource load of the node, where , , are the weight coefficients of the CPU resource utilization rate , the weight coefficient of the memory resource utilization rate and the GPU resource utilization rate respectively weight coefficients, and .

[0020] In one embodiment, the method of using a graph neural network to model the communication structure between GPUs inside a node, the internal communication performance of the node, and the communication performance between nodes to generate an embedding vector representing the node interaction characteristics specifically includes:

[0021] Regarding the cluster nodes as vertices in the graph structure, the internal communication performance of the nodes as vertex attributes, and the communication performance between nodes as edge attributes, extracting the interaction features between nodes through graph convolution operations to generate an embedding vector.

[0022] In one embodiment, the method of splicing the embedding vector with the actual resource load of the node and the resource requirements of the containers to be scheduled to obtain the cluster state specifically includes:

[0023] Cluster state ;

[0024] Wherein, represents the utilization rate of the node resource volume, represents the remaining resource volume of the node, represents the resource requirements of the container, represents the deployment state of the container, are respectively the embedding vector representing the internal communication performance of the node and the embedding vector representing the communication performance between nodes.

[0025] In one embodiment, the method of inputting the cluster state into a deep Q-network, defining the mapping relationship from the container to the node through the action space, and optimizing the container scheduling scheme of the deep Q-network decision based on the reward function specifically includes:

[0026] Define each action as scheduling a container to a node;

[0027] Based on two optimization objectives, the reward function By linearly weighting the two optimization objectives to comprehensively measure the scheduling scheme:

[0028] ;

[0029] In the formula, represents the degree of similarity of the communication performance of the nodes where the jobs have been deployed, represents the degree of similarity of the actual loads of each node in the cluster; and are respectively the weight coefficients of the two optimization objectives, ;

[0030] ;

[0031] ;

[0032] Wherein, represents the minimum communication performance of the nodes where the job has been deployed, represents the maximum communication performance of the nodes where the job has been deployed; represents the minimum load of the cluster nodes, represents the maximum load of the cluster nodes.

[0033] In one embodiment, it further includes the continuous optimization process of the deep Q network: the scheduling effect corresponding to the container scheduling scheme generated for the real-time cluster state is fed back to the deep Q network model, and the container scheduling scheme is continuously optimized through online incremental training.

[0034] A multi-objective distributed deep learning container scheduling system, comprising:

[0035] An acquisition module: acquiring the communication structure between GPUs inside each node in the distributed cluster and the internal communication performance of the node through GPU topology awareness technology, obtaining the inter-node communication performance through network awareness technology, and monitoring the real resource load of the node through load awareness technology;

[0036] An encoding module: using a graph neural network to model the communication structure between GPUs inside the node, the internal communication performance of the node, and the inter-node communication performance, generating an embedding vector representing the node interaction characteristics, and splicing the embedding vector with the real resource load of the node and the resource requirements of the container to be scheduled to obtain the cluster state;

[0037] An offline training module: inputting the cluster state into the deep Q network, defining the mapping relationship from the container to the node through the action space, and optimizing the container scheduling scheme determined by the deep Q network based on the reward function to minimize the communication overhead and balance the node load, thereby realizing the training of the deep Q network;

[0038] An inference and deployment module: deploying the deep Q network model completed in offline training to the scheduler, and generating a container scheduling scheme according to the real-time cluster state.

[0039] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method in any one of the embodiments are implemented.

[0040] Compared with the prior art, the beneficial technical effects of the present invention are:

[0041] The present invention makes up for the deficiencies of existing container orchestrators in terms of communication and load awareness. By using a graph neural network to model the communication topology inside and outside the cluster and combining load awareness technology to monitor the real-time utilization of node resources, the scheduler can accurately perceive the communication performance and node load status in the cluster, thereby optimizing the scheduling strategy. Further, the present invention uses a reinforcement learning algorithm to achieve the overall planning and multi-objective optimization of global resources. Through the combination of offline training and online inference, the trained model can quickly generate an optimal scheduling strategy in the actual environment, significantly reducing the solution time. In summary, the scheduling method proposed by the present invention not only ensures the real-time performance of scheduling, but also effectively alleviates the communication bottleneck in the training of distributed deep learning jobs and achieves load balancing among nodes, thereby significantly improving the overall performance of the cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a flowchart of the method in an embodiment of the present invention.

[0043] Figure 2 It is a schematic diagram of the framework of the scheduling method in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The following describes in detail a preferred embodiment of the present invention with reference to the accompanying drawings.

[0045] Aiming at the deficiencies of mainstream container orchestrators in communication awareness and load awareness, the present invention uses GPU topology awareness, network awareness, and load awareness technologies to perform fine-grained awareness and modeling of the node status in a distributed cluster. Specifically, the GPU topology awareness technology is used to model the communication structure between GPUs within a node, clarify the connection method between GPUs within each node, so as to improve the awareness accuracy of internal communication performance; the network awareness technology is used to capture the communication characteristics between nodes in the cluster, including performance indicators such as communication bandwidth, latency, and packet loss rate, to improve the awareness ability of the scheduler for the communication performance between nodes; the load awareness technology monitors the resource utilization of each node in the cluster in real time through monitoring components such as Metrics Server to ensure that the scheduler can perceive the real load of the node during task scheduling. Secondly, to solve the problems of the lack of globality in traditional container scheduling methods and the long solution time of traditional multi-objective optimization algorithms, the present invention proposes a multi-objective container scheduling method based on deep reinforcement learning (DRL). First, use graph neural networks (GNN) to model the communication topology in the cluster, extract the complex relationships between nodes and encode them into state vectors, so as to provide an accurate representation of communication performance for the scheduling strategy; then, learn the optimal container scheduling strategy through the deep Q network (DQN) algorithm to achieve the goal of reducing communication overhead and improving the load balance degree of the cluster. This method can not only coordinate cluster resources from a global perspective, but also quickly generate scheduling decisions in a dynamic environment, significantly improving scheduling efficiency and system performance.

[0046] The present invention mainly collects the communication topology, node load, and container information between cluster nodes, and uses graph neural networks for cluster state encoding. Then, the encoded state information is used as input, and a feasible scheduling scheme is generated through decision-making by a reinforcement learning agent to achieve efficient scheduling of containers.

[0047] As Figure 1 shown, a multi-objective distributed deep learning container scheduling method in the present invention includes the following steps:

[0048] S1: Collect the communication structure between GPUs within each node in the distributed cluster and the internal communication performance of the node through the GPU topology awareness technology, obtain the communication performance between nodes through the network awareness technology, and monitor the real resource load of the node through the load awareness technology.

[0049] S2: Use a graph neural network to model the communication structure between GPUs inside a node, the internal communication performance of the node, and the communication performance between nodes, generate an embedding vector representing the interaction characteristics of the node, and splice the embedding vector with the real resource load of the node and the resource requirements of the containers to be scheduled to obtain the cluster state.

[0050] S3: Input the cluster state into a deep Q-network, define the mapping relationship from containers to nodes through the action space, and optimize the container scheduling scheme decided by the deep Q-network based on the reward function to minimize the communication overhead and balance the node load, thus realizing the training of the deep Q-network.

[0051] S4: Deploy the deep Q-network model completed with offline training to the scheduler, and generate a container scheduling scheme according to the real-time cluster state.

[0052] In the technical solution of the present invention, a node is an independent computing unit in a distributed cluster, representing a physical server or virtual machine in the cluster, and having independent computing resources (CPU, GPU, memory) and network interfaces.

[0053] In one embodiment, collecting the communication structure between GPUs inside each node in the distributed cluster and the internal communication performance of the node through the GPU topology awareness technology in step S1 specifically includes:

[0054] Quantify the internal communication performance weight according to whether the connection method between GPUs inside the node is NVLink or PCIe, and measure the internal communication performance of the node based on the communication structure between GPUs.

[0055] In one embodiment, obtaining the communication performance between nodes through the network awareness technology specifically includes:

[0056] Calculate and measure the communication performance between nodes through the following formula :

[0057] ;

[0058] In the formula, and respectively represent the bandwidth weight and the latency weight, , , respectively represent the bandwidth value, the latency value, and the packet loss rate.

[0059] In one embodiment, monitoring the real resource load of the node through the load awareness technology specifically includes:

[0060] ;

[0061] is the true resource load of the node, where , , are the CPU resource utilization rate , the memory resource utilization rate and the weight coefficients of the GPU resource utilization rate , and .

[0062] Traditional container orchestrators lack the awareness of the communication performance between nodes and the actual load of nodes in the cluster, resulting in the inability to fully optimize communication efficiency and load balancing during scheduling decisions. Therefore, the present invention first collects data on nodes in the cluster through various technical means, and collects these useful metrics to lay a foundation for subsequent algorithm design. Specifically, a GPU device plugin is used to collect the topology information between GPUs inside each node, a network performance testing tool is used to collect the network performance metrics between cluster nodes, and the actual load conditions of each node (such as the utilization rates of CPU, GPU, and memory) are monitored in real time through monitoring components such as Kubernetes' Metrics Server. These metrics provide comprehensive data support for cluster state encoding and subsequent scheduling scheme design.

[0063] Specifically, as Figure 2 shown, an improved GPU device plugin can be deployed on the cluster nodes, and the device plugin is used to collect the topology structure and communication performance between GPUs inside each node in the cluster, such as the connection method (NVLink or PCIe) and the bandwidth situation, and use this to measure the communication performance inside the node; then, a network measurement tool is used to regularly measure the network performance metrics between nodes, including communication bandwidth, latency, and packet loss rate, etc., and use this to measure the communication performance between nodes. In addition, the actual resource load conditions of each node, including the actual utilization rates of CPU, memory, and GPU resources, are monitored in real time through Kubernetes Metrics Server. After collecting these performance metrics, they need to be preprocessed and quantified according to a certain scoring method. For example, weights can be assigned to the communication performance according to different topology connection methods between GPUs; the communication cost between nodes can be scored, and the smaller the value of the communication cost, the lower the communication cost.

[0064] In one embodiment, in step S2, using a graph neural network to model the communication structure between GPUs inside the node, the internal communication performance of the node, and the communication performance between nodes, and generating an embedding vector representing the node interaction characteristics specifically includes:

[0065] Regarding the cluster nodes as vertices in the graph structure, the internal communication performance of the node as vertex attributes, and the communication performance between nodes as edge attributes, and extracting the interaction features between nodes through graph convolution operations to generate an embedding vector.

[0066] To effectively handle variable-length and high-dimensional state data, it is necessary to encode the cluster state. The present invention uses a graph neural network to process the collected communication performance metrics and encodes the communication topologies inside and outside the cluster nodes into states. The graph neural network can directly model the irregular communication topologies in the cluster, capture the complex interaction characteristics between nodes (such as bandwidth, latency, etc.), and encode them into state vectors. This method can not only retain the global topology information but also has strong scalability and can adapt to the dynamic changes in the cluster scale and communication topology. In addition, the embedding vectors output by the graph neural network, the actual resource loads of each node, and the resource demand vectors of each container are concatenated to form the complete state input of the deep Q network. In this way, the state input not only includes the global characteristics of the communication topologies inside and outside the cluster but also combines the resource load information of each node, thus providing a comprehensive and accurate state description for the deep Q network model. This combination can better support the learning of scheduling strategies, enabling the deep reinforcement learning model to consider both communication performance and the resource load of nodes when performing task scheduling to achieve multi-objective optimization of alleviating communication bottlenecks and load balancing.

[0067] The graph neural network can effectively process irregular and variable-length cluster structures, extract the interaction characteristics in complex networks, and generate embedding vectors. In addition, the actual load information of nodes and the information of containers to be scheduled also need to be part of the state. Finally, the embedding vectors generated by the graph neural network are concatenated with the resource load and container information vectors to form the complete state input for the decision-making process of the subsequent reinforcement learning model.

[0068] In one embodiment, the concatenating the embedding vector with the actual resource load of the node and the resource demand of the container to be scheduled to obtain the cluster state specifically includes:

[0069] Cluster state ;

[0070] Wherein, represents the utilization rate of node resource volume, represents the remaining resource volume of the node, represents the resource demand of the container, represents the deployment state of the container, are the embedding vectors representing the internal communication performance of the node and the embedding vectors representing the communication performance between nodes, respectively.

[0071] Specifically, the cluster state It is necessary to comprehensively express the information of the current cluster and the tasks to be scheduled so that the agent can judge the system state. To ensure that the agent fully considers the communication cost during the decision-making process, the present invention not only includes the utilization rate of resources on each node in the design of the state and the remaining resource amount , but also incorporates key communication topology information, encodes this information through a graph neural network to generate a structured embedding vector , which respectively represents the communication performance inside and outside the node, and takes it as part of the state of the deep reinforcement learning algorithm. In addition, the resource requirements of the containers , relevant job information, and their current deployment status are integrated into the cluster state.

[0072] In one embodiment, inputting the cluster state into the deep Q network in step S3, defining the mapping relationship from the container to the node through the action space, and optimizing the container scheduling scheme of the deep Q network decision based on the reward function specifically includes:

[0073] Define each action as scheduling a container to a node;

[0074] Based on two optimization objectives, the reward function linearly weights the two optimization objectives to comprehensively measure the scheduling scheme:

[0075] ;

[0076] In the formula, represents the degree of similarity of the communication performance of the nodes where the jobs have been deployed, represents the degree of similarity of the actual loads of each node in the cluster; and are respectively the weight coefficients of the two optimization objectives, ;

[0077] ;

[0078] ;

[0079] Among them, represents the minimum communication performance of the nodes where the jobs have been deployed, represents the maximum communication performance of the nodes where the jobs have been deployed; represents the minimum load of the cluster nodes, represents the maximum load of the cluster nodes.

[0080] "Nodes where the jobs have been deployed" refers to the cluster nodes that have been assigned and are running the current job containers during the scheduling process.

[0081] In the learning stage of the scheduling strategy, the present invention uses a deep Q-network model to design an agent. In the design of the model actions, each action is regarded as scheduling a container to a node, and the action set represents a scheduling scheme for a group of containers. Then, when the number of containers to be scheduled is m and the total number of nodes is N, the action can be expressed as:

[0082] ;

[0083] .

[0084] In addition, in the design of the model reward, since the multi-container scheduling problem belongs to a long-sequence decision-making problem and there is a significant correlation in the scheduling between containers, it is difficult to accurately evaluate the overall quality of the agent's decision before all containers are completely scheduled. Therefore, we adopt a method of reward sharing, and use the final reward as the shared reward for all steps in an epoch. Based on two optimization objectives, we linearly weight multiple objective function values to comprehensively measure the pros and cons of the scheduling scheme.

[0085] In the process of generating the scheduling strategy, the key step of the present invention is to use the encoded cluster state as the input, and then perform inference and decision-making through a reinforcement learning agent. The state vector includes the communication topology characteristics of the nodes in the cluster, the resource load conditions, and the container resource requirements. These information are encoded and embedded by a graph neural network and concatenated with the resource vector to form a complete input representation. Based on these inputs, the agent continuously learns the optimal scheduling strategy through the deep Q-network model in offline training to achieve the goal of minimizing communication overhead and ensuring node load balance. In the online inference stage, the agent uses these state vectors in real time to generate the optimal scheduling scheme for containers. When a new container needs to be scheduled, the agent will select the most suitable node to deploy the container according to the overall state of the current cluster to ensure the communication efficiency and load balance of the system. Through such a design, the present invention can perform intelligent scheduling efficiently in a dynamic environment and realize the real-time optimization of the scheduling scheme.

[0086] In one embodiment, it further includes a continuous optimization process of the deep Q-network: the scheduling effect corresponding to the container scheduling scheme generated for the real-time cluster state is fed back to the deep Q-network model, and the container scheduling scheme is continuously optimized through online incremental training.

[0087] Once the agent completes the reasoning and decision-making, the generated container scheduling plan will be actually executed by the scheduler. Based on the scheduling decisions output by the agent, the scheduler assigns the containers to be scheduled to appropriate nodes for running to optimize the overall performance of the cluster. The scheduling plan output by the reasoning is a set of mappings of containers to nodes. For example, {0:3, 1:2, 2:1} means that container 0 should be scheduled to node 3, container 1 should be scheduled to node 2, and container 2 should be scheduled to node 1. During the execution process, the system will provide real-time feedback on the scheduling effect and use this feedback information for continuous optimization of the reinforcement learning model, thereby continuously improving the scheduling strategy to ensure the efficiency and stability of the scheduling.

[0088] Offline training can fully learn the scheduling strategy in a complex simulation environment, thereby providing efficient decision support for online reasoning. Compared with traditional multi-objective optimization algorithms, the combination of offline training and online reasoning enables the scheduler to quickly infer the optimal container scheduling strategy according to the real-time cluster state, greatly improving the scheduling efficiency and ensuring the stable operation of the system in a dynamic environment.

[0089] It should be understood that although the steps in the flowchart of the accompanying drawings of the specification are shown sequentially according to the indication of the arrows, these steps do not necessarily need to be executed sequentially according to the order indicated by the arrows. Unless there is a clear description in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings of the specification may include multiple steps or multiple stages. These steps or stages do not necessarily need to be executed at the same moment, but can be executed at different moments. The execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0090] The present invention also provides a multi-objective distributed deep learning container scheduling system. Since the implementation solutions and methods for the system to solve problems are similar, the implementation of the specific system in the embodiments of this specification can refer to the implementation of the foregoing method, and the repeated parts will not be elaborated. As used hereinafter, the term "module" or "modular unit" may be a combination of software and / or hardware that can achieve a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0091] A multi-objective distributed deep learning container scheduling system includes:

[0092] Collection module: It collects the communication structure between GPUs within each node in the distributed cluster and the internal communication performance of the node through GPU topology awareness technology, obtains the inter-node communication performance through network awareness technology, and monitors the real resource load of the node through load awareness technology;

[0093] Encoding module: It uses a graph neural network to model the communication structure between GPUs within the node, the internal communication performance of the node, and the inter-node communication performance, generates an embedding vector representing the interaction characteristics of the node, and splices the embedding vector with the real resource load of the node and the resource requirements of the container to be scheduled to obtain the cluster state;

[0094] Offline training module: It inputs the cluster state into the deep Q network, defines the mapping relationship from the container to the node through the action space, optimizes the container scheduling scheme decided by the deep Q network based on the reward function to minimize the communication overhead and balance the node load, and realizes the training of the deep Q network;

[0095] Inference and deployment module: It deploys the deep Q network model completed by offline training to the scheduler and generates a container scheduling scheme according to the real-time cluster state.

[0096] The present invention also provides a computer-readable storage medium including instructions, such as a memory including instructions. The above instructions can be executed by a processor to complete the above method. The storage medium can be a computer-readable storage medium. For example, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0097] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to include all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention, and any reference signs in the claims should not be regarded as limiting the claims involved.

[0098] In addition, it should be understood that although this specification is described according to the embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A multi-objective distributed deep learning container scheduling method, characterized in that Including: Collect the communication structure between GPUs inside each node in the distributed cluster and the internal communication performance of the node through GPU topology awareness technology, obtain the inter-node communication performance through network awareness technology, and monitor the real resource load of the node through load awareness technology; Use a graph neural network to model the communication structure between GPUs inside the node, the internal communication performance of the node, and the inter-node communication performance, generate an embedding vector representing the node interaction characteristics, and splice the embedding vector with the real resource load of the node and the resource requirements of the container to be scheduled to obtain the cluster state; Cluster status ; Indicates the utilization rate of node resource volume, Indicates the remaining resource volume of the node, Indicates the container resource requirement, Indicates the container deployment status, Are respectively the embedding vector representing the internal communication performance of the node and the embedding vector representing the communication performance between nodes; Input the cluster state into the deep Q network, define the mapping relationship from the container to the node through the action space, and optimize the container scheduling scheme decided by the deep Q network based on the reward function to minimize the communication overhead and balance the node load, and realize the training of the deep Q network; Deploy the deep Q network model completed by offline training to the scheduler, and generate a container scheduling scheme according to the real-time cluster state.

2. The multi-objective distributed deep learning container scheduling method according to claim 1, wherein The collection of the communication structure between GPUs inside each node in the distributed cluster and the internal communication performance of the node through GPU topology awareness technology specifically includes: Quantify the internal communication performance weight according to the connection method between GPUs inside the node, and measure the internal communication performance of the node based on the communication structure between GPUs.

3. The multi-objective distributed deep learning container scheduling method according to claim 1, characterized in that The obtaining of the inter-node communication performance through network awareness technology specifically includes: The communication performance between nodes is measured by the following formula :[[]]END]] ; Wherein, and respectively represent the bandwidth weight and the delay weight, , , respectively represent the bandwidth value, the delay value and the packet loss rate.

4. A multi-objective distributed deep learning container scheduling method according to claim 1, characterized in that The monitoring of the real resource load of the node through load awareness technology specifically includes: ; is the true resource load of the node, where , , are the weight coefficients of CPU resource utilization , the weight coefficient of memory resource utilization , and the weight coefficient of GPU resource utilization , respectively, and .

5. A multi-objective distributed deep learning container scheduling method according to claim 1, characterized in that The use of a graph neural network to model the communication structure between GPUs inside the node, the internal communication performance of the node, and the inter-node communication performance, and generate an embedding vector representing the node interaction characteristics specifically includes: Take the cluster nodes as vertices in the graph structure, the internal communication performance of the node as vertex attributes, the inter-node communication performance as edge attributes, extract the inter-node interaction features through graph convolution operations, and generate an embedding vector.

6. The multi-objective distributed deep learning container scheduling method according to claim 1, characterized in that The input of the cluster state into the deep Q network, the definition of the mapping relationship from the container to the node through the action space, and the optimization of the container scheduling scheme decided by the deep Q network based on the reward function specifically include: Define each action as scheduling a container to a node; Based on two optimization objectives, the reward function By linearly weighting the two optimization objectives to comprehensively measure the scheduling scheme: ; In the formula, represents the similarity degree of the communication performance of the nodes where the jobs are deployed, represents the similarity degree of the actual loads of the nodes in the cluster; and are the weight coefficients of the two optimization objectives respectively, ; ; ; Among them, represents the minimum communication performance of the nodes where the job has been deployed, represents the maximum communication performance of the nodes where the job has been deployed; represents the minimum load of the cluster nodes, represents the maximum load of the cluster nodes.

7. A multi-objective distributed deep learning container scheduling method according to claim 1, characterized in that, It also includes the continuous optimization process of the deep Q network: feedback the scheduling effect corresponding to the container scheduling scheme generated for the real-time cluster state to the deep Q network model, and continuously optimize the container scheduling scheme through online incremental training.

8. A multi-objective distributed deep learning container scheduling system, characterized in that Including: Collection module: Collect the communication structure between GPUs inside each node in the distributed cluster and the internal communication performance of the node through GPU topology awareness technology, obtain the inter-node communication performance through network awareness technology, and monitor the real resource load of the node through load awareness technology; Encoding module: Use a graph neural network to model the communication structure between GPUs inside the node, the internal communication performance of the node, and the inter-node communication performance, generate an embedding vector representing the node interaction characteristics, and splice the embedding vector with the real resource load of the node and the resource requirements of the container to be scheduled to obtain the cluster state; Cluster status ; Indicates the utilization rate of node resource volume, Indicates the remaining resource volume of the node, Indicates the container resource requirement, Indicates the container deployment status, They are the embedding vector representing the internal communication performance of the node and the embedding vector representing the communication performance between nodes respectively; Offline training module: Input the cluster status into the deep Q-network, define the mapping relationship from containers to nodes through the action space, optimize the container scheduling scheme of the deep Q-network decision based on the reward function to minimize the communication overhead and balance the node load, and implement the training of the deep Q-network; Inference and deployment module: Deploy the deep Q-network model completed by offline training to the scheduler, and generate a container scheduling scheme according to the real-time cluster status.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-target container scheduling method and system for distributed learning operation of cloud platform

    CN118796470A

  • Computing power network system

    US20240064175A1