Rescheduling Method, Device, System, Equipment, Medium and Product for Training Tasks
By monitoring the cluster state in large model training and rescheduling to topology with lower communication costs, the problem of degradation in cluster communication efficiency and stability is solved, and more efficient and reliable training task execution is achieved.
Patent Information
- Application Number
- CN202510245733.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-04
AI Technical Summary
During the training process of large-scale model, as the training work continues, the communication efficiency of the cluster decreases, resulting in a decrease in training efficiency. In addition, the existing technology can easily cause errors in the entire cluster when a failure occurs, affecting stability and reliability.
By monitoring the cluster status, eliminating the exception handling unit in case of failures and rebuilding new units, combined with network topology awareness strategy, rescheduling training tasks to topology with lower communication costs at low usage, adapting to dynamic changes, and improving communication efficiency and stability.
It improves the communication efficiency and stability of large-scale model training, optimizes resource utilization, avoids the impact of single point failure on training, and ensures the reliable execution of training tasks.
Smart Images

Figure CN119759545B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence (AI), and more specifically, to a method, device, system, equipment, medium and product for rescheduling training tasks. Background Art
[0002] A large model refers to an artificial intelligence model with a large number of parameters, such as billions or even hundreds of billions of parameters. Large models can handle complex tasks and generate high-quality results, and are widely used in fields such as natural language processing and computer vision. The parallel strategies for large model training have evolved from early data parallelism to the widely used 3D parallelism (a combination of data parallelism, model parallelism, and pipeline parallelism), forming a multi-dimensional parallel training framework.
[0003] As the model scale continues to increase, the communication bandwidth requirements for the clusters used to train large models are also getting higher and higher. In current large model training, training tasks are assigned to specific topologies in the cluster according to preset rules before training starts. However, as the training work continues, the communication efficiency of the topologies executing the training tasks may decline, thus affecting the training efficiency. In addition, when a failure occurs (such as GPU failure or node failure), error handling is directly executed, resulting in the entire cluster being affected, posing challenges to the stability and reliability of training. Summary of the Invention
[0004] The present invention provides a method, device, system, equipment, medium and product for rescheduling training tasks, which helps to improve training efficiency.
[0005] The technical solution of the embodiment of the present invention is as follows:
[0006] A method for rescheduling training tasks, comprising:
[0007] When it is determined that the rescheduling time point is reached, determining a first metric of the current topology to which the training task has been scheduled, the current topology including a plurality of processing units, and the first metric characterizing the communication cost of the current topology;
[0008] Determining a rescheduling topology formed after at least one processing unit in the current topology is virtually migrated;
[0009] Determining a second metric of the rescheduling topology, the second metric characterizing the communication cost of the rescheduling topology;
[0010] When the second metric is less than the first metric, rescheduling the training task to the rescheduling topology.
[0011] In one embodiment, the determination of reaching the rescheduling time includes at least one of the following:
[0012] Determine that the utilization rate of the cluster for executing the training task is lower than a predetermined threshold or the overall topology of the cluster has changed;
[0013] Determine that a user-triggered rescheduling instruction is received;
[0014] Determine that a predetermined rescheduling time point is reached.
[0015] In one embodiment, the determination of the rescheduling topology formed after at least one processing unit in the current topology is virtually migrated includes:
[0016] Virtually migrate the processing unit in the first node from its current position in the first node to another position different from the current position in the first node; or
[0017] Virtually migrate the processing unit in the first node from its current position in the first node to the second node.
[0018] In one embodiment, the virtual migration of the processing unit in the first node from its current position in the first node to the second node includes:
[0019] Virtually migrate the processing unit in the first node from its current position in the first node to the second node connected to the same switch node as the first node; or
[0020] Virtually migrate the processing unit in the first node from its current position in the first node to the second node connected to a different switch node from the first node.
[0021] In one embodiment, the determination of the first metric of the current topology to which the training task has been scheduled includes: determining the individual communication cost of each processing unit in the current topology with the remaining processing units in the current topology; calculating a first summation result of the individual communication costs of all processing units in the current topology; and determining the first metric based on the first summation result;
[0022] The determination of the second metric of the rescheduling topology includes: determining the individual communication cost of each processing unit in the rescheduling topology with the remaining processing units in the rescheduling topology; calculating a second summation result of the individual communication costs of all processing units in the rescheduling topology; and determining the second metric based on the second summation result.
[0023] In one embodiment, the rescheduling of the training task to the rescheduling topology includes:
[0024] Evict the at least one processing unit;
[0025] Reconstruct the at least one processing unit at the location in the rescheduling topology to which the at least one processing unit is virtually migrated.
[0026] In one embodiment, it includes:
[0027] When a processing unit in an abnormal state is detected in the current topology, evict the processing unit in the abnormal state;
[0028] Reconstruct the processing unit at the current location of the current topology where the processing unit in the abnormal state is located.
[0029] In one embodiment, it includes:
[0030] When a processing unit in an abnormal state is detected in the current topology, determine the abnormal processing topology formed after the virtual migration of the processing unit in the abnormal state;
[0031] Determine a third metric of the abnormal processing topology, where the third metric characterizes the communication cost of the abnormal processing topology;
[0032] When the third metric is less than the first metric, evict the processing unit in the abnormal state and reconstruct the processing unit at the location in the abnormal processing topology to which the processing unit in the abnormal state is virtually migrated;
[0033] When the third metric is greater than or equal to the first metric, evict the processing unit in the abnormal state and reconstruct the processing unit at the current location of the current topology where the processing unit in the abnormal state is located.
[0034] A rescheduling device for a training task, including:
[0035] A first determination module, configured to determine a first metric of the current topology to which the training task has been scheduled when determining that the rescheduling time has arrived, where the current topology includes multiple processing units, and the first metric characterizes the communication cost of the current topology;
[0036] A second determination module, configured to determine a rescheduling topology formed after virtual migration of at least one processing unit in the current topology;
[0037] A third determination module, configured to determine a second metric of the rescheduling topology, where the second metric characterizes the communication cost of the rescheduling topology;
[0038] A rescheduling module, configured to reschedule the training task to the rescheduling topology when the second metric is less than the first metric.
[0039] A rescheduling system for training tasks, comprising:
[0040] A cluster containing multiple nodes;
[0041] A scheduler for scheduling training tasks to the cluster to form a current topology, the current topology including multiple processing units in respective nodes;
[0042] A rescheduler for, when determining that the rescheduling opportunity has arrived, determining a first metric of the current topology, the first metric characterizing the communication cost of the current topology; determining a rescheduled topology formed after at least one processing unit in the current topology is virtually migrated; determining a second metric of the rescheduled topology, the second metric characterizing the communication cost of the rescheduled topology; and when the second metric is less than the first metric, rescheduling the training task to the rescheduled topology.
[0043] An electronic device, comprising:
[0044] A memory;
[0045] A processor;
[0046] Wherein an application program executable by the processor is stored in the memory, for causing the processor to execute the rescheduling method of the training task as described in any one of the above.
[0047] A computer-readable storage medium, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the processor is caused to execute the rescheduling method of the training task as described in any one of the above.
[0048] A program product, including a computer program, which when executed by a processor implements the rescheduling method of the training task as described in any one of the above.
[0049] As can be seen from the above technical solution, when determining the rescheduling time point, the first index of the current topology to which the training task has been scheduled is determined. The current topology includes multiple processing units, and the first index represents the communication cost of the current topology; the rescheduling topology formed after at least one processing unit in the current topology is virtually migrated is determined; the second index of the rescheduling topology is determined, and the second index represents the communication cost of the rescheduling topology; when the second index is less than the first index, the training task is rescheduled to the rescheduling topology. Thus, by rescheduling the training task to be executed by a topology with a lower communication cost, various dynamic changes in the cluster (such as load changes and cluster resource changes, etc.) can be adapted, thereby improving communication efficiency and training efficiency. In addition, when an anomaly is detected, the processing unit is reconstructed based on multiple mechanisms instead of simply reporting an error, thereby improving the stability and reliability of the training task. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 FIG. is a schematic flowchart of a method for rescheduling a training task according to an embodiment of the present invention.
[0051] Figure 2 FIG. is a schematic diagram of virtual migration of a virtual processing unit within a node according to an embodiment of the present invention.
[0052] Figure 3 FIG. is a schematic diagram of virtual migration of a processing unit between nodes according to an embodiment of the present invention.
[0053] Figure 4 FIG. is a schematic diagram of virtual migration of a processing unit in a cluster based on a leaf-spine structure according to an embodiment of the present invention.
[0054] Figure 5 FIG. is a schematic diagram of a rescheduling process of a training task according to an embodiment of the present invention.
[0055] Figure 6 FIG. is a schematic structural diagram of a rescheduling system of a training task according to an embodiment of the present invention.
[0056] Figure 7 FIG. is a schematic structural diagram of a rescheduling device of a training task according to an embodiment of the present invention.
[0057] Figure 8 FIG. is a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings.
[0059] For the sake of simplicity and intuitiveness in description, the solutions of the present invention will be elaborated below by describing several representative embodiments. A large number of details in the embodiments are only used to help understand the solutions of the present invention. However, it is obvious that the technical solutions of the present invention can be implemented without being limited to these details. In order to avoid unnecessarily obscuring the solutions of the present invention, some embodiments are not described in detail but only the framework is given. Hereinafter, "including" means "including but not limited to", and "according to..." means "at least according to..., but not limited to only according to...". Due to the language habits of Chinese, when the quantity of a component is not specifically indicated hereinafter, it means that the component can be one or more, or can be understood as at least one.
[0060] The parallel strategies for large model training have evolved from early data parallelism to the widely used 3D parallelism nowadays. 3D parallelism comprehensively utilizes three technologies: data parallelism, model parallelism, and pipeline parallelism, forming a multi-dimensional parallel training framework. Each parallel strategy targets specific bottleneck problems, and their combination can improve training efficiency and resource utilization. In data parallelism, gradient synchronization is mainly performed, and the communication volume is proportional to the model size. In model parallelism, communication is mainly achieved through tensor slicing, and the communication volume is relatively large. In pipeline parallelism, activation value transmission is mainly performed, and the communication volume is relatively small. The scheduler in the cluster used for training large models can adopt different bandwidth allocation strategies for different parallel dimensions according to these characteristics.
[0061] In the existing 3D parallel training of large models, as the training work continues, many dynamic changes may occur in the cluster (such as load changes and cluster resource changes, etc.), resulting in a decrease in the communication efficiency of the current topology for executing training tasks, thereby affecting the training efficiency. Moreover, in the existing 3D parallel training technology of large models, there are limitations in terms of fault tolerance. For example, taking the Spine-Leaf switch network architecture as an example for illustration, the Spine layer switch may become a single point of failure. Once a failure occurs, the entire cluster directly reports an error, challenging the stability and reliability of the training.
[0062] In an embodiment of the present invention, a rescheduling strategy based on network topology awareness is adopted to overcome or mitigate the above problems. By continuously monitoring the running state of the cluster, when a failure is detected, instead of simply reporting an error directly, abnormal running units are evicted, and new running units are created to replace the evicted ones, which can avoid the impact of single-point failures on training and improve the stability and reliability of training tasks. Moreover, the topology allocation status of training tasks can be detected when the cluster utilization rate is low, and the training tasks can be rescheduled to a new topology with lower communication costs, so that the resource allocation is more reasonable, can adapt to various changes during the training process, improve communication efficiency and training efficiency, and optimize resource utilization.
[0063] The above disclosure details the technical defects existing in the related art, the reasons for these technical defects, and the thought analysis process for overcoming these technical defects. In fact, the recognition of the above technical defects is not common knowledge in the art, but a novel discovery by the inventor in the research. In addition, the reason tracing of the technical defects and the thought analysis process for overcoming these technical defects are also the step-by-step analysis results of the inventor in the actual research process, and none of them are common knowledge in the art.
[0064] Figure 1 FIG. is a schematic flowchart of a rescheduling method for a training task according to an embodiment of the present invention. This method can be executed by a controller in a cluster that executes a large model training task. As Figure 1 shown, the method includes:
[0065] Step 101: When it is determined that the rescheduling opportunity has arrived, determine a first metric of the current topology to which the training task has been scheduled. The current topology includes multiple processing units, and the first metric characterizes the communication cost of the current topology.
[0066] In one embodiment, determining that the rescheduling opportunity has arrived includes at least one of the following:
[0067] (1) Determine that the utilization rate of the cluster used to execute the training task is lower than a predetermined threshold or the overall topology structure of the cluster has changed.
[0068] For example, the utilization rate of the cluster can be measured based on indicators such as the GPU computing power utilization rate of the cluster, the GPU video memory utilization rate of the cluster, and the cluster throughput. Moreover, the change in the overall topology structure of the cluster can include: the addition or deletion of nodes in the cluster, the change in the hardware performance of nodes, or the change in the switch network structure used to connect nodes, and so on.
[0069] (2) Determine that a rescheduling instruction triggered by the user has been received.
[0070] For example, a rescheduling instruction sent by the user through a human-computer interaction interface is received.
[0071] (3) Determine the arrival at the predetermined rescheduling time point.
[0072] For example, set periodic rescheduling time points. When each periodic time point is reached, the arrival at the rescheduling opportunity is automatically determined.
[0073] It can be seen that by executing the Figure 1 process during the low-usage period or periodic time points of the cluster, the complexity and overhead of rescheduling can be reduced, and an efficient and stable training process can be achieved.
[0074] The scheduler in the cluster can, based on various predetermined scheduling principles, schedule the training tasks to the processing units in the cluster (the processing units are usually located in nodes that are physical servers), thereby forming the current topology. For example, the scheduling principles can include: resource matching principle, affinity and anti-affinity principle, priority and fairness principle, or load balancing and communication optimization principle, and so on.
[0075] The current topology contains the specific processing units in the cluster for executing the training task and the specific deployment locations of these specific processing units in the cluster (for example, which specific node the processing unit is located in). These specific processing units in the current topology execute the training task.
[0076] The nodes in the cluster, as physical servers, support the training tasks of large models through their powerful computing resources, performing complex mathematical operations and data processing. Moreover, the cluster realizes communication between nodes through various types of switch network architectures. For example, the switch network architecture can be implemented as a Fat Tree structure, Leaf-Spine structure, Multi-Plane structure, SuperPOD structure, Multi-Track structure, and so on.
[0077] The processing unit is usually located in a node that is a physical server. The processing unit (e.g., a Pod in Kubernetes) is the smallest deployable and schedulable unit in a cluster and can be implemented as a software module. Specifically, the processing unit may contain one or more closely related containers that share network and storage resources. In large model training, the roles of the processing unit mainly include: (1) Encapsulating training tasks: The processing unit can encapsulate one or more containers for encapsulating and running training tasks. For example, a processing unit may contain one container for executing model training and another container for data loading or logging. (2) Resource sharing: The containers in the processing unit share the same network namespace and storage volume, which enables convenient communication and data sharing between containers. (3) Scheduling and management: The scheduler allocates the running units to appropriate nodes according to resource requirements and the cluster status. (4) Lifecycle management: The processing unit has its own lifecycle and can be created, started, stopped, and deleted.
[0078] The first metric of the current topology to which the training task has been scheduled is used to characterize the communication cost of the current topology, where the communication cost of the current topology is the cumulative cost of communication between all processing units in the current topology. Similarly, the second metric is used to characterize the communication cost of the rescheduled topology, where the communication cost of the rescheduled topology is the cumulative cost of communication between all processing units in the rescheduled topology.
[0079] In one implementation, determine the individual communication cost between each processing unit in the current topology and the remaining processing units in the current topology; based on a predetermined cumulative method and the individual communication cost between each processing unit in the current topology and the remaining processing units in the current topology, determine the first metric. Similarly, determine the individual communication cost between each processing unit in the rescheduled topology and the remaining processing units in the rescheduled topology; based on a predetermined cumulative method and the individual communication cost between each processing unit in the rescheduled topology and the remaining processing units in the rescheduled topology, determine the second metric.
[0080] The individual communication costs of all processing units in the current topology and the rescheduled topology can be accumulated based on various optional types of cumulative methods to obtain the first metric and the second metric respectively. For example, the cumulative methods may include: simple summation method; weighted summation method; linear summation method; cumulative operation based on exponential function; cumulative operation based on power function; cumulative operation based on logarithmic function; cumulative operation based on composite function, and so on.
[0081] Considering that simple summation has the advantage of being easy to implement, it is preferable to use the simple summation method to determine the first index and the second index. In one embodiment, determining the first index of the current topology to which the training task has been scheduled includes: determining the individual communication cost between each processing unit in the current topology and the remaining processing units in the current topology; calculating the first summation result of the individual communication costs of all processing units in the current topology; and determining the first index based on the first summation result.
[0082] In the embodiments of the present invention, various technical parameters can be used to measure the individual communication cost. For example, the technical parameters may include: communication time, communication bandwidth occupancy, network latency, network jitter, and network resource occupancy, etc. The individual communication cost between nodes can be determined based on a single or multiple technical parameters. In a specific implementation, the first index or the second index can be calculated through matrix operations. The specific process includes: determining a communication matrix based on the topological structure, where the communication matrix lists the communication relationships between all nodes in the topology, including the data volume and the communication direction; calculating the individual communication cost between each pair of nodes according to technical parameters such as communication time and bandwidth occupancy; and accumulating the individual communication costs of all pairs of nodes to obtain the total communication cost of the topology, which is the first index or the second index.
[0083] Example: Assume that the current topology to which the training task has been scheduled includes: processing units 1 to 4.
[0084] First, calculate:
[0085] a1: the communication cost between processing unit 1 and processing units 2, 3, and 4;
[0086] a2: the communication cost between processing unit 2 and processing units 1, 3, and 4;
[0087] a3: the communication cost between processing unit 3 and processing units 1, 2, and 4;
[0088] a4: the communication cost between processing unit 4 and processing units 1, 2, and 3.
[0089] Then, determine (a1 + a2 + a3 + a4) as the first index.
[0090] Step 102: Determine the rescheduled topology formed after at least one processing unit in the current topology is virtually migrated.
[0091] Here, at least one processing unit in the current topology can be virtually migrated to form a rescheduling topology. Among them: Any processing unit in the current topology can be virtually migrated to any node with an idle position. Preferably, in a best-effort manner, the processing unit is virtually migrated to the same node as the remaining processing units in the current topology, so as to minimize the communication cost of the topology.
[0092] In one embodiment, step 102 includes:
[0093] (1) Virtually migrate the processing unit in the first node from its current position in the first node to another position different from the current position in the first node.
[0094] Here, the processing units in the current topology are virtually moved within the node to determine the rescheduling topology formed after the virtual migration.
[0095] Figure 2 It is a schematic diagram of migrating virtual processing units within a node according to an embodiment of the present invention.
[0096] In Figure 2 , the current topology includes processing unit 0 to processing unit 3, and processing unit 0 to processing unit 3 are all in the same node 0. The processing unit 3 can be virtually migrated to another processing unit 4 in node 0, thereby forming a rescheduling topology: processing unit 0, processing unit 1, processing unit 2, and processing unit 4.
[0097] (2) Virtually migrate the processing unit in the first node from its current position in the first node to the second node.
[0098] Here, the processing units in the current topology are virtually moved between different nodes of the processing unit to determine the rescheduling topology formed after the virtual migration.
[0099] For example, the virtual movement of the processing unit between different nodes can include:
[0100] (a) Virtually migrate the processing unit in the first node from its current position in the first node to the second node connected to the same switch node as the first node.
[0101] (b) Virtually migrate the processing unit in the first node from its current position in the first node to the second node connected to a different switch node from the first node.
[0102] Figure 3 It is a schematic diagram of virtually migrating a processing unit between nodes according to an embodiment of the present invention.
[0103] In Figure 3Among them, the current topology includes processing units 0 to 3, and processing units 0 to 3 are all in the same node 0. The processing unit 3 can be virtually migrated to the processing unit 4 in node 1, thereby forming a rescheduled topology: processing unit 0 in node 0, processing unit 1 in node 0, processing unit 2 in node 0, and processing unit 4 in node 1. Among them: Node 0 and node 1 can be connected to a common switch node, or node 0 and node 1 are respectively connected to different switch nodes.
[0104] Step 103: Determine a second metric of the rescheduled topology, where the second metric characterizes the communication cost of the rescheduled topology.
[0105] In one embodiment, determining the second metric of the rescheduled topology includes: determining the individual communication cost of each processing unit in the rescheduled topology with the remaining processing units in the rescheduled topology; calculating a second summation result of the individual communication costs of all processing units in the rescheduled topology; and determining the second metric based on the second summation result.
[0106] For example: Figure 3 As shown, assume that the rescheduled topology includes: processing unit 0 in node 0, processing unit 1 in node 0, processing unit 2 in node 0, and processing unit 4 in node 1.
[0107] First, calculate:
[0108] b1: The communication cost between processing unit 0 in node 0 and processing unit 1 in node 0, processing unit 2 in node 0, and processing unit 4 in node 1;
[0109] b2: The communication cost between processing unit 1 in node 0 and processing unit 0 in node 0, processing unit 2 in node 0, and processing unit 4 in node 1;
[0110] b3: The communication cost between processing unit 2 in node 0 and processing unit 0 in node 0, processing unit 1 in node 0, and processing unit 4 in node 1;
[0111] b4: The communication cost between processing unit 4 in node 1 and processing unit 0 in node 0, processing unit 1 in node 0, and processing unit 2 in node 0.
[0112] Then, determine (b1 + b2 + b3 + b4) as the second metric.
[0113] Step 104: When the second metric is less than the first metric, reschedule the training task to the rescheduled topology.
[0114] Here, when the second metric is less than the first metric, the training task is rescheduled to a rescheduling topology, and thus the training task continues to be executed by the rescheduling topology. Therefore, rescheduling the training task to a rescheduling topology with lower communication cost can improve communication efficiency and training efficiency. When an anomaly is detected, instead of simply reporting an error, the processing unit is reconstructed, which can improve stability and reliability.
[0115] In step 102, based on the different virtual migration positions of different processing units, multiple rescheduling topologies may be formed. At this time, the second metric of each rescheduling topology can be calculated to form a second metric set that includes the second metrics of all rescheduling topologies. Then, the smallest second metric is selected from the second metric set. In step 104, when the smallest second metric is less than the first metric, the training task is rescheduled to the rescheduling topology corresponding to the smallest second metric.
[0116] In one embodiment, step 104 includes: evicting at least one processing unit (that is, the at least one processing unit that is virtually migrated in step 102 to form the rescheduling topology); reconstructing at least one processing unit at the position where the at least one processing unit is virtually migrated in the rescheduling topology. Therefore, based on the rescheduling topology, the processing unit is reconstructed, and the processing of the training task can continue to be performed.
[0117] Figure 4 Schematic diagram of virtual migration of processing units in a leaf-spine structure-based cluster according to an embodiment of the present invention. In a leaf-spine structure, a leaf switch group usually refers to a group of switches located at the bottom layer of the network topology, directly connecting to nodes and providing network access functions. The leaf switch group is connected to the spine switch through uplink links. The leaf switch group may include at least one leaf switch. In Figure 4 it, it is assumed that a single processing unit runs in a single node. The leaf switch group 1 is connected to nodes 1 to 8. The leaf switch group 4 is connected to nodes 9 to 16. The leaf switch groups 1 to 4 are respectively connected to the spine switch.
[0118] For the training task, the current topology scheduled by the scheduling unit includes: processing unit 1 in node 1, processing unit 2 in node 2, and processing unit 9 in node 9.
[0119] It is calculated that:
[0120] d1: Communication cost between processing unit 1 and processing unit 2: (S1 + S2);
[0121] d2: Communication cost between processing unit 1 and processing unit 9: (S1 + C1 + C2 + S9).
[0122] d3: Communication cost between processing unit 2 and processing unit 1: (S1 + S2);
[0123] d4: Communication cost between processing unit 2 and processing unit 9: (S2 + C1 + C2 + S9).
[0124] Therefore, the communication cost of the current topology is: d1 + d2 + d3 + d4. Among them: C1 is the communication cost between leaf switch group 1 and the spine switch; C2 is the communication cost between leaf switch group 4 and the spine switch; S2 is the communication cost between processing unit 2 and leaf switch group 1; S9 is the communication cost between processing unit 9 and leaf switch group 4.
[0125] During the rescheduling process, processing unit 9 is virtually moved to the same node as the rest of the processing units in the current topology (for example, at the position of processing unit 8 in node 1), thus obtaining the rescheduled topology: processing unit 1 in node 1, processing unit 2 in node 2, and processing unit 8 in node 1.
[0126] It is calculated that:
[0127] e1: Communication cost between processing unit 1 and processing unit 2: (S1 + S2);
[0128] e2: Communication cost between processing unit 1 and processing unit 8: (S1 + S8).
[0129] e3: Communication cost between processing unit 2 and processing unit 1: (S1 + S2);
[0130] e4: Communication cost between processing unit 2 and processing unit 8: (S1 + S8).
[0131] Among them: S8 is the communication cost between processing unit 8 and leaf switch group 1.
[0132] Therefore, the communication cost of the rescheduled topology includes: e1 + e2 + e3 + e4.
[0133] When e1 + e2 + e3 + e4 is greater than or equal to d1 + d2 + d3 + d4, the training task is not rescheduled to be executed by the rescheduled topology, that is, the training task is still executed by processing unit 1 in node 1, processing unit 2 in node 2, and processing unit 9 in node 9. When e1 + e2 + e3 + e4 is less than d1 + d2 + d3 + d4, the training task can be rescheduled to be executed by the rescheduled topology. At this time, first evict processing unit 9 in node 9, and reconstruct the evicted processing unit at the position of processing unit 8 in node 1, so that the training task is executed by processing unit 1 in node 1, processing unit 2 in node 2, and processing unit 8 in node 1.
[0134] In the above Figure 1 a rescheduling method for training tasks is described. Considering that large model training usually includes multiple parallel training tasks (for example, data parallel tasks, model parallel tasks, pipeline parallelism), the process shown in Figure 1 can be executed for each training task respectively to complete large model training.
[0135] In one embodiment, Figure 1 the method shown further includes: when a processing unit in an abnormal state is detected in the current topology, evicting the processing unit in the abnormal state; and reconstructing the processing unit at the current position of the processing unit in the abnormal state in the current topology.
[0136] Therefore, when an abnormal processing unit is detected, instead of simply reporting an error, the processing unit in the abnormal state is evicted, and the processing unit is reconstructed at the current position in the current topology, thereby improving the stability and reliability of the training task.
[0137] In one embodiment, it includes: when a processing unit in an abnormal state is detected in the current topology, determining an abnormal processing topology formed after the virtual migration of the processing unit in the abnormal state; determining a third metric of the abnormal processing topology, where the third metric characterizes the communication cost of the abnormal processing topology; when the third metric is less than the first metric, evicting the processing unit in the abnormal state and reconstructing the processing unit at the position where the processing unit in the abnormal state is virtually migrated in the abnormal processing topology; when the third metric is greater than or equal to the first metric, evicting the processing unit in the abnormal state and reconstructing the processing unit at the current position of the processing unit in the abnormal state in the current topology.
[0138] Therefore, when an abnormality is detected, instead of simply reporting an error, the processing unit in the abnormal state is evicted, and based on the abnormal processing topology formed after the virtual migration of the processing unit in the abnormal state, the third metric of the abnormal processing topology is calculated. Then, based on the third metric, the reconstruction position of the processing unit is determined, thereby further considering the communication cost factor during the reconstruction process and improving the communication efficiency and training efficiency.
[0139] Figure 5 is a schematic diagram demonstrating the rescheduling process of the training task according to an embodiment of the present invention. As Figure 5 shown, the method includes:
[0140] Step 201: Detect the state of the cluster, where the cluster trains a large model in a parallel manner of multiple training tasks.
[0141] Step 202: Determine whether there are abnormal processing units in the cluster. If so, execute Step 209 and subsequent steps; otherwise, execute Step 203 and subsequent steps. For example, it is possible to determine whether there are abnormalities in the nodes in the cluster to determine whether the processing units in the nodes are abnormal.
[0142] Step 203: Determine whether the rescheduling time has arrived. If so, execute Step 204 and subsequent steps; otherwise, return to execute Step 201.
[0143] Step 204: Traverse all training tasks.
[0144] Step 205: For the currently traversed training task, calculate the communication cost a1 of the current topology of this training task.
[0145] Step 206: For the currently traversed training task, calculate the communication cost a2 of the rescheduling topology of this training task. Among them: in the process of forming the rescheduling topology, in a best-effort manner, virtualize and migrate the processing units to the same node as the remaining processing units in the current topology.
[0146] Step 207: Determine whether a2 is less than a1. If so, execute Step 208; otherwise, return to execute Step 204 and subsequent steps.
[0147] Step 208: Evict and reconstruct the processing units so that the training task is executed by the rescheduling topology, and end this process.
[0148] Step 209: Evict the abnormal processing units.
[0149] Step 210: Reconstruct the abnormal processing units.
[0150] Figure 6 It is a schematic structural diagram of a rescheduling system for training tasks according to an embodiment of the present invention. As Figure 6 shown, the rescheduling system includes: a cluster including a plurality of nodes; a scheduler for scheduling training tasks to the cluster to form a current topology, the current topology including a plurality of processing units in their respective nodes; a rescheduler for when it is determined that the rescheduling time has arrived, determining a first metric of the current topology, the first metric characterizing the communication cost of the current topology; determining a rescheduling topology formed after at least one processing unit in the current topology is virtually migrated; determining a second metric of the rescheduling topology, the second metric characterizing the communication cost of the rescheduling topology; and when the second metric is less than the first metric, rescheduling the training task to the rescheduling topology.
[0151] Nodes in the cluster, acting as physical servers, usually contain high-performance CPUs and GPUs for executing compute-intensive tasks. One or more running units can run in each node. Each running unit can contain a separate container for running a training task, or can contain multiple containers responsible for different functions such as training, data loading, logging, etc. Based on the switch network architecture, nodes in the cluster can communicate with each other.
[0152] The scheduler schedules training tasks to the cluster based on various predefined scheduling principles to form the current topology. For example, the scheduling principles can include: (1) Resource matching principle: The scheduler first checks the resource conditions of each node in the cluster, including the availability of CPUs, memory, GPUs, etc., to ensure that the selected node can meet the resource requirements of the training task. For example, for deep learning training tasks that require a large amount of GPU resources, the scheduler will preferentially select nodes with sufficient GPU resources. Moreover, on the premise of meeting the task resource requirements, the scheduler tries to select nodes with lower resource utilization rates to achieve balanced utilization of cluster resources. This helps to avoid situations where some nodes are overcrowded with resources while other nodes have idle resources, and improves the overall resource utilization efficiency. (2) Affinity and anti-affinity principles: Through node affinity rules, training tasks can be scheduled to nodes with specific labels. For example, tasks that require a specific operating system or software environment can be scheduled to nodes with the corresponding labels to ensure the smooth running of the tasks. (3) Priority and fairness principles: For different training tasks, different priorities can be set according to their importance and urgency. The scheduler will preferentially schedule high-priority tasks to ensure that critical tasks can be executed in a timely manner. For example, for model training tasks that need to be completed quickly, a higher priority can be set. The scheduler tries to ensure that all tasks can obtain fair scheduling opportunities. Even if some tasks cannot be scheduled temporarily due to insufficient resources, they will be included in the scheduling queue and scheduled when resources are available. (4) Load balancing and communication optimization principle: The scheduler tries to evenly distribute training tasks to each node in the cluster to avoid some nodes being overloaded while other nodes are underloaded. In a distributed training scenario, the scheduler considers the communication overhead between nodes and tries to schedule running units that need to communicate frequently to the same node or adjacent nodes, which can reduce network transmission latency and improve training efficiency.
[0153] In one embodiment, a rescheduler is used to evict a processing unit in an abnormal state when a processing unit in an abnormal state is detected in the current topology; and to reconstruct the processing unit at the current position of the current topology where the processing unit in the abnormal state is located.
[0154] In one embodiment, a rescheduler is configured to, when a processing unit in an abnormal state is detected in the current topology, determine an abnormal processing topology formed after the virtual migration of the processing unit in the abnormal state; determine a third metric of the abnormal processing topology, where the third metric characterizes the communication cost of the abnormal processing topology; when the third metric is less than the first metric, evict the processing unit in the abnormal state and reconstruct the processing unit at the location to which the processing unit in the abnormal state is virtually migrated in the abnormal processing topology; when the third metric is greater than or equal to the first metric, evict the processing unit in the abnormal state and reconstruct the processing unit at the current location of the processing unit in the abnormal state in the current topology.
[0155] When the switch architecture in the cluster is implemented as a Fat Tree structure or a Spine-Leaf structure, the rescheduler and the scheduler can be arranged on the leaf nodes.
[0156] In summary, when the rescheduling timing is determined, a first metric of the current topology to which the training task has been scheduled is determined. The current topology includes multiple processing units, and the first metric characterizes the communication cost of the current topology; a rescheduling topology formed after the virtual migration of at least one processing unit in the current topology is determined; a second metric of the rescheduling topology is determined, where the second metric characterizes the communication cost of the rescheduling topology; when the second metric is less than the first metric, the training task is rescheduled to the rescheduling topology. Thus, by rescheduling the training task to be executed by a topology with a lower communication cost, various dynamic changes in the cluster (such as load changes and cluster resource changes, etc.) can be adapted, thereby improving the communication efficiency and training efficiency. In addition, when an abnormality is detected, the processing unit is reconstructed based on multiple mechanisms instead of simply reporting an error, thereby improving the stability and reliability of the training task.
[0157] Figure 7 It is a schematic structural diagram of a rescheduling device for a training task according to an embodiment of the present invention. As Figure 7 shown, the rescheduling device 700 for the training task includes: a first determination module 701, configured to determine a first metric of the current topology to which the training task has been scheduled when it is determined that the rescheduling timing is reached. The current topology includes multiple processing units, and the first metric characterizes the communication cost of the current topology; a second determination module 702, configured to determine a rescheduling topology formed after the virtual migration of at least one processing unit in the current topology; a third determination module 703, configured to determine a second metric of the rescheduling topology, where the second metric characterizes the communication cost of the rescheduling topology; a rescheduling module 704, configured to reschedule the training task to the rescheduling topology when the second metric is less than the first metric.
[0158] In one embodiment, determining the arrival of the rescheduling opportunity includes at least one of the following: determining that the utilization rate of the cluster for executing the training task is lower than a predetermined threshold or the overall topology of the cluster has changed; determining that a user-triggered rescheduling instruction has been received; determining that a predetermined rescheduling time point has been reached.
[0159] In one embodiment, the second determination module 702 is configured to virtually migrate the processing unit in the first node from the current position in the first node to another position different from the current position in the first node; or to virtually migrate the processing unit in the first node from the current position in the first node to the second node.
[0160] In one embodiment, the second determination module 702 is configured to virtually migrate the processing unit in the first node from the current position in the first node to the second node connected to the same switch node as the first node; or to virtually migrate the processing unit in the first node from the current position in the first node to the second node connected to a different switch node from the first node.
[0161] In one embodiment, determining the first metric of the current topology to which the training task has been scheduled includes: determining the individual communication cost of each processing unit in the current topology with the remaining processing units in the current topology; calculating a first summation result of the individual communication costs of all the processing units in the current topology; determining the first metric based on the first summation result; determining the second metric of the rescheduling topology includes: determining the individual communication cost of each processing unit in the rescheduling topology with the remaining processing units in the rescheduling topology; calculating a second summation result of the individual communication costs of all the processing units in the rescheduling topology; determining the second metric based on the second summation result.
[0162] In one embodiment, the rescheduling module 704 is configured to evict at least one processing unit; and reconstruct at least one processing unit at the position in the rescheduling topology to which at least one processing unit has been virtually migrated.
[0163] In one embodiment, the rescheduling module 704 is configured to evict the processing unit in an abnormal state when a processing unit in an abnormal state is detected in the current topology; and reconstruct the processing unit at the current position of the processing unit in an abnormal state in the current topology.
[0164] In one embodiment, the rescheduling module 704 is configured to, when a processing unit in an abnormal state is detected in the current topology, determine an abnormal processing topology formed after the virtual migration of the processing unit in the abnormal state; determine a third metric of the abnormal processing topology, where the third metric characterizes the communication cost of the abnormal processing topology; when the third metric is less than the first metric, evict the processing unit in the abnormal state, and reconstruct the processing unit at the position where the processing unit in the abnormal state is virtually migrated in the abnormal processing topology; when the third metric is greater than or equal to the first metric, evict the processing unit in the abnormal state, and reconstruct the processing unit at the current position of the processing unit in the abnormal state in the current topology.
[0165] An embodiment of the present invention also provides an electronic device with a processor-memory architecture. Figure 8 It is a structural diagram of the electronic device according to the embodiment of the present invention. As Figure 8 shown, the electronic device includes a processor 801, a memory 802, and a computer program stored on the memory 802 and executable on the processor 801. When the computer program is executed by the processor 801, it implements the rescheduling method for the training task as described above. Among them, the memory 802 can be specifically implemented as various storage media such as electrically erasable programmable read-only memory (EEPROM), flash memory, programmable read-only memory (PROM), etc. The processor 801 can be implemented as including one or more central processing units or one or more field-programmable gate arrays, where the field-programmable gate array integrates one or more central processing unit cores. Specifically, the central processing unit or the central processing unit core can be implemented as a CPU, GPU, GPGPU, MCU, or DSP, etc.
[0166] It should be noted that not all steps and modules in the above-mentioned processes and structural diagrams are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted according to needs. The division of each module is only for the convenience of description in terms of functional division. In actual implementation, one module can be implemented by multiple modules, and the functions of multiple modules can also be implemented by the same module. These modules can be located in the same device or in different devices.
[0167] The hardware modules in each embodiment can be implemented mechanically or electronically. For example, a hardware module can include specially designed permanent circuits or logic devices (such as dedicated processors, such as FPGAs or ASICs) for performing specific operations. For instance, specific operations can be completed in various types of chips (e.g., artificial intelligence chips). A hardware module can also include programmable logic devices or circuits (such as including general-purpose processors or other programmable processors) temporarily configured by software for executing specific operations. As for whether to specifically adopt a mechanical approach, or use dedicated permanent circuits, or use temporarily configured circuits (such as configured by software) to implement the hardware module, it can be determined based on cost and time considerations.
[0168] The present invention also provides a machine-readable storage medium storing instructions for causing a machine to execute the methods described in this application. Specifically, a system or device equipped with a storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program code stored in the storage medium. In addition, some or all of the actual operations can also be completed by an operating system operating on the computer based on the instructions of the program code. The program code read from the storage medium can also be written to the memory provided in the expansion board inserted into the computer or to the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU etc. installed on the expansion board or expansion unit are caused to execute some and all of the actual operations, thereby implementing the functions of any one of the above embodiments. Embodiments of the storage medium for providing the program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer or cloud via a communication network.
[0169] In this document, "schematic" means "serving as an example, instance or illustration", and any illustration or embodiment described as "schematic" in this document should not be construed as a more preferred or more advantageous technical solution. To make the drawings concise, only the parts related to the present invention are schematically shown in each drawing, and do not represent their actual structure as a product. Additionally, to make the drawings concise and easy to understand, in some drawings, for components with the same structure or function, only one of them is schematically illustrated, or only one of them is labeled. In this document, "a" does not mean that the quantity of the parts related to the present invention is limited to "only one", and "a" does not exclude the case where the quantity of the parts related to the present invention is "more than one". In this document, "upper", "lower", "front", "rear", "left", "right", "inner", "outer", etc. are only used to represent the relative positional relationship between relevant parts, rather than defining the absolute positions of these relevant parts.
[0170] The above is only a preferred embodiment of the present invention, and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A rescheduling method for training tasks, characterized in that, Including: When it is determined that the rescheduling time point is reached, determine a first metric of the current topology to which the training task has been scheduled, where the current topology includes multiple processing units, and the first metric characterizes the communication cost of the current topology, and the processing units are located in nodes; Determine a rescheduling topology formed after at least one processing unit in the current topology is virtually migrated; Determine a second metric of the rescheduling topology, where the second metric characterizes the communication cost of the rescheduling topology; When the second metric is less than the first metric, reschedule the training task to the rescheduling topology; The determining the rescheduling topology formed after at least one processing unit in the current topology is virtually migrated includes: Virtually migrate the processing unit in the first node from its current position in the first node to another position different from the current position in the first node; Or Virtually migrate the processing unit in the first node from its current position in the first node to a second node.
2. The method according to claim 1, characterized in that, The determining that the rescheduling time point is reached includes at least one of the following: Determine that the utilization rate of the cluster used to execute the training task is lower than a predetermined threshold or the overall topology of the cluster has changed; Determine that a user-triggered rescheduling instruction is received; Determine that a predetermined rescheduling time point is reached.
3. The method according to claim 1, characterized in that, The virtually migrating the processing unit in the first node from its current position in the first node to a second node includes: Virtually migrate the processing unit in the first node from its current position in the first node to a second node connected to the same switch node as the first node; or Virtually migrate the processing unit in the first node from its current position in the first node to a second node connected to a different switch node from the first node.
4. The method according to claim 1, wherein The determining the first metric of the current topology to which the training task has been scheduled includes: determining the individual communication cost of each processing unit in the current topology with the remaining processing units in the current topology; calculating a first summation result of the individual communication costs of all processing units in the current topology; and determining the first metric based on the first summation result; The determining the second metric of the rescheduling topology includes: determining the individual communication cost of each processing unit in the rescheduling topology with the remaining processing units in the rescheduling topology; calculating a second summation result of the individual communication costs of all processing units in the rescheduling topology; and determining the second metric based on the second summation result.
5. The method according to any one of claims 1-4, characterized in that, The rescheduling the training task to the rescheduling topology includes: Evicting the at least one processing unit; Reconstructing the at least one processing unit at the position in the rescheduling topology to which the at least one processing unit is virtually migrated.
6. The method according to any one of claims 1 to 4, characterized in that, Including: When a processing unit in an abnormal state is detected in the current topology, evict the processing unit in the abnormal state; Reconstruct the processing unit at the current position of the processing unit in the abnormal state in the current topology.
7. The method according to any one of claims 1-4, characterized in that, Including: When a processing unit in an abnormal state is detected in the current topology, determine an abnormal processing topology formed after the virtual migration of the processing unit in the abnormal state; Determine a third metric of the abnormal processing topology, where the third metric characterizes the communication cost of the abnormal processing topology; When the third metric is less than the first metric, evict the processing unit in the abnormal state, and reconstruct the processing unit at the position in the abnormal processing topology to which the processing unit in the abnormal state is virtually migrated; When the third metric is greater than or equal to the first metric, evict the processing unit in the abnormal state, and reconstruct the processing unit at the current position of the processing unit in the abnormal state in the current topology.
8. A rescheduling device for training tasks, characterized in that, Includes: A first determination module, configured to determine a first metric of the current topology to which a training task has been scheduled when determining that a rescheduling opportunity has arrived, where the current topology includes multiple processing units, and the first metric characterizes the communication cost of the current topology, and the processing units are located in nodes; A second determination module, configured to determine a rescheduling topology formed after virtual migration of at least one processing unit in the current topology; A third determination module, configured to determine a second metric of the rescheduling topology, where the second metric characterizes the communication cost of the rescheduling topology; A rescheduling module, configured to reschedule the training task to the rescheduling topology when the second metric is less than the first metric; The determination of the rescheduling topology formed after virtual migration of at least one processing unit in the current topology includes: Virtually migrate the processing unit in the first node from its current position in the first node to another position different from the current position in the first node; Or Virtually migrate the processing unit in the first node from its current position in the first node to a second node.
9. A rescheduling system for training tasks, characterized in that, Includes: A cluster, including multiple nodes; A scheduler, configured to schedule a training task to the cluster to form a current topology, where the current topology includes multiple processing units in their respective nodes, and the processing units are located in nodes; A rescheduler, configured to determine a first metric of the current topology when determining that a rescheduling opportunity has arrived, where the first metric characterizes the communication cost of the current topology; Determine a rescheduling topology formed after virtual migration of at least one processing unit in the current topology; Determine a second metric of the rescheduling topology, where the second metric characterizes the communication cost of the rescheduling topology; when the second metric is less than the first metric, reschedule the training task to the rescheduling topology; The determination of the rescheduling topology formed after virtual migration of at least one processing unit in the current topology includes: Virtually migrate the processing unit in the first node from its current position in the first node to another position different from the current position in the first node; Or Virtually migrate the processing unit in the first node from its current position in the first node to a second node.
10. An electronic device, characterized in that, Includes: A memory; A processor; Among them, an application program executable by the processor is stored in the memory, and is used to cause the processor to execute the rescheduling method of the training task described in any one of claims 1-7.
11. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by the processor, the processor is caused to execute the rescheduling method of the training task described in any one of claims 1-7.
12. A program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the rescheduling method of the training task described in any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Scheduling strategy determination method and system for pipeline parallel training
CN116450312A