Task scheduling method and device, equipment and medium

CN121603495BActive Publication Date: 2026-09-22TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511696499.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-09-22
Estimated Expiration
2045-11-18

AI Technical Summary

Benefits of technology

[0015]本申请实施例提供的任务调度方法、装置、设备及介质,通过构建待调度的多个子任务对应的任务间通信矩阵以及待调度的多个任务处理节点对应的网络亲和度矩阵,分别对多个子任务和多个任务处理节点进行聚类处理,并对聚类得到的任务簇集合和节点簇集合进行映射,得到高质量的初始映射结果,针对初始映射结果通过执行邻域动作和通信开销进行优化迭代,通过两个阶段的粗调度和精调度的结合,在保障调度质量的同时有效控制响应时间,有助于降低整体通信成本,减少通信瓶颈,兼容不同的网络拓扑结构,提高整个深度学习任务的效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121603495B_ABST
    Figure CN121603495B_ABST
Patent Text Reader

Abstract

The application provides a task scheduling method and device, equipment and medium, the method comprises the following steps: constructing a task intercommunication matrix corresponding to a plurality of subtasks to be scheduled and a network affinity matrix corresponding to a plurality of task processing nodes to be scheduled; based on the task intercommunication matrix and the network affinity matrix, the plurality of subtasks and the plurality of task processing nodes are respectively clustered, and the task cluster set and the node cluster set obtained by clustering are mapped to determine the initial mapping result between the plurality of subtasks and the plurality of task processing nodes; the neighborhood action is performed on the initial mapping result to obtain the optimized mapping result; based on the comparison result between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, the initial mapping result is iteratively optimized until the iteration stopping condition is met, and the target mapping result is obtained. Through the combination of coarse scheduling and fine scheduling, the response time is effectively controlled while the scheduling quality is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to a task scheduling method, apparatus, device, and medium. Background Technology

[0002] With the proliferation of large-scale deep learning tasks (such as models with billions to hundreds of billions of parameters), a single machine often cannot meet the load requirements. Distributed parallelism has become a new option, which means that deep learning tasks are distributed across different machines to work together. Deep learning tasks are divided into multiple parallel methods and distributed training is carried out in a multi-machine environment.

[0003] When scheduling tasks, random allocation or matching based on computing resources is often used, which results in long search times and low scheduling quality. Summary of the Invention

[0004] In view of this, this application provides a task scheduling method, apparatus, device and medium that effectively controls response time while ensuring scheduling quality through a combination of two stages of coarse scheduling and fine scheduling.

[0005] Specifically, this application is implemented through the following technical solution: According to a first aspect of this application, a task scheduling method is provided, the method comprising: Construct an inter-task communication matrix corresponding to multiple subtasks to be scheduled and a network affinity matrix corresponding to multiple task processing nodes to be scheduled. The multiple subtasks to be scheduled are obtained by partitioning the deep learning distributed task to be scheduled. The inter-task communication matrix is ​​used to indicate the task communication volume between every two subtasks in the multiple subtasks. The network affinity matrix is ​​used to indicate the bandwidth and congestion index between every two task processing nodes in the multiple task processing nodes. Based on the inter-task communication matrix and the network affinity matrix, clustering is performed on the multiple subtasks and the multiple task processing nodes, and the clustered task cluster set and node cluster set are mapped to determine the initial mapping result between the multiple subtasks and the multiple task processing nodes. The initial mapping result is processed by a neighborhood action to obtain an optimized mapping result. The neighborhood action is used to adjust the mapping relationship between some subtasks and task processing nodes. Based on the comparison between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, the initial mapping result is optimized iteratively until the iteration stopping condition is met. The optimized mapping result when the iteration stopping condition is met is determined as the target mapping result between the multiple subtasks and the multiple task processing nodes.

[0006] In one optional implementation, the execution of the neighborhood action includes: For any two task clusters that have the same number of subtasks, swap the node clusters mapped to the two task clusters; Alternatively, for any two subtasks in a task cluster that are communicating, swap the task processing nodes mapped to those two subtasks. Alternatively, at least one subtask from any task cluster can be mapped to another idle node cluster.

[0007] In one optional implementation, the communication overhead corresponding to the initial mapping result is determined through the following steps: When scheduling tasks according to the initial mapping result, determine the bandwidth and congestion index between every two task processing nodes in the plurality of task processing nodes, and construct the network affinity matrix corresponding to the initial mapping result; Based on the inter-task communication matrix and the network affinity matrix corresponding to the initial mapping result, the communication overhead corresponding to the initial mapping result is determined.

[0008] In one optional implementation, the step of optimizing the initial mapping result based on the comparison result between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, until the iteration stopping condition is met, includes: Based on the comparison between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, the acceptance probability of the optimized mapping result is determined, and the acceptance probability is used to indicate the likelihood of accepting the optimized mapping result; The accepted mapping result is used as the initial mapping result, and the neighborhood action is performed on the initial mapping result until the iteration stopping condition is met.

[0009] In one optional implementation, determining the acceptance probability of the optimized mapping result based on a comparison between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result includes: If the communication overhead corresponding to the initial mapping result is greater than the communication overhead corresponding to the optimized mapping result, the probability of accepting the optimized mapping result is determined to be 1. If the communication overhead corresponding to the initial mapping result is less than or equal to the communication overhead corresponding to the optimized mapping result, the acceptance probability of the optimized mapping result is determined based on the communication overhead corresponding to the initial mapping result, the communication overhead corresponding to the optimized mapping result, and the overhead parameter.

[0010] In one optional implementation, when the initial mapping result is optimized and iterated using simulated annealing, the overhead parameter used in the first round of iteration is a preset overhead parameter, and the overhead parameters used in the remaining rounds of iteration are generated by reducing the overhead parameter used in the previous round of iteration, with each round of iteration including the same number of iterations.

[0011] In one optional implementation, the iteration stopping condition is that the cost parameter used is reduced to the termination cost parameter, and the optimized mapping result when the iteration stopping condition is met is the mapping result accepted in the last iteration of the current iteration in which the optimization iteration is performed using the termination cost parameter.

[0012] According to a second aspect of this application, a task scheduling apparatus is provided, the apparatus comprising: The matrix construction module is used to construct the inter-task communication matrix corresponding to the multiple sub-tasks to be scheduled and the network affinity matrix corresponding to the multiple task processing nodes to be scheduled. The multiple sub-tasks to be scheduled are obtained by partitioning the deep learning distributed task to be scheduled. The inter-task communication matrix is ​​used to indicate the task communication volume between every two sub-tasks in the multiple sub-tasks. The network affinity matrix is ​​used to indicate the bandwidth and congestion index between every two task processing nodes in the multiple task processing nodes. The clustering module is used to perform clustering processing on the multiple subtasks and the multiple task processing nodes based on the inter-task communication matrix and the network affinity matrix, respectively, and to map the clustered task cluster set and node cluster set to determine the initial mapping result between the multiple subtasks and the multiple task processing nodes. An action execution module is used to perform neighborhood actions on the initial mapping result to obtain an optimized mapping result. The neighborhood actions are used to adjust the mapping relationship between some subtasks and task processing nodes. The optimization iteration module is used to optimize and iterate the initial mapping result based on the comparison result between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, until the iteration stopping condition is met, and the optimized mapping result when the iteration stopping condition is met is determined as the target mapping result between the multiple subtasks and the multiple task processing nodes.

[0013] According to a third aspect of this application, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the task scheduling method described in the first aspect above.

[0014] According to a fourth aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the task scheduling method described in the first aspect above.

[0015] The task scheduling method, apparatus, device, and medium provided in this application construct an inter-task communication matrix corresponding to multiple sub-tasks to be scheduled and a network affinity matrix corresponding to multiple task processing nodes to be scheduled. They then perform clustering processing on the multiple sub-tasks and multiple task processing nodes, and map the clustered task cluster set and node cluster set to obtain a high-quality initial mapping result. The initial mapping result is then optimized iteratively by executing neighborhood actions and reducing communication overhead. Through the combination of two stages of coarse and fine scheduling, the scheduling quality is guaranteed while effectively controlling response time, which helps reduce overall communication costs, minimizes communication bottlenecks, is compatible with different network topologies, and improves the efficiency of the entire deep learning task.

[0016] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this disclosure.

[0017] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a task scheduling method in an exemplary embodiment of this application; Figure 2 This is a communication diagram between subtasks illustrated in an exemplary embodiment of this application; Figure 3 This is a schematic diagram illustrating the bandwidth between task processing nodes according to an exemplary embodiment of this application; Figure 4 This is a schematic diagram illustrating a congestion index between task processing nodes according to an exemplary embodiment of this application; Figure 5 This is a schematic diagram illustrating the distribution of eigenvalues ​​of a symmetric normalized Laplace matrix according to an exemplary embodiment of this application; Figure 6 This is a schematic diagram illustrating a method for determining a task cluster, as shown in an exemplary embodiment of this application; Figure 7 This is a schematic diagram illustrating a task scheduling process according to an exemplary embodiment of this application; Figure 8a This is an iterative schematic diagram illustrating a communication overhead according to an exemplary embodiment of this application; Figure 8bThis is an iterative schematic diagram illustrating another communication overhead, as shown in an exemplary embodiment of this application; Figure 8c This is an iterative schematic diagram illustrating yet another communication overhead, as shown in an exemplary embodiment of this application; Figure 8d This is an iterative schematic diagram illustrating another communication overhead according to an exemplary embodiment of this application; Figure 9 This is a comparison diagram of communication overhead under different scheduling strategies, illustrating an exemplary embodiment of this application; Figure 10 This is a schematic diagram of a task scheduling device shown in an exemplary embodiment of this application; Figure 11 This is a schematic diagram of the structure of a computer device shown in an exemplary embodiment of this application. Detailed Implementation

[0019] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0020] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0021] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0022] Research has found that when scheduling tasks, random allocation is often used, or matching is done based on computing resources, or simple algorithms such as greedy allocation algorithms are used for allocation. This results in long search times and low scheduling quality.

[0023] Based on the above research, this application provides a task scheduling method. By constructing the inter-task communication matrix corresponding to multiple sub-tasks to be scheduled and the network affinity matrix corresponding to multiple task processing nodes to be scheduled, the method performs clustering processing on multiple sub-tasks and multiple task processing nodes respectively, and maps the clustered task cluster set and node cluster set to obtain a high-quality initial mapping result. The initial mapping result is then optimized iteratively by executing neighborhood actions and communication overhead. Through the combination of coarse scheduling and fine scheduling in two stages, the method effectively controls the response time while ensuring scheduling quality, which helps to reduce the overall communication cost and reduce communication bottlenecks.

[0024] To facilitate understanding of this embodiment, a task scheduling method disclosed in this application will first be described in detail. The execution entity of the task scheduling method provided in this application is generally an electronic device with a certain computing power. This electronic device can be a server, which can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms. In some possible implementations, this task scheduling method can be implemented by the processor calling computer-readable instructions stored in memory.

[0025] The following description, in conjunction with the accompanying drawings, illustrates a task scheduling method provided by an embodiment of this application.

[0026] See Figure 1 The diagram shown is a flowchart illustrating a task scheduling method according to an exemplary embodiment of this application. Figure 1 As shown in the figure, the task scheduling method provided in this embodiment includes steps S101 to S104, wherein: S101: Construct the inter-task communication matrix corresponding to the multiple sub-tasks to be scheduled and the network affinity matrix corresponding to the multiple task processing nodes to be scheduled. The multiple sub-tasks to be scheduled are obtained by partitioning the deep learning distributed task to be scheduled. The inter-task communication matrix is ​​used to indicate the task communication volume between every two sub-tasks in the multiple sub-tasks. The network affinity matrix is ​​used to indicate the bandwidth and congestion index between every two task processing nodes in the multiple task processing nodes.

[0027] Here, the deep learning distributed task to be scheduled can be, for example, model training, neuromorphic computing, etc. The multiple subtasks obtained by dividing the deep learning distributed task can be assigned to different task processing nodes for collaborative work.

[0028] In some possible implementations, the plurality of subtasks are obtained through the following steps: Based on the task structure of the deep learning distributed task to be scheduled and the predetermined parallel strategy, the deep learning distributed task is divided into multiple subtasks, and the parallel strategy is used to indicate the task division method.

[0029] In this step, based on the task structure of the deep learning distributed task, the overall task execution flow and the dependencies between each execution step can be clarified. Then, a pre-determined parallel strategy is invoked, and the deep learning distributed task is divided into multiple subtasks according to the partitioning method indicated by the parallel strategy. Optionally, the input data, output data, and resource requirement information of each subtask can also be labeled for subsequent mapping.

[0030] In this way, dividing tasks according to their structure and parallel strategies in deep learning distributed tasks can improve the effectiveness of subtask partitioning and provide a high-quality data foundation for subsequent task scheduling.

[0031] In some possible implementations, the parallel strategy includes at least one of data parallelism, pipeline parallelism, and tensor parallelism. Data parallelism (DP) often requires global gradient synchronization, such as Ring-Allreduce, and under this strategy, multiple data parallel groups can be obtained. Pipeline parallelism (PP) often requires inter-stage activation or gradient exchange during the forward propagation and / or backward propagation stages, and under this strategy, multiple task execution stages can be obtained. Tensor parallelism (TP) involves dense communication in high-speed intra-machine channels (such as NVLink), and under this strategy, multiple task execution stages can be obtained.

[0032] Optionally, a single-strategy approach can be used for task partitioning. Alternatively, a hybrid strategy can be employed, combining the advantages of data parallelism, pipelined parallelism, and tensor parallelism to maximize computational resource utilization. The hybrid strategy approach is more complex than the single-strategy approach because it requires consideration of the interactions between different strategy dimensions.

[0033] For example, if a three-dimensional hybrid strategy of data parallelism, pipeline parallelism, and tensor parallelism is adopted, the number of subtasks is represented by the following formula (1): (1) in, Indicates the number of subtasks. This indicates the number of subtasks divided under a data parallelism strategy. This indicates the number of subtasks divided under the pipeline parallelism strategy. This represents the number of subtasks divided under the tensor parallel strategy.

[0034] Here, for strategies that are not adopted, the number of subtasks can be 1.

[0035] For example, a data parallel strategy can be divided into 3 data parallel groups, a pipelined parallel strategy into 3 task execution stages, and a tensor parallel strategy into 2 task execution stages, resulting in a total of 18 subtasks: A1 (A11, A12), A2 (A21, A22), A3 (A31, A32), B1 (B11, B12), B2 (B21, B22), B3 (B31, B32), C1 (C11, C12), C2 (C21, C22), and C3 (C31, C32). This is understandable. The first data parallel group includes A1 (A11, A12), A2 (A21, A22), and A3 (A31, A32); the second data parallel group includes B1 (B11, B12), B2 (B21, B22), and B3 (B31, B32); and the third data parallel group includes C1 (C11, C12), C2 (C21, C22), and C3 (C31, C32). Each data parallel group has three task execution stages. Taking the first data parallel group as an example, it includes the three task execution stages A1, A2, and A3. Each task execution stage has two task execution steps. Taking task execution stage A1 as an example, it includes the two task execution steps A11 and A12.

[0036] In some possible implementations, the inter-task communication matrix corresponding to the multiple subtasks to be scheduled is constructed through the following steps: Based on the task communication volume between every two subtasks in the multiple subtasks to be scheduled, the inter-task communication matrix is ​​constructed.

[0037] Specifically, the task communication between two subtasks can be determined through the following steps: The execution position of each subtask under each strategy is determined. The execution position of the subtask under the data parallel strategy indicates the data parallel group to which the subtask belongs. The execution position of the subtask under the pipeline parallel strategy indicates the execution stage to which the subtask belongs. The execution position of the subtask under the tensor parallel strategy indicates the execution link to which the subtask belongs. Based on the execution positions of the two subtasks, the task communication volume between the two subtasks is determined.

[0038] In the above steps, for each of the two subtasks, the task execution position of each subtask under each task strategy can be determined. Based on the respective task execution positions of the two subtasks, the basic communication volume between the two subtasks is estimated, and the basic communication volume is determined as the task communication volume.

[0039] In this way, when determining the task communication volume, the task execution position of the subtask under each strategy is taken into consideration, so that the subtasks with high communication volume can be assigned to the same cluster during subsequent clustering processing.

[0040] In some possible implementations, the method further includes: If two subtasks belong to different data parallel groups and are in the same task execution stage, determine the data parallel communication volume between the two subtasks. The task communication volume between the two subtasks is determined based on the basic task communication volume and the data parallel communication volume between the two subtasks.

[0041] Following the example above, subtasks A11 and B11 belong to different data parallel groups and have the same task execution stage. When determining the task communication volume between subtasks A11 and B11, data parallel communication volume can be added on the basis of the basic task communication volume.

[0042] In this way, for two subtasks that belong to different data parallel groups and have the same task execution stage, adding data parallel communication volume on top of the basic communication volume helps to improve the comprehensiveness and accuracy of determining task communication volume.

[0043] In some other possible implementations, the method further includes: If two subtasks belong to the same data parallel group and their execution phases are adjacent, determine the pipeline parallel communication volume between the two subtasks. The task communication volume between the two subtasks is determined based on the basic task communication volume and the pipeline parallel communication volume between the two subtasks.

[0044] Following the example above, subtasks A11 and A21 belong to the same data parallel group and their task execution stages are adjacent. When determining the task communication volume between subtasks A11 and A21, pipeline parallel communication volume can be added on the basis of the basic task communication volume.

[0045] In this way, for two subtasks that are adjacent in execution phase and share the same data parallel group, adding pipelined parallel communication on top of the basic communication volume helps to improve the comprehensiveness and accuracy of determining task communication volume.

[0046] In the above embodiments, the data parallel group partitioning under the data parallel strategy and the task execution stage partitioning under the pipelined parallel strategy are considered. The task execution stage partitioning under the tensor parallel strategy is not considered because the tensor parallel strategy has higher communication requirements than the data parallel strategy and the pipelined parallel strategy, and is generally completed within a single machine. However, this embodiment is aimed at multiple task processing nodes, that is, multiple machines. At the machine level, the communication volume of the tensor parallel strategy can be ignored to reduce the scale.

[0047] For example, it can be done through Subtasks sub-tasks The task communication between each pair of subtasks By splicing them together, you can generate 3D inter-task communication matrix .

[0048] To more intuitively illustrate the communication volume between subtasks, please refer to [link / reference]. Figure 2 This is a communication diagram between subtasks illustrated in an exemplary embodiment of this application. Figure 2 As shown in the diagram, taking 61 subtasks as an example, the horizontal axis represents the target subtask, and the vertical axis represents the source subtask. During communication between subtasks, data is sent from the source subtask to the target subtask; that is, the source subtask acts as the starting point of communication, and the target subtask acts as the destination. To visually represent the communication volume between each pair of subtasks, the intensity of the color is used to indicate the volume; the darker the color, the greater the communication volume.

[0049] In practical applications, each task processing node is a machine device, including but not limited to a Graphics Processing Unit (GPU), a Central Processing Unit (CPU), a Field Programmable Gate Array (FPGA), an Artificial Intelligence (AI) accelerator card, and an edge computing box. Multiple task processing nodes to be scheduled can be nodes within the same topology, such as topologies with a clear grouping structure like Fat-Tree or Dragonfly, or relatively balanced topologies like Ramanujan or Jellyfish.

[0050] In some possible implementations, the network affinity matrix corresponding to the multiple task processing nodes to be scheduled is constructed through the following steps: The network affinity matrix is ​​constructed based on the bandwidth and congestion metrics between every two task processing nodes among the multiple task processing nodes to be scheduled.

[0051] In some possible implementations, the bandwidth and congestion metrics between the two task processing nodes are determined through the following steps: Determine the connection relationship between two task processing nodes, including direct connection, indirect connection, and no connection; Based on the connection relationship between the two task processing nodes, determine the bandwidth and congestion indicators between the two task processing nodes.

[0052] In the above steps, within the topology formed by the multiple task processing nodes, if one of the two task processing nodes can directly reach the other, the connection between the two task processing nodes can be determined as a direct connection; if one of the two task processing nodes needs to go through other nodes to reach the other, the connection between the two task processing nodes can be determined as an indirect connection; if one of the two processing nodes is unreachable from the other, the connection between the two task processing nodes can be determined as a disconnection. Based on the connection relationships between the two task processing nodes, the bandwidth and congestion indicators between the two task processing nodes are determined.

[0053] In this way, when determining the bandwidth and congestion indicators between two task processing nodes, the connection between the two task processing nodes is taken into account, ensuring that the determined bandwidth and congestion indicators can reflect the actual node topology, improving the accuracy of the determined bandwidth and congestion indicators, and enhancing the adaptability of task scheduling.

[0054] In some possible implementations, determining the bandwidth and congestion metrics between the two task processing nodes based on the connection relationship between them includes: When the connection relationship is a direct connection, the original bandwidth, retransmission rate, network interface queue length and packet loss rate between the two task processing nodes are normalized. The normalized original bandwidth is determined as the bandwidth between the two task processing nodes; Based on the normalized retransmission rate, the normalized network interface queue length, and the normalized packet loss rate, the congestion index between the two task processing nodes is determined.

[0055] In the above steps, if the connection relationship is a direct connection, the raw bandwidth, retransmission rate, network interface queue length, and packet loss rate between the two task processing nodes can be collected.

[0056] For example, see also Figure 3This is a schematic diagram illustrating bandwidth between task processing nodes as an exemplary embodiment of this application. The bandwidth shown in this example is the raw bandwidth, such as... Figure 3 As shown, gray dots represent task processing nodes. To visually demonstrate the bandwidth between pairs of task processing nodes, the intensity of the line connecting the gray dots is used to represent the bandwidth; the darker the line, the greater the bandwidth value.

[0057] The raw bandwidth, retransmission rate, network interface queue length, and packet loss rate between the two task processing nodes are normalized.

[0058] When normalizing the original bandwidth, specifically, the ratio between the original bandwidth and the bandwidth threshold can be determined, and the smaller value between this ratio and 1 is determined as the original bandwidth after normalization, thereby normalizing the bandwidth to the range of [0,1].

[0059] For example, the original bandwidth after normalization can be determined by the following formula (2): (2) in, This represents the normalized raw bandwidth between task processing node u and task processing node v; a larger value indicates better bandwidth. This represents the raw bandwidth between task processing node u and task processing node v. This indicates the bandwidth threshold.

[0060] The specific value of the bandwidth threshold can be set according to the actual task scheduling needs, and there is no restriction here. For example, it can be 10Gbps. The original bandwidth exceeding the bandwidth threshold is considered to fully meet the requirements.

[0061] The specific methods for normalizing the retransmission rate, network interface queue length, and packet loss rate are similar to those for normalizing the original bandwidth, and can be referred to the description in the above embodiments, which will not be repeated here.

[0062] In this embodiment, the normalized original bandwidth can be determined as the bandwidth between the two task processing nodes. Based on the normalized retransmission rate, the normalized network interface queue length, and the normalized packet loss rate, the congestion index between the two task processing nodes can be determined.

[0063] When determining the congestion index between two task processing nodes based on the normalized retransmission rate, normalized network interface queue length, and normalized packet loss rate, specifically, the normalized retransmission rate, normalized network interface queue length, and normalized packet loss rate can be weighted and summed to obtain the congestion index between the two task processing nodes. The specific values ​​of the weights corresponding to the normalized retransmission rate, normalized network interface queue length, and normalized packet loss rate can be set according to the actual task scheduling needs and are not restricted here.

[0064] For example, the congestion index in the case of direct connection can be determined by the following formula (3): (3) in, This represents the congestion metric between directly connected task processing nodes u and v. This represents the normalized retransmission rate between directly connected task processing nodes u and v. This represents the weight corresponding to the retransmission rate after normalization. This represents the normalized network interface queue length between directly connected task processing nodes u and v. This represents the weight corresponding to the normalized network interface queue length. This represents the normalized packet loss rate between directly connected task processing nodes u and v. This represents the weight corresponding to the packet loss rate after normalization.

[0065] For example, see also Figure 4 This is a schematic diagram illustrating a congestion index between task processing nodes, as shown in an exemplary embodiment of this application. Figure 4 As shown, gray dots represent task processing nodes. To visually represent the congestion indicators between pairs of task processing nodes, the intensity of the lines connecting the gray dots is used to indicate the congestion level; the darker the line, the higher the congestion level.

[0066] In this way, for directly connected task processing nodes, the original bandwidth, retransmission rate, network interface queue length, and packet loss rate between the two task processing nodes are normalized to eliminate the difference in units. Then, the normalized original bandwidth is taken as the required bandwidth. The normalized retransmission rate, normalized network interface queue length, and normalized packet loss rate are weighted and summed to obtain the congestion index, thus avoiding the quality of a certain value from affecting the accuracy of the congestion index.

[0067] In some possible implementations, determining the bandwidth and congestion metrics between the two task processing nodes based on the connection relationship between them includes: In the case where the connection relationship is an indirect connection, the original bandwidth, retransmission rate, network interface queue length and packet loss rate corresponding to the connection path between the two task processing nodes are normalized. The normalized original bandwidth corresponding to the shortest connection path between two task processing nodes is determined as the bandwidth between the two task processing nodes. Based on the normalized retransmission rate, normalized network interface queue length, and normalized packet loss rate of the connection path between the two task processing nodes, the congestion index between the two task processing nodes is determined.

[0068] In the above steps, if the connection relationship is an indirect connection, such as communication between two task processing nodes requiring forwarding through a switch, a connection path between the two task processing nodes can be determined. The connection path includes multiple directly connected task processing nodes. The original bandwidth, retransmission rate, network interface queue length, and packet loss rate between each pair of directly connected task processing nodes on the connection path can be collected, and the original bandwidth, retransmission rate, network interface queue length, and packet loss rate between each pair of directly connected task processing nodes on the connection path can be normalized. Here, the specific method of normalization is similar to that in the aforementioned embodiments (e.g., formula (2)), and can be referred to the description in the aforementioned embodiments, which will not be repeated here.

[0069] The effective bandwidth between two indirectly connected task processing nodes depends on the bottleneck link on the shortest connection path. In this embodiment of the disclosure, the original bandwidth with the smallest value is selected from the normalized original bandwidth between every two directly connected task processing nodes on the connection path. That is, the normalized original bandwidth corresponding to the shortest connection path between the two task processing nodes is determined, and this value is determined as the bandwidth between the two task processing nodes.

[0070] For example, the bandwidth in the case of indirect connection can be determined by the following formula (4): (4) in, This represents the bandwidth between indirectly connected task processing nodes u and v, where P represents the connection path between task processing nodes u and v, and task processing nodes x and y are two directly connected task processing nodes on the connection path P. This represents the normalized raw bandwidth between task processing node x and task processing node y.

[0071] The congestion index between two indirectly connected task processing nodes depends on the congestion index between every two directly connected task processing nodes on the connection path between the two task processing nodes. When determining the congestion index between two task processing nodes based on the normalized retransmission rate, normalized network interface queue length, and normalized packet loss rate corresponding to the connection path between the two task processing nodes, specifically, for every two directly connected task processing nodes on the connection path between the two task processing nodes, the congestion index between the two task processing nodes can be determined based on the normalized retransmission rate, normalized network interface queue length, and normalized packet loss rate between the two task processing nodes. Here, the specific method for determining the congestion index is similar to that in the aforementioned embodiments (e.g., formula (3)), and can be referred to the description in the aforementioned embodiments, which will not be repeated here. Then, based on the congestion index between every two directly connected task processing nodes on the connection path, the congestion index between the two task processing nodes is determined.

[0072] For example, the congestion index in the case of indirect connection can be determined by the following formula (5): (5) in, This represents the congestion index between indirectly connected task processing nodes u and v, where P represents the connection path between task processing nodes u and v, and task processing nodes x and y are two directly connected task processing nodes on the connection path P. This represents the congestion index between task processing node x and task processing node y.

[0073] In this way, for indirectly connected task processing nodes, the bandwidth between the two task processing nodes is determined based on the normalized raw bandwidth corresponding to the shortest connection path between them, and the congestion index between the two task processing nodes is determined based on the congestion index between each pair of directly connected task processing nodes on the connection path between them. This helps to improve the accuracy of the determined bandwidth and congestion index and enhance network affinity.

[0074] In some possible implementations, determining the bandwidth and congestion metrics between the two task processing nodes based on the connection relationship between them includes: If the connection is not established, the bandwidth and congestion metrics between the two task processing nodes are set to 0.

[0075] Here, since the two task processing nodes are not connected, there will be no communication between them, and therefore no corresponding bandwidth and congestion indicators. Thus, the bandwidth and congestion indicators between the two task processing nodes can be set to 0. This way, subsequent clustering mapping will automatically skip such pairings, avoid assigning communication-intensive subtasks to task processing nodes that cannot communicate with each other, reduce invalid waiting and timeout retries, save detection overhead, and improve the stability of task scheduling.

[0076] In some possible implementations, constructing a network affinity matrix based on bandwidth and congestion metrics between every two task processing nodes among the multiple task processing nodes to be scheduled includes: Based on the bandwidth and congestion indicators between every two task processing nodes in the multiple task processing nodes to be scheduled, the network affinity score between every two task processing nodes in the multiple task processing nodes is determined. A network affinity matrix is ​​constructed based on the network affinity score between every two task processing nodes among the multiple task processing nodes.

[0077] Optionally, the bandwidth and congestion index between each pair of task processing nodes can be smoothed to obtain the smoothed bandwidth and smoothed congestion index between each pair of task processing nodes. Based on the smoothed bandwidth and smoothed congestion index between each pair of task processing nodes in the multiple task processing nodes, the network affinity score between each pair of task processing nodes in the multiple task processing nodes can be determined.

[0078] When smoothing the bandwidth between two task processing nodes, for each time step other than the first time step, the bandwidth of the current time step after smoothing is determined based on the bandwidth of that time step and the bandwidth of the previous time step. The bandwidth of the previous time step can be the bandwidth of the previous time step after smoothing.

[0079] For example, bandwidth smoothing can be performed as shown in the following formula (6): (6) in, This represents the bandwidth at the current time step after smoothing. Represents the smoothing factor. Indicates the bandwidth at the current time step. This indicates the bandwidth of the previous time step.

[0080] The specific value of the smoothing factor can be set according to the actual task scheduling needs, and there is no restriction here.

[0081] The specific method for smoothing the congestion index between two task processing nodes is similar to the method for smoothing the bandwidth between two task processing nodes, and can be referred to the description in the above embodiments, which will not be repeated here.

[0082] For example, the network affinity score between two task processing nodes can be determined by the following formula (7): (7) in, This represents the network affinity score between task processing node u and task processing node v. The network affinity score ranges from [0,1]. A higher network affinity score indicates better network quality. This represents the congestion penalty coefficient, which indicates the degree of negative impact on network quality. This represents the bandwidth after smoothing between task processing node u and task processing node v. This represents the congestion index after smoothing between task processing node u and task processing node v.

[0083] The specific value of the congestion penalty coefficient can be set according to the actual task scheduling needs, and there is no restriction here.

[0084] The network affinity score between each pair of task processing nodes is calculated. By splicing them together, you can generate The network affinity matrix is ​​Q, and A is the number of task processing nodes.

[0085] S102: Based on the inter-task communication matrix and the network affinity matrix, cluster the multiple subtasks and the multiple task processing nodes respectively, and map the clustered task cluster set and node cluster set to determine the initial mapping result between the multiple subtasks and the multiple task processing nodes.

[0086] In this step, the inter-task communication matrix can be used to cluster the multiple subtasks to determine a task cluster set. The network affinity matrix can be used to cluster the multiple task processing nodes to determine a node cluster set. The task cluster set and the node cluster set are mapped to determine the initial mapping result between the multiple subtasks and the multiple task processing nodes. This enables rapid coarse tuning, provides a fast and high-quality initial solution in large-scale scenarios, and generates an initial allocation scheme that meets resource constraints, so that subsequent fine tuning can be optimized in a smaller and more meaningful search space.

[0087] Specifically, based on the inter-task communication matrix and the network affinity matrix, clustering is performed on the multiple subtasks and the multiple task processing nodes, and the clustered task cluster set and node cluster set are mapped to determine the initial mapping result between the multiple subtasks and the multiple task processing nodes, including: Based on the inter-task communication matrix, the multiple subtasks are clustered to obtain a set of task clusters, and based on the network affinity matrix, the multiple task processing nodes are clustered to obtain a set of node clusters. Each task cluster in the set of task clusters includes at least one subtask, and each node cluster in the set of node clusters includes at least one task processing node. Based on the resource requirements of each subtask and the resource supply information of each task processing node, the task cluster set and the node cluster set are mapped to determine the initial mapping result between the multiple subtasks and the multiple task processing nodes.

[0088] In the above steps, the inter-task communication matrix can be used to cluster the multiple subtasks to determine a task cluster set, and the network affinity matrix can be used to cluster the multiple task processing nodes to determine a node cluster set. Based on the judgment that the resource supply information of each task processing node can meet the resource demand information of each subtask, the task cluster set and the node cluster set are mapped to determine the mapping result between the multiple subtasks and the multiple task processing nodes.

[0089] The resource demand information and the resource supply information include, but are not limited to, the number of GPUs, video memory, specific accelerator card types, and memory.

[0090] In some possible implementations, the step of clustering the multiple subtasks based on the inter-task communication matrix to obtain a task cluster set includes: Based on the inter-task communication matrix, spectral clustering is performed on the multiple subtasks to obtain a task spectral clustering matrix. Normalize each row of the task spectrum clustering matrix to obtain a normalized task spectrum clustering matrix; K-means clustering is performed on the normalized task spectrum clustering matrix to obtain multiple task clusters; Based on the target communication volume of each task cluster, determine the communication index of each task cluster; The multiple task clusters are sorted in descending order according to the communication index to generate the task cluster set.

[0091] In the above steps, the inter-task communication matrix can be used to perform spectral clustering on the multiple subtasks to obtain a task spectral clustering matrix. Spectral clustering is a graph-based clustering algorithm that mainly utilizes the similarity between data points to construct a graph structure, and then performs clustering based on the eigenvectors of the Laplacian matrix.

[0092] In some possible implementations, the step of performing spectral clustering on the plurality of subtasks based on the inter-task communication matrix to obtain a task spectral clustering matrix includes: The inter-task communication matrix is ​​used as the weight matrix for spectral clustering, and each element in the inter-task communication matrix is ​​used to indicate the weight between every two graph nodes in the undirected connected graph. Based on the inter-task communication matrix, determine the degree matrix of the undirected connected graph; Based on the degree matrix and the inter-task communication matrix, a symmetric normalized Laplace matrix is ​​determined. Calculate multiple eigenvalues ​​of the symmetric normalized Laplace matrix, and select the target eigenvalues ​​with the smallest predetermined number from the multiple eigenvalues; The task spectrum clustering matrix is ​​generated based on the feature vectors corresponding to the target feature values.

[0093] The process of determining the spectral clustering matrix will be explained next.

[0094] Specifically, we can define an undirected connected graph G = (V, E), where V represents the set of vertices and E represents the set of edges. Then, we determine a weight matrix W, where each element of the weight matrix W indicates the weight between any two nodes in the undirected connected graph G. For example... W represents the weight between graph node a and graph node b in an undirected connected graph G. Based on the weight matrix W, the degree matrix D of the undirected connected graph G can be determined. In the degree matrix D, only the elements on the diagonal from the top left to the bottom right have specific values; all other elements are 0.

[0095] For example, the degree matrix can be determined by the following formula (8): (8) in, This represents the element at the a-th row and a-th column in the degree matrix D. Let represent the weight between graph node a and graph node b in an undirected connected graph G, and n represent the number of graph nodes in the undirected connected graph G.

[0096] After determining the degree matrix D, the symmetric normalized Laplace matrix can be determined based on the degree matrix D and the weight matrix W. For example, the symmetric normalized Laplace matrix can be determined by the following formula (9): (9) in, Let L denote the symmetric normalized Laplacian matrix, D denote the degree matrix, and L denote the unnormalized Laplacian matrix. The unnormalized Laplacian matrix is ​​determined based on the difference between the degree matrix and the weight matrix (L=DW).

[0097] In this way, the orthogonality of eigenvectors can be guaranteed by using a symmetric normalized Laplace matrix.

[0098] Determining the symmetric normalized Laplace matrix Then, calculate the symmetric normalized Laplace matrix. Multiple eigenvalues.

[0099] For example, the eigenvalues ​​of a symmetric normalized Laplace matrix can be determined by the following formula (10): (10) in, Represents the symmetric normalized Laplace matrix. Let x represent the eigenvalues ​​of the symmetric normalized Laplacian matrix, and let x represent the eigenvectors corresponding to the eigenvalues.

[0100] Determining the symmetric normalized Laplace matrix After obtaining multiple feature values, the target feature value with the smallest preset number k is selected from the multiple feature values.

[0101] The preset quantity k is based on the symmetric normalized Laplace matrix. The characteristic gap is determined, and the characteristic gap is used to represent the symmetric normalized Laplacian matrix. The difference between two adjacent feature values ​​is calculated after sorting multiple feature values ​​in ascending order of numerical value. Specifically, the largest feature gap can be determined from multiple feature gaps. This can be understood as the difference between two feature values. The feature value that appears earlier in the ascending order of numerical value is then identified, and its corresponding order value is set as a preset quantity k.

[0102] For example, see also Figure 5 This is a schematic diagram illustrating the distribution of eigenvalues ​​of a symmetric normalized Laplace matrix, as shown in an exemplary embodiment of this application. Figure 5 As shown in the figure, the horizontal axis represents the index of the feature value, and the vertical axis represents the specific value of the feature value. In this example, the maximum feature gap is the feature gap between the 5th and 6th feature values. Therefore, the preset quantity k=5, and the first 5 feature values ​​in ascending order of numerical value are used as the target feature values.

[0103] Based on the above, it can be seen that in calculating the symmetric normalized Laplacian matrix... When there are multiple eigenvalues, the eigenvector corresponding to each eigenvalue can also be calculated (refer to formula (10)). Each eigenvector is a... A dimensional vector. The feature vectors corresponding to the target feature values ​​are used as column vectors to form... A spectral clustering matrix U of dimension, where k is the preset number.

[0104] When determining the task spectral clustering matrix, the inter-task communication matrix is ​​used as the weight matrix for spectral clustering processing. The spectral clustering matrix U obtained according to the above example is the task spectral clustering matrix.

[0105] In this way, spectral clustering can capture the natural grouping structure in the network (such as nonlinearity and nonconvexity), adapt to dynamic and irregular traffic scenarios, and is particularly effective when there is congestion or traffic hotspots. Using spectral clustering helps to improve the clustering effect, enhance the accuracy of the task spectral clustering matrix, and spectral clustering is fast, making it suitable for online rapid scheduling decisions.

[0106] After obtaining the task spectrum clustering matrix, each row of the task spectrum clustering matrix can be normalized to obtain a normalized task spectrum clustering matrix.

[0107] For example, the elements in the normalized task spectrum clustering matrix can be represented by the following formula (11): (11) in, This represents the element at the c-th row and d-th column in the normalized task spectrum clustering matrix. This represents the element at the c-th row and d-th column in the task spectrum clustering matrix, where k represents the preset number.

[0108] Here, the task spectrum clustering matrix is The matrix is ​​dimensional, and correspondingly, the normalized task spectrum clustering matrix is ​​still . A matrix of dimension k is subjected to K-means clustering in k dimensions to obtain multiple task clusters.

[0109] For example, see also Figure 6 This is a schematic diagram illustrating a method for determining task clusters, as shown in an exemplary embodiment of this application. In this example, to more clearly demonstrate the clustering effect, the k-dimensional data is mapped to a 2-dimensional representation, such as... Figure 6 As shown in the diagram, a dot represents a subtask, and multiple task clusters can be obtained based on the distribution density of the subtasks.

[0110] After obtaining multiple task clusters, the communication index of each task cluster can be determined based on the target communication volume of each task cluster.

[0111] In some possible implementations, the target traffic includes intra-cluster traffic, which can be understood as indicating all traffic of the corresponding task cluster, including intra-cluster and extra-cluster traffic.

[0112] The step of determining the communication index for each task cluster based on the target communication volume of each task cluster includes: For each task cluster, the communication index of the task cluster is determined based on the intra-cluster communication volume of the task cluster, the target communication volume of the task cluster, and the maximum target communication volume in each task cluster.

[0113] In this step, for each task cluster, the cohesion ratio of the task cluster can be determined based on the intra-cluster communication volume and the target communication volume of the task cluster; the communication index of the task cluster can be determined based on the cohesion ratio of the task cluster, the target communication volume of the task cluster, and the maximum target communication volume in each task cluster.

[0114] Specifically, the cohesion ratio of a task cluster can be determined based on the ratio of the intra-cluster communication volume to the target communication volume of the task cluster.

[0115] For example, the cohesion ratio of a task cluster can be determined by the following formula (12): (12) in, This represents the cohesion ratio of the e-th task cluster. This represents the intra-cluster communication of the e-th task cluster. This represents the target communication volume of the e-th task cluster.

[0116] For example, the communication index of a task cluster can be determined by the following formula (13): (13) in, This represents the communication index of the e-th task cluster. This represents the cohesion ratio of the e-th task cluster. This represents the intra-cluster communication of the e-th task cluster. This represents the maximum target communication volume in each task cluster. Indicates the first communication coefficient. This represents the second communication coefficient.

[0117] The specific values ​​of the first communication coefficient and the second communication coefficient can be set according to the actual task scheduling needs, and there are no restrictions here.

[0118] After determining the communication index of each task cluster, the multiple task clusters can be sorted in descending order according to the communication index to generate the task cluster set.

[0119] For example, the set of task clusters can be represented by the following formula (14): (14) in, This represents a set of task clusters. In this example, the task cluster set includes... A cluster of tasks This represents the first task cluster sorted in descending order of communication index. This represents the second task cluster sorted in descending order of communication index. This indicates the last task cluster sorted in descending order of communication index.

[0120] In this way, using the inter-task communication matrix as input, spectral clustering is first used to obtain a task spectral clustering matrix. Then, row normalization is performed on the task spectral clustering matrix to eliminate dimensional differences. Next, K-means clustering is used to divide the subtasks into multiple task clusters. Then, based on the target communication volume of each task cluster, a communication index is determined, and the task clusters are arranged in descending order of communication index to form a set of task clusters. Thus, the task clusters with the most frequent communication are prioritized, and can be assigned to appropriate node clusters during subsequent mapping, which helps reduce cross-node traffic, lower end-to-end latency, and improve bandwidth utilization.

[0121] In some possible implementations, the clustering of the plurality of task processing nodes based on the network affinity matrix to obtain a node cluster set includes: Based on the network affinity matrix, spectral clustering is performed on the multiple task processing nodes to obtain a node spectral clustering matrix. The nodes are normalized by normalizing each row of the node spectrum clustering matrix to obtain the normalized node spectrum clustering matrix. K-means clustering is performed on the normalized node spectrum clustering matrix to obtain multiple node clusters; The node clusters are sorted in descending order according to the number of task processing nodes included in the node clusters to generate the node cluster set.

[0122] Here, the specific process of obtaining the node spectral clustering matrix based on the network affinity matrix is ​​similar to the process of obtaining the task spectral clustering matrix based on the inter-task communication matrix. The network affinity matrix is ​​used as the weight matrix for spectral clustering processing. For details, please refer to the description in the above embodiments, which will not be repeated here.

[0123] The specific process of obtaining the normalized node spectrum clustering matrix is ​​similar to that of obtaining the normalized task spectrum clustering matrix. For details, please refer to the description in the above embodiments, which will not be repeated here.

[0124] The specific process of obtaining multiple node clusters is similar to that of obtaining multiple task clusters, and can be referred to the description in the above embodiments, which will not be repeated here.

[0125] After obtaining multiple task clusters, the number of task processing nodes included in each task cluster can be determined. The multiple node clusters are then sorted in descending order according to the number of task processing nodes included in each node cluster to generate the node cluster set.

[0126] For example, the set of node clusters can be represented by the following formula (15): (15) in, This represents a set of node clusters. In this example, the set of node clusters includes... A cluster of nodes, This represents the first cluster of nodes sorted in descending order by the number of task processing nodes. This represents the second cluster of nodes sorted in descending order by the number of task processing nodes. This represents the last cluster of nodes sorted in descending order by the number of task processing nodes.

[0127] In this way, using the network affinity matrix as input, spectral clustering is first used to obtain the node spectral clustering matrix. Then, the node spectral clustering matrix is ​​row-normalized to eliminate dimensional differences. Subsequently, K-means clustering is used to divide the task processing nodes into multiple node clusters, which are then arranged in descending order of size to form a set of node clusters. As a result, the node cluster with the most task processing nodes is sorted first, and can be assigned to suitable task clusters during subsequent mapping, which helps to reduce cross-node traffic, reduce end-to-end latency, and improve bandwidth utilization.

[0128] In the above embodiments, spectral clustering is performed on the multiple subtasks based on the inter-task communication matrix, and spectral clustering is performed on the multiple task processing nodes based on the network affinity matrix. In other embodiments, other graph clustering or heuristic grouping methods can also be used to process the multiple subtasks based on the inter-task communication matrix and the multiple task processing nodes based on the network affinity matrix to obtain the corresponding matrices.

[0129] In some possible implementations, the step of mapping the task cluster set and the node cluster set based on the resource requirement information of each subtask and the resource supply information of each task processing node, and determining the mapping result between the plurality of subtasks and the plurality of task processing nodes, includes: Based on the resource requirements of each subtask and the resource supply information of each task processing node, each task cluster and each node cluster is mapped sequentially according to the arrangement order of each task cluster in the task cluster set and the arrangement order of each node cluster in the node cluster set, until the allocation of the multiple subtasks or the allocation of the multiple task processing nodes is completed, thereby obtaining the mapping result between the multiple subtasks and the multiple task processing nodes.

[0130] In the above steps, the task clusters can be sequentially assigned to the node clusters according to the arrangement order of each task cluster in the task cluster set and the arrangement order of each node cluster in the node cluster set. By checking whether the task clusters and node clusters meet the resource constraints, the task clusters can be assigned to the node clusters in sequence until the multiple sub-tasks are assigned or the multiple task processing nodes are assigned, so as to obtain the mapping result between the multiple sub-tasks and the multiple task processing nodes.

[0131] Optionally, if a matching error occurs during the mapping process, a rollback can be triggered, returning to the steps of dividing subtasks and determining task processing nodes. The inter-task communication matrix and network affinity matrix can be reconstructed and clustered to obtain a new set of task clusters and node clusters for remapping. An alarm signal can also be issued for manual verification.

[0132] In this way, prioritizing the allocation of task clusters with higher communication volume to node clusters with higher network affinity and richer resources helps to reduce communication overhead, reduce end-to-end latency, improve resource utilization, and increase the efficiency of the entire deep learning task.

[0133] In some possible implementations, the mapping of task clusters and node clusters includes: If the resource supply of the node cluster can meet the resource requirements of the task cluster, the node cluster and the task cluster are mapped. If the resource supply of the node cluster cannot meet the resource requirements of the task cluster, the task cluster is divided into multiple sub-task clusters according to the intra-cluster communication density. The sub-task clusters are then mapped sequentially according to their resource requirements in descending order.

[0134] For example, for the first task cluster in the task cluster set, it is determined whether the resource supply of the first node cluster in the node cluster set can meet the resource requirements of the first task cluster. If the resource supply of the first node cluster can meet the resource requirements of the first task cluster, the first node cluster and the first task cluster are mapped. Then, for the second task cluster in the task cluster set, it is determined whether the resource supply of the second node cluster in the node cluster set can meet the resource requirements of the second task cluster. For example, if the first task cluster includes 10 sub-tasks and the first node cluster includes 10 task processing nodes, then the first node cluster and the first task cluster can be mapped.

[0135] If the resource supply of the first node cluster cannot meet the resource requirements of the first task cluster, the second task cluster is divided into multiple sub-task clusters according to the intra-cluster communication density of the second task cluster. These sub-task clusters are then sorted in descending order of their resource requirements. The first sub-task cluster and the first node cluster are mapped together. Next, resource constraint checks are performed on the second sub-task cluster and the second node cluster to attempt mapping. For example, if the first task cluster includes 10 sub-tasks and the first node cluster includes 6 task processing nodes, the first node cluster can be split into two sub-task clusters: the first sub-task cluster includes 6 sub-tasks, and the second sub-task cluster includes 4 sub-tasks. The first sub-task cluster is then mapped to the first node cluster, and the second sub-task cluster is mapped to the second node cluster.

[0136] Optionally, during the mapping process, in addition to resource constraint judgment, it is also possible to detect whether the task processing nodes in the node cluster are suitable for executing the tasks in the task cluster, and adjust the mapping accordingly based on the detection results.

[0137] In this way, when the resource supply of the node cluster is sufficient to meet the resource requirements of the entire task cluster, mapping is performed directly. If the resource supply is insufficient, the task cluster is partially split and mapped in sequence according to the resource requirements, giving priority to the allocation of task clusters with larger communication volumes, while preserving local high affinity characteristics.

[0138] The embodiments disclosed herein can provide a high-quality global initial solution through a coarse-tuning process, significantly reducing the initial search cost and convergence time. On this basis, subsequent effective local perturbation search further reduces communication costs and the occurrence of local optima.

[0139] S103: Perform a neighborhood action on the initial mapping result to obtain an optimized mapping result. The neighborhood action is used to adjust the mapping relationship between some subtasks and task processing nodes.

[0140] In this step, neighborhood actions are performed on the initial mapping results to quickly generate candidate solutions near the initial mapping results. The mapping relationship between some subtasks and task processing nodes is adjusted without overturning the global mapping, thereby converging to a better configuration in a shorter time. This effectively reduces the impact of sporadic hotspots and idle nodes left over from coarse-grained clustering, reduces cross-link traffic, and lowers end-to-end latency.

[0141] In some possible implementations, the execution of the neighborhood action includes: For any two task clusters that have the same number of subtasks, swap the node clusters mapped to the two task clusters; Alternatively, for any two subtasks in a task cluster that are communicating, swap the task processing nodes mapped to those two subtasks. Alternatively, at least one subtask from any task cluster can be mapped to another idle node cluster.

[0142] Here, the communication of the topology structure is often regular. There are task clusters of equal size in the task cluster set, that is, there are task clusters with the same number of subtasks. Two task clusters with the same number of subtasks can be randomly selected and the node clusters mapped by the two task clusters can be swapped to achieve random and complete swapping of task clusters.

[0143] A task cluster often includes a pair of subtasks that communicate with each other. Alternatively, two subtasks that communicate with each other can be randomly selected from a task cluster, and their corresponding task processing nodes can be swapped to achieve the exchange of high-communication task pairs.

[0144] A task cluster may contain idle node clusters. A task cluster can be randomly selected, and at least one subtask from that cluster can be mapped to another idle node cluster. Optionally, all subtasks from the task cluster can be allocated to other idle node clusters, achieving random migration of task clusters. Here, if the resource supply of an idle node cluster can meet the resource requirements of the task cluster, all subtasks from that task cluster can be allocated to that idle node cluster. For example, if the task cluster includes 5 subtasks and the idle node cluster includes 5 task processing nodes, then all subtasks from the task cluster can be allocated to that idle node cluster. If the resource supply of an idle node cluster cannot meet the resource requirements of the task cluster, cross-node cluster scheduling can be performed, allocating all subtasks from the task cluster to different idle node clusters. For example, if the task cluster includes 5 subtasks, and an idle node cluster includes 4 task processing nodes, then the 4 subtasks in the task cluster can be evenly distributed to the idle node cluster with 4 task processing nodes, and the remaining 1 subtask in the task cluster can be distributed to another idle node cluster with 1 task processing node. Alternatively, some subtasks in the task cluster can be evenly distributed to other idle node clusters, realizing random migration of some task clusters.

[0145] In this way, neighborhood actions, using single-point migration, pair swapping, and task cluster migration as means, fully consider network affinity, can generate candidate solutions and significantly affect communication costs. Through local adjustments, cross-node traffic can be continuously reduced while maintaining global layout stability, thereby reducing communication overhead.

[0146] S104: Based on the comparison result between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, the initial mapping result is optimized and iterated until the iteration stopping condition is met. The optimized mapping result when the iteration stopping condition is met is determined as the target mapping result between the multiple subtasks and the multiple task processing nodes.

[0147] In this step, based on the initial mapping result, controlled local perturbation is performed, and new mappings are continuously tried until convergence is achieved, resulting in a mapping result that satisfies resource constraints and has low communication overhead. This allows for obtaining a mapping result with the lowest possible communication overhead within a limited time. In some possible implementations, the communication overhead corresponding to the initial mapping result is determined through the following steps: When scheduling tasks according to the initial mapping result, determine the bandwidth and congestion index between every two task processing nodes in the plurality of task processing nodes, and construct the network affinity matrix corresponding to the initial mapping result; Based on the inter-task communication matrix and the network affinity matrix corresponding to the initial mapping result, the communication overhead corresponding to the initial mapping result is determined.

[0148] Here, since the initial mapping result is determined, the bandwidth and congestion index between every two task processing nodes will change after the task processing nodes allocate sub-tasks according to the initial mapping result. Therefore, when re-determining the bandwidth and congestion index between every two task processing nodes among the multiple task processing nodes according to the initial mapping result, the network affinity matrix corresponding to the initial mapping result is reconstructed. The task communication volume between sub-tasks will not change. Thus, based on the inter-task communication matrix and the network affinity matrix corresponding to the initial mapping result, the communication overhead corresponding to the initial mapping result is determined.

[0149] For example, the communication overhead can be determined by the following formula (16): (16) in, This represents the communication overhead corresponding to the mapping result S. Subtasks sub-tasks Inter-task communication volume Indicates the number of subtasks. This indicates the number of subtasks divided under a data parallelism strategy. This indicates the number of subtasks divided under the pipeline parallelism strategy. This represents the number of subtasks partitioned under the tensor parallel strategy. This represents the network affinity score between task processing node u and task processing node v when scheduling tasks according to the mapping result S. In the mapping result, each task processing node is assigned a subtask.

[0150] In this way, by determining the communication overhead, subtasks with high communication demands can be assigned to task processing nodes with high bandwidth, low latency, and higher network affinity, which helps to reduce overall communication costs and reduce communication bottlenecks.

[0151] Here, the specific method for determining the communication overhead corresponding to the optimized mapping result is similar to the method for determining the communication overhead corresponding to the initial mapping result. Both involve re-determining the corresponding network affinity matrix and then combining it with the inter-task communication matrix. For details, please refer to the description in the above embodiments, which will not be repeated here.

[0152] In some possible implementations, the step of optimizing the initial mapping result based on the comparison between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, until the iteration stopping condition is met, includes: Based on the comparison between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, the acceptance probability of the optimized mapping result is determined, and the acceptance probability is used to indicate the likelihood of accepting the optimized mapping result; The accepted mapping result is used as the initial mapping result, and the neighborhood action is performed on the initial mapping result until the iteration stopping condition is met.

[0153] In the above steps, the communication overhead corresponding to the initial mapping result is compared with the communication overhead corresponding to the optimized mapping result to obtain a comparison result. Based on the comparison result, the acceptance probability of the optimized mapping result is determined. According to the acceptance probability, it is determined whether to accept the optimized mapping result. If the optimized mapping result is accepted, it is used as the new initial mapping result, and a neighborhood action is performed on the new initial mapping result to continue iterative optimization until the iteration stopping condition is met. If the optimized mapping result is not accepted, a neighborhood action is performed on the original initial mapping result to continue iterative optimization.

[0154] The neighborhood actions used in each iteration can be the same or different; there is no restriction here.

[0155] For example, if the initial mapping result obtained after coarse adjustment is S0, and the optimized mapping result after iterating on S0 is S1, the acceptance probability corresponding to S1 is executed. If S1 is accepted, S1 continues to be iterated; if S1 is not accepted, S0 continues to be iterated.

[0156] For example, after S0, S1, S2, S3, and S4, the current iteration yields S5. The acceptance probability corresponding to S5 is executed. If S5 is not accepted, the acceptance probability corresponding to S4 is executed. If S4 is also not accepted, the acceptance probability corresponding to S3 is executed, until there is an accepted mapping result or the process reverts to S0.

[0157] In this way, by determining the acceptance probability based on the two communication overheads being compared, and continuing to iterate on the accepted mapping results, the diversity of the search can be increased, the situation of getting trapped in local optima can be reduced, and thus, while maintaining scheduling stability, a better configuration can be continuously explored, premature convergence can be avoided, communication costs can be gradually reduced, node load and link utilization can tend to be balanced, and finally a higher quality and more robust target mapping result can be obtained.

[0158] In some possible implementations, determining the acceptance probability of the optimized mapping result based on a comparison between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result includes: If the communication overhead corresponding to the initial mapping result is greater than the communication overhead corresponding to the optimized mapping result, the probability of accepting the optimized mapping result is determined to be 1. If the communication overhead corresponding to the initial mapping result is less than or equal to the communication overhead corresponding to the optimized mapping result, the acceptance probability of the optimized mapping result is determined based on the communication overhead corresponding to the initial mapping result, the communication overhead corresponding to the optimized mapping result, and the overhead parameter.

[0159] In the above steps, if the communication overhead corresponding to the initial mapping result is greater than the communication overhead corresponding to the optimized mapping result, it means that the communication overhead has been reduced in this iteration, which is an optimized solution and can be directly accepted. Therefore, the acceptance probability of the optimized mapping result is 1.

[0160] If the communication overhead corresponding to the initial mapping result is less than or equal to the communication overhead corresponding to the optimized mapping result, it means that the current iteration has not reduced the communication overhead. To resolve this, the acceptance probability of the optimized mapping result can be determined based on the communication overhead corresponding to the initial mapping result, the communication overhead corresponding to the optimized mapping result, and the overhead parameter.

[0161] Optionally, simulated annealing is suitable for solving complex optimization problems with multiple local optima. Its core idea is to simulate the gradual cooling of a substance at high temperatures, allowing the system to gradually reach its lowest energy state, thus finding the global optimum. In this embodiment, simulated annealing can be used for fine-tuning. When using simulated annealing, the overhead parameter is temperature.

[0162] For example, when the communication overhead corresponding to the initial mapping result is less than or equal to the communication overhead corresponding to the optimized mapping result, the acceptance probability of the optimized mapping result can be determined by the following formula (17): (17) in, Indicates the initial mapping result The corresponding communication overhead is less than or equal to the optimized mapping result. When considering the corresponding communication overhead, optimize the mapping results. The corresponding acceptance probability, Indicates the optimized mapping result The corresponding communication overhead, Indicates the initial mapping result The corresponding communication overhead, T, represents the overhead parameter used in this iteration. The larger the overhead parameter, the higher the probability of accepting degraded solutions.

[0163] In this way, optimal solutions are always accepted, ensuring that every positive improvement is retained. Degenerate solutions are calculated and accepted, and small rebounds can be tolerated temporarily while continuing to explore better configurations. This approach can quickly lock in significant gains while avoiding premature entrapment in local optima. The iterative process maintains a balance between convergence stability and global search capability, thereby continuously reducing overall communication costs and improving the quality of mapping results.

[0164] In some possible implementations, when the initial mapping result is optimized iteratively using simulated annealing, the overhead parameter used in the first round of iteration is a preset overhead parameter. The overhead parameters used in the remaining iterations are generated by reducing the overhead parameters used in the previous iteration. Optionally, the overhead parameters used in the next iteration can be determined based on the cooling coefficient and the overhead parameters used in the previous iteration.

[0165] For example, reducing the overhead parameter can be represented by the following formula (18): (18) in, This indicates the cost parameters used in the third iteration. This represents the cooling coefficient, which ranges from [0,1]. This indicates the cost parameters used in the second iteration.

[0166] The specific value of the overhead parameter can be set according to the actual task scheduling needs, and there is no restriction here.

[0167] Here, each round of iteration includes the same number of iterations, that is, the number of iterations under each cost parameter is the same, for example, N. The specific value of the number of iterations can be set according to the actual task scheduling needs, and there is no restriction here.

[0168] In this way, the first iteration starts with a preset cost parameter to ensure that the entire optimization iteration process has a sufficient probability of accepting degenerate solutions and to maintain a wide search range. In each subsequent iteration, the cost parameter is reduced by the same amount while the number of iterations remains unchanged, gradually reducing the probability of accepting degenerate solutions, so as to shift from a broad random search to local fine optimization.

[0169] In the above embodiments, fine-tuning is performed by simulated annealing. In other embodiments, other heuristic algorithms that can approximate the target through iteration can also be used, such as particle swarm optimization, genetic algorithm, differential evolution algorithm, etc., and overhead parameters matching the algorithm used can be set.

[0170] In some possible implementations, the iteration stopping condition is that the cost parameter used is reduced to a termination cost parameter, and the optimized mapping result when the iteration stopping condition is met is the mapping result accepted in the last iteration of the current iteration in which the optimization iteration is performed using the termination cost parameter.

[0171] For example, taking simulated annealing as an example, with N iterations, the temperature decreases to the termination temperature. At that time, The mapping result received during the Nth iteration is determined as the optimized mapping result that satisfies the iteration stopping condition, which is the final target mapping result.

[0172] In this way, as the overhead parameters decrease, the algorithm gradually shifts from extensive random search to local fine-grained optimization. The optimized mapping result that meets the iteration stopping condition is determined as the target mapping result between the multiple subtasks and the multiple task processing nodes, ensuring that tasks with high communication requirements are assigned to paths with better network affinity, thereby effectively reducing the overall communication cost.

[0173] For a clearer illustration of the task scheduling process, see [link to relevant documentation]. Figure 7 This is a schematic diagram illustrating a task scheduling process as shown in an exemplary embodiment of this application. Figure 7 As shown, a task communication matrix corresponding to multiple subtasks to be scheduled and a network affinity matrix corresponding to multiple task processing nodes to be scheduled are constructed. Based on the task communication matrix and the network affinity matrix, clustering is performed on the multiple subtasks and the multiple task processing nodes respectively, and the clustered task cluster set and node cluster set are mapped to determine the initial mapping result between the multiple subtasks and the multiple task processing nodes. A neighborhood action is performed on the initial mapping result to obtain an optimized mapping result. Based on the comparison result between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, the initial mapping result is optimized iteratively until the iteration stopping condition is met. The optimized mapping result that meets the iteration stopping condition is determined as the target mapping result between the multiple subtasks and the multiple task processing nodes. The specific steps are described in the foregoing embodiment and will not be repeated here.

[0174] As the model and task scale increase, initial mapping methods such as one-time spectral clustering may not be able to reach the global optimum, especially when resource constraints require some tasks to be allocated across clusters, or when the network topology is relatively uniform, making the initial spectral clustering solution close to a random solution. In such cases, further local search optimization is needed to avoid local optima. Existing heuristic algorithms (such as simulating annealing directly on random initial solutions) often have long search times, slow convergence, and are greatly affected by the quality of the initial solution. The embodiments disclosed in this publication combine topology-aware scheduling with coarse tuning (such as spectral clustering) and fine tuning (such as simulated annealing) to control response time while ensuring scheduling quality.

[0175] To more clearly demonstrate the effect of task scheduling, please refer to the example provided. Figure 8a , Figure 8b , Figure 8c and Figure 8d . Figure 8a This is an iterative schematic diagram illustrating communication overhead as an exemplary embodiment of this application. Figure 8b This is an iterative schematic diagram illustrating another communication overhead as an exemplary embodiment of this application. Figure 8c This is an iterative schematic diagram illustrating yet another communication overhead, as shown in an exemplary embodiment of this application. Figure 8d This is an iterative schematic diagram illustrating another communication overhead as an exemplary embodiment of this application. Figures 8a-8d The diagram illustrates the iterative communication overhead under different data parallelism strategies. Figure 8a The demonstration shows a 4×4 strategy (dividing into 4 subtasks under both data parallelism and pipeline parallelism strategies). Figure 8b The demonstration shows a 4×8 strategy (dividing into 4 subtasks under a data parallel strategy and 8 subtasks under a pipelined parallel strategy). Figure 8c The demonstration shows a 4×16 strategy (divided into 4 subtasks under the data parallel strategy and 16 subtasks under the pipeline parallel strategy). Figure 8d This demonstrates an 8×16 strategy (divided into 4 subtasks under a data parallel strategy and 16 subtasks under a pipelined parallel strategy). The horizontal axis represents the number of iterations, and the vertical axis represents the communication overhead. The orange line represents the case where fine-tuning is performed based on a randomly obtained initial mapping result, and the blue line represents the case where fine-tuning is performed based on the initial mapping result obtained in this embodiment. Under each parallel strategy, the communication overhead obtained by fine-tuning based on the initial mapping result obtained in this embodiment is smaller overall and converges faster than the communication overhead obtained by fine-tuning based on a randomly obtained initial mapping result, thus helping to reduce search time. See also Figure 9 , Figure 9This diagram illustrates a comparison of communication overhead under different scheduling strategies, as shown in an exemplary embodiment of this application. Figure 9 As shown in the diagram, the horizontal axis represents the network topology, and the vertical axis represents the communication overhead. From left to right, the four topologies are Ramanujan, Fat-Tree, Dragonfly, and Jellyfish. The red line represents the communication overhead corresponding to the random mapping result, the blue line represents the communication overhead corresponding to the initial mapping result obtained after coarse adjustment, and the green line represents the communication overhead corresponding to the target mapping result obtained after coarse and fine adjustments. Under each topology, the communication overhead corresponding to the initial mapping result is less than that corresponding to the random mapping result, and the communication overhead corresponding to the target mapping result is less than that corresponding to the initial mapping result. Using this embodiment, the overall communication cost can be reduced, communication bottlenecks can be minimized, and it is compatible with different network topologies, achieving significant optimization effects.

[0176] The task scheduling method provided in this application constructs an inter-task communication matrix corresponding to multiple sub-tasks to be scheduled and a network affinity matrix corresponding to multiple task processing nodes to be scheduled. It then performs clustering processing on the multiple sub-tasks and multiple task processing nodes, and maps the clustered task cluster set and node cluster set to obtain a high-quality initial mapping result. The initial mapping result is then optimized iteratively by executing neighborhood actions and communication overhead. Through the combination of coarse scheduling and fine scheduling in two stages, the method effectively controls the response time while ensuring scheduling quality, which helps to reduce overall communication costs, reduce communication bottlenecks, be compatible with different network topologies, and improve the efficiency of the entire deep learning task.

[0177] Corresponding to the aforementioned embodiments of the task scheduling method, this application also provides embodiments of a task scheduling apparatus.

[0178] Please see Figure 10 This is a schematic diagram illustrating a task scheduling device according to an exemplary embodiment of this application. Figure 10 As shown in the figure, the task scheduling device 1000 provided in this application embodiment includes: The matrix construction module 1001 is used to construct the inter-task communication matrix corresponding to the multiple sub-tasks to be scheduled and the network affinity matrix corresponding to the multiple task processing nodes to be scheduled. The multiple sub-tasks to be scheduled are obtained by dividing the deep learning distributed task to be scheduled. The inter-task communication matrix is ​​used to indicate the task communication volume between every two sub-tasks in the multiple sub-tasks. The network affinity matrix is ​​used to indicate the bandwidth and congestion index between every two task processing nodes in the multiple task processing nodes. The clustering processing module 1002 is used to perform clustering processing on the multiple subtasks and the multiple task processing nodes based on the inter-task communication matrix and the network affinity matrix, respectively, and to map the clustered task cluster set and node cluster set to determine the initial mapping result between the multiple subtasks and the multiple task processing nodes. The action execution module 1003 is used to perform neighborhood actions on the initial mapping result to obtain an optimized mapping result. The neighborhood actions are used to adjust the mapping relationship between some subtasks and task processing nodes. The optimization iteration module 1004 is used to optimize and iterate the initial mapping result based on the comparison result between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, until the iteration stopping condition is met, and the optimized mapping result when the iteration stopping condition is met is determined as the target mapping result between the multiple subtasks and the multiple task processing nodes.

[0179] In some possible implementations, the action execution module 1003, when executing neighborhood actions, is specifically used for: For any two task clusters that have the same number of subtasks, swap the node clusters mapped to the two task clusters; Alternatively, for any two subtasks in a task cluster that are communicating, swap the task processing nodes mapped to those two subtasks. Alternatively, at least one subtask from any task cluster can be mapped to another idle node cluster.

[0180] In some possible implementations, the optimization iteration module 1004 determines the communication overhead corresponding to the initial mapping result through the following steps: When scheduling tasks according to the initial mapping result, determine the bandwidth and congestion index between every two task processing nodes in the plurality of task processing nodes, and construct the network affinity matrix corresponding to the initial mapping result; Based on the inter-task communication matrix and the network affinity matrix corresponding to the initial mapping result, the communication overhead corresponding to the initial mapping result is determined.

[0181] In some possible implementations, the optimization iteration module 1004 optimizes the initial mapping result based on a comparison between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, until the iteration stopping condition is met. Specifically, it is used to: Based on the comparison between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, the acceptance probability of the optimized mapping result is determined, and the acceptance probability is used to indicate the likelihood of accepting the optimized mapping result; The accepted mapping result is used as the initial mapping result, and the neighborhood action is performed on the initial mapping result until the iteration stopping condition is met.

[0182] In some possible implementations, when the optimization iteration module 1004 determines the acceptance probability of the optimized mapping result based on a comparison between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, it is specifically used for: If the communication overhead corresponding to the initial mapping result is greater than the communication overhead corresponding to the optimized mapping result, the probability of accepting the optimized mapping result is determined to be 1. If the communication overhead corresponding to the initial mapping result is less than or equal to the communication overhead corresponding to the optimized mapping result, the acceptance probability of the optimized mapping result is determined based on the communication overhead corresponding to the initial mapping result, the communication overhead corresponding to the optimized mapping result, and the overhead parameter.

[0183] In some possible implementations, when the initial mapping result is optimized and iterated using simulated annealing, the overhead parameter used in the first round of iteration is a preset overhead parameter, and the overhead parameters used in the remaining rounds of iteration are generated by reducing the overhead parameter used in the previous round of iteration, with each round of iteration including the same number of iterations.

[0184] In some possible implementations, the iteration stopping condition is that the cost parameter used is reduced to a termination cost parameter, and the optimized mapping result when the iteration stopping condition is met is the mapping result accepted in the last iteration of the current iteration in which the optimization iteration is performed using the termination cost parameter.

[0185] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0186] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0187] Based on the same technical concept, this application also provides a computer device 1100, referring to... Figure 11The diagram shown is a schematic representation of the structure of a computer device according to an exemplary embodiment of this application, comprising: The processor is 1110, the memory is 1120, and the bus is 1130. The memory 1120 is used to store execution instructions and includes main memory 1121 and external memory 1122. The main memory 1121, also known as internal memory, is used to temporarily store the operation data in the processor 1110 and the data exchanged with external memory 1122 such as hard disk. The processor 1110 exchanges data with external memory 1122 through main memory 1121.

[0188] In this embodiment, the memory 1120 is specifically used to store application code that executes the solution of this application, and its execution is controlled by the processor 1110. That is, when the electronic device 1100 is running, the processor 1110 communicates with the memory 1120 through the bus 1130, or the processor 1110 communicates with the memory 1120 through other means, so that the processor 1110 executes the application code stored in the memory 1120, and then executes the steps of the task scheduling method described in any of the foregoing embodiments.

[0189] The memory 1120 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0190] Processor 1110 may be an integrated circuit chip with signal processing capabilities. The aforementioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor.

[0191] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 1100. In other embodiments of this application, the electronic device 1100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0192] This disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the task scheduling method described in the above-described method embodiments. The storage medium may be a volatile or non-volatile computer-readable storage medium.

[0193] This disclosure also provides a computer program product, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the task scheduling method provided in any of the above embodiments of this disclosure. For details, please refer to the above method embodiments, which will not be repeated here.

[0194] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium, which can be a volatile or non-volatile computer-readable storage medium. In another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0195] Furthermore, embodiments of the subject matter and functional operation described in this specification can be implemented in the following ways: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. Alternatively or additionally, program instructions may be encoded on artificially generated propagation signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information and transmit it to a suitable receiving device for execution by the data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or combinations thereof.

[0196] The processing and logic flow described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform corresponding functions by operating on input data and generating output. The processing and logic flow can also be executed by dedicated logic circuitry—such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the device can also be implemented as dedicated logic circuitry.

[0197] Suitable computers for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as disks, magneto-optical disks, or optical disks, or the computer will be operatively coupled to such mass storage devices to receive data from or transfer data to them, or both. However, a computer is not required to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.

[0198] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into dedicated logic circuitry.

[0199] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.

[0200] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0201] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0202] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A task scheduling method, characterized in that, The method includes: Construct an inter-task communication matrix corresponding to multiple subtasks to be scheduled and a network affinity matrix corresponding to multiple task processing nodes to be scheduled. The multiple subtasks to be scheduled are obtained by partitioning the deep learning distributed task to be scheduled. The inter-task communication matrix is ​​used to indicate the task communication volume between every two subtasks in the multiple subtasks. The network affinity matrix is ​​used to indicate the bandwidth and congestion index between every two task processing nodes in the multiple task processing nodes. Based on the inter-task communication matrix and the network affinity matrix, clustering is performed on the multiple subtasks and the multiple task processing nodes, and the clustered task cluster set and node cluster set are mapped to determine the initial mapping result between the multiple subtasks and the multiple task processing nodes. The initial mapping result is processed by a neighborhood action to obtain an optimized mapping result. The neighborhood action is used to adjust the mapping relationship between some subtasks and task processing nodes. Based on the comparison between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, the initial mapping result is optimized iteratively until the iteration stopping condition is met. The optimized mapping result when the iteration stopping condition is met is determined as the target mapping result between the multiple subtasks and the multiple task processing nodes.

2. The method according to claim 1, characterized in that, The execution of neighborhood actions includes: For any two task clusters that have the same number of subtasks, swap the node clusters mapped to the two task clusters; Alternatively, for any two subtasks in a task cluster that are communicating, swap the task processing nodes mapped to those two subtasks. Alternatively, at least one subtask from any task cluster can be mapped to another idle node cluster.

3. The method according to claim 1, characterized in that, The communication overhead corresponding to the initial mapping result is determined by the following steps: When scheduling tasks according to the initial mapping result, determine the bandwidth and congestion index between every two task processing nodes in the plurality of task processing nodes, and construct the network affinity matrix corresponding to the initial mapping result; Based on the inter-task communication matrix and the network affinity matrix corresponding to the initial mapping result, the communication overhead corresponding to the initial mapping result is determined.

4. The method according to claim 1, characterized in that, The optimization and iteration of the initial mapping result based on the comparison between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, until the iteration stopping condition is met, includes: Based on the comparison between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, the acceptance probability of the optimized mapping result is determined, and the acceptance probability is used to indicate the likelihood of accepting the optimized mapping result; The accepted mapping result is used as the initial mapping result, and the neighborhood action is performed on the initial mapping result until the iteration stopping condition is met.

5. The method according to claim 4, characterized in that, The step of determining the acceptance probability of the optimized mapping result based on the comparison between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result includes: If the communication overhead corresponding to the initial mapping result is greater than the communication overhead corresponding to the optimized mapping result, the probability of accepting the optimized mapping result is determined to be 1. If the communication overhead corresponding to the initial mapping result is less than or equal to the communication overhead corresponding to the optimized mapping result, the acceptance probability of the optimized mapping result is determined based on the communication overhead corresponding to the initial mapping result, the communication overhead corresponding to the optimized mapping result, and the overhead parameter.

6. The method according to claim 5, characterized in that, When the initial mapping result is optimized and iterated using simulated annealing, the overhead parameter used in the first round of iteration is the preset overhead parameter, and the overhead parameter used in the remaining rounds of iteration is generated by reducing the overhead parameter used in the previous round of iteration. Each round of iteration includes the same number of iterations.

7. The method according to claim 5, characterized in that, The iteration stopping condition is that the cost parameter used is reduced to the termination cost parameter, and the optimized mapping result when the iteration stopping condition is met is the mapping result accepted in the last iteration of the current iteration when the optimization iteration is performed using the termination cost parameter.

8. A task scheduling device, characterized in that, The device includes: The matrix construction module is used to construct the inter-task communication matrix corresponding to the multiple sub-tasks to be scheduled and the network affinity matrix corresponding to the multiple task processing nodes to be scheduled. The multiple sub-tasks to be scheduled are obtained by partitioning the deep learning distributed task to be scheduled. The inter-task communication matrix is ​​used to indicate the task communication volume between every two sub-tasks in the multiple sub-tasks. The network affinity matrix is ​​used to indicate the bandwidth and congestion index between every two task processing nodes in the multiple task processing nodes. The clustering module is used to perform clustering processing on the multiple subtasks and the multiple task processing nodes based on the inter-task communication matrix and the network affinity matrix, respectively, and to map the clustered task cluster set and node cluster set to determine the initial mapping result between the multiple subtasks and the multiple task processing nodes. An action execution module is used to perform neighborhood actions on the initial mapping result to obtain an optimized mapping result. The neighborhood actions are used to adjust the mapping relationship between some subtasks and task processing nodes. The optimization iteration module is used to optimize and iterate the initial mapping result based on the comparison result between the communication overhead corresponding to the initial mapping result and the communication overhead corresponding to the optimized mapping result, until the iteration stopping condition is met, and the optimized mapping result when the iteration stopping condition is met is determined as the target mapping result between the multiple subtasks and the multiple task processing nodes.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the task scheduling method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the task scheduling method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Computing power resource partitioning method and device based on hypergraph clustering, equipment and medium

    CN120386635A

  • Intelligent instrument multi-task real-time optimization method and system based on dynamic resource scheduling

    CN120578512A