Link state sensing neural network model training flow path selection method

By dynamically adjusting path selection through in-band network telemetry and a greedy algorithm, the problem that static path selection and random hashing mechanisms cannot adapt to dynamic link states and inter-task communication contention is solved, thereby improving the communication efficiency and bandwidth utilization of distributed neural network model training.

CN120602393APending Publication Date: 2025-09-05UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510729476.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In the existing technology, static path selection strategies and random hash mechanisms cannot adapt to dynamic link states and inter-task communication contention problems, resulting in low communication efficiency.

Method used

Through in-band network telemetry technology, the link status is monitored in real time. Combined with the greedy algorithm and the shortest remaining processing time priority strategy, the path selection of training tasks is dynamically adjusted, the path with the least congestion is prioritized, and the communication contention between tasks is reduced.

Benefits of technology

It significantly improves the communication efficiency of distributed neural network model training, reduces tail latency, improves bandwidth utilization, and adapts to real-time changes in link status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602393A_ABST
    Figure CN120602393A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer networks, and discloses a link state sensing neural network model training flow path selection method, which comprises the following steps: collecting path information between each pair of hosts by sending a detection data packet carrying an in-band network telemetry field; calculating a priority index based on the number of sent bytes of the training task, wherein the priority index is determined by the proportion of the number of sent bytes of the training task to the total number of bytes needing to be sent and the number of sent bytes; a greedy algorithm is adopted, the training tasks serve as current training tasks according to priority indexes from high to low, paths with the lowest congestion degree are preferentially allocated to the current training tasks, and if the paths with the lowest congestion degree are occupied, suboptimal paths are allocated to the current training tasks; the congestion degree of the path is evaluated through the link utilization rate of each link in the path; according to the method, the link state and the task characteristics are monitored in real time, the path selection strategy is dynamically adjusted, and the communication efficiency of distributed neural network model training is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer networks, and in particular to a link state-aware neural network model training traffic path selection method. Background Art

[0002] In recent years, with the continuous expansion of neural network model training and the prevalence of distributed computing, network transmission has become an increasingly significant bottleneck limiting training efficiency. Most existing multi-tenant clusters attempt to reduce resource contention (including communication contention) and improve the efficiency of large-scale neural network training by placing training tasks on different hosts through a jobscheduler.

[0003] The core goal of a task scheduler is to rationally allocate computing resources to optimize the execution of neural network model training tasks. Most existing task scheduler communications (including inter-host network links and PCIe / NVLinks within hosts) are treated as black boxes, focusing on the allocation of GPU computing resources and the optimization of computing contention. Although a few task schedulers attempt to consider communication contention, due to the unpredictability of training tasks in a cluster (such as training task scale, start / end time, etc.), task schedulers are usually unable to completely avoid or eliminate communication contention. Therefore, existing task scheduling schemes focus more on the scheduling of computing resources, while the communication contention problem is often ignored or limited to local optimization.

[0004] Because task scheduling cannot completely solve the communication contention problem between training tasks, communication schedulers have emerged. Communication schedulers mainly solve the communication contention problem between multiple flows and training tasks by optimizing the scheduling of network flows. Based on their application scope and function, communication schedulers can be divided into the following three categories:

[0005] General co-flow schedulers: General co-flow schedulers primarily address the mathematical aspects of multi-flow scheduling, including path selection and priority assignment. Their goal is to optimize the average completion time of training tasks or minimize training task latency. These schedulers focus on solving general data flow scheduling problems and do not delve into how conflicts and competition between different flows affect end-to-end performance or iteration time in model training tasks. Because they do not optimize for the unique traffic characteristics of training tasks (such as bursty traffic and large-scale data transmission), they struggle to achieve optimal performance in model training scenarios.

[0006] Intra-job communication schedulers: Intra-job communication schedulers aim to optimize the scheduling of multiple flows within a single model training task, typically adjusting flow priorities or selecting different paths based on traffic patterns. These schedulers optimize communication by identifying the characteristics and priorities of different flows within a training task, but they often ignore the communication contention issues between training tasks. As a result, their performance can be significantly impacted when multiple training tasks are executed in parallel, especially when competition for network resources is intense.

[0007] Inter-job communication schedulers: Inter-job communication schedulers focus on solving the communication contention problem between different model training tasks. These schedulers usually try to reduce the contention between training tasks by predicting the traffic patterns of training tasks and using time offset or other scheduling strategies.

[0008] In summary, most current path selection-based communication schedulers focus primarily on scheduling and optimizing flows within training tasks. They attempt to reduce latency and improve bandwidth utilization by selecting the optimal path for each flow within a training task or adjusting flow priorities. However, these approaches focus on optimizing communication within a single training task and do not fully consider communication contention between multiple training tasks. To address this issue, future research needs to focus more on communication contention between training tasks and explore dynamic path selection strategies based on link state awareness. Summary of the Invention

[0009] To address the above technical problems, the present invention provides a link-state-aware neural network model training traffic path selection method, aiming to address the problem in existing technologies where static path selection strategies and random hashing mechanisms are unable to adapt to dynamic link states and inter-task communication contention. This method regularly monitors link states using in-band network telemetry (INT) information, while also monitoring changes in model training tasks. Priority indicators are defined based on the real-time communication volume of training tasks, and the current optimal path is always allocated based on the priority. The path selection strategy is dynamically adjusted throughout the entire process, significantly improving the communication efficiency of distributed neural network model training.

[0010] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0011] The present invention provides a link state-aware neural network model training traffic path selection method, comprising:

[0012] Step 1: Collect path information between each pair of hosts by sending probe packets carrying in-band network telemetry fields. The path information includes the link utilization of each link in the path.

[0013] Step 2: Calculate the priority index based on the number of bytes sent for the training task: the priority index is determined by the ratio of the number of bytes sent for the training task to the total number of bytes to be sent and the number of bytes sent;

[0014] Step 3: Using a greedy algorithm, the training tasks are selected as the current training tasks in descending order of priority. The path with the least congestion is assigned to the current training task first. If the path with the least congestion is occupied, the suboptimal path is assigned to the current training task. The congestion level of the path is evaluated by the link utilization of each link in the path.

[0015] When a new training task is added or a training task is completed, steps 1 to 3 are re-executed to update the path allocation of the training task.

[0016] In one embodiment, the in-band network telemetry fields include switch egress port, queue depth, timestamp, and link utilization;

[0017] The switch egress port represents the port number from which the detection data packet leaves the switch, and is used to locate the congested egress link;

[0018] The queue depth represents the current queue depth of the switch egress port;

[0019] The timestamp indicates the time when the detection data packet passes through the switch;

[0020] The link utilization rate indicates the utilization rate of the link where the detection data packet is located.

[0021] In one embodiment, the priority indicator is determined by the ratio of the number of bytes sent to the total number of bytes to be sent and the number of bytes sent for the training task, specifically including:

[0022]

[0023] P i is the priority index of the i-th training task, byte ratio The number of bytes sent for the training task sent The total number of bytes to be sent total ratio; n is the total number of training tasks.

[0024] Compared with the prior art, the beneficial technical effects of the present invention are:

[0025] 1. This invention uses in-band network telemetry (INT) technology to collect fine-grained link status information such as link utilization, queue depth, and hop-by-hop delay in real time, providing an accurate decision-making basis for path selection and avoiding the problem of unreasonable path allocation caused by insufficient information in traditional methods.

[0026] 2. The present invention combines the shortest remaining processing time first (SRPT) scheduling strategy and the task priority sorting mechanism to give priority to tasks with less remaining communication volume, effectively reducing tail delay and improving overall training efficiency.

[0027] 3. The present invention adopts a greedy algorithm to dynamically allocate paths, ensuring that high-priority tasks give priority to the optimal path, while allocating suboptimal paths to low-priority tasks, alleviating the communication contention problem between tasks and improving bandwidth utilization.

[0028] 4. The present invention can dynamically adjust path allocation when a task is completed or a new task arrives, adapting to real-time changes in link status and further optimizing the completion time of training tasks.

[0029] 5. The present invention solves the combinatorial optimization problem of path selection in large-scale clusters through a heuristic algorithm, while ensuring the solution efficiency and significantly reducing the impact of communication contention on training performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is a flow chart of a method in an embodiment of the present invention;

[0031] Figure 2 A schematic diagram of inter-task communication contention in an embodiment of the present invention;

[0032] Figure 3 Schematic diagram of the algorithm principle in an embodiment of the present invention;

[0033] Figure 4 Schematic diagram of multi-task path allocation in an embodiment of the present invention. DETAILED DESCRIPTION

[0034] A preferred embodiment of the present invention will be described in detail below with reference to the accompanying drawings.

[0035] During the training of a neural network model, frequent data exchange and communication are usually required, and a large number of network flows are generated. Computing nodes usually share network resources, so how to reasonably schedule these network flows has become an important issue in optimizing the efficiency of distributed training. Some studies use static path selection strategies or use the equivalent multi-path routing (ECMP) hash mechanism to randomly select paths without considering the characteristics of model training and the continuous changes in link status during training. This unreasonable path allocation will cause communication contention problems and ultimately lead to a decrease in training speed. In response to these shortcomings, the present invention provides a link state-aware neural network model training traffic path selection method, which realizes dynamic adjustment of the path by real-time monitoring of changes in link status and model training tasks, solves the problem that unreasonable path allocation will cause communication contention and ultimately lead to a decrease in training speed, and can achieve the effect of improving the communication efficiency of distributed training of neural network models. In summary, the link state-aware neural network model training traffic path selection method provided by the present invention can make more full use of the bandwidth resources of the intelligent computing center, while shortening the training time of the neural network model.

[0036] Problem definition: Communication contention in this paper refers to the phenomenon where multiple data flows or training tasks compete for the same network link or shared resources (such as switch ports, queue bandwidth, etc.), resulting in increased transmission delays for some flows, decreased throughput, or uneven resource utilization. Its core manifestations are:

[0037] (1) Resource conflict: When multiple flows pass through the same bottleneck link, insufficient bandwidth allocation causes queuing delays or packet loss.

[0038] (2) Load imbalance: Some links are overloaded while others are idle, resulting in reduced overall bandwidth utilization.

[0039] (3) Increased tail latency: Contention causes key flows (such as gradient synchronization flows) to take longer to complete, slowing down the efficiency of distributed training.

[0040] Since the communication mode of a single training task is usually fixed and regular, the communication contention within the training task can be alleviated by scheduling within the training task (such as gradient fusion, pipeline optimization); while when multiple tasks are concurrent, the data streams of different training tasks collide randomly on the shared link, resulting in sudden congestion, which cannot be solved by intra-task optimization. Therefore, the present invention mainly considers alleviating inter-task communication contention through path selection strategies, and ignores intra-task contention. Figure 2 As shown in Figure 3, when different tasks select the same bottleneck link, communication contention occurs, while other links may be idle or have low bandwidth utilization.

[0041] Problem modeling: In order to quantify the training overhead of the analysis task under different paths, it is necessary to perform mathematical modeling on the whole. Since model training is usually periodic and iterative, this invention only discusses the optimization problem within one iteration, and does not consider the resource scheduling problem across iterations. First, the symbols used are explained. Generally, four types of sets in a multi-tenant cluster are considered: computing node set, training task set, link set, and switch set. Computing node set W = {w1, w2, ..., w m} consists of m computing nodes, and the training task set J = {j1, j2, ..., j n} consists of n training tasks, j n is the nth training task, each task j i Need to execute k i subtasks.

[0042] Since the placement and allocation of tasks are not considered in this invention, it is assumed that for each training task j i , its placement and number of subtasks k i All are known conditions. Each device node (including computing nodes and switch nodes) is connected by a link. Because each subtask needs to be assigned to a computing node for training, a binary variable is used. As a 0-1 binary variable, it represents the training task j i Is subtask k of the computation node w l implement, Represents task j i Subtask k is on computing node w l Training time on training task j i The completion time is C i To express, since in the process of parallel computing, the completion time of a task depends on the slowest completion time of all subtasks, C i can be Calculated. represents the path selected by subtask k for communication, and each path can be represented as a set of links it passes through.

[0043] Based on the above system description, this problem can be modeled as follows. First, the overall optimization goal is to minimize the average task completion time of the cluster, which can be defined as:

[0044]

[0045] Constraints: There must be at least two or more tasks being trained simultaneously in the cluster:

[0046] n≥2.

[0047] For each subtask k, it must be assigned to and can only be assigned to one computing node for training:

[0048]

[0049] In order to make the workload of the task not exceed the load capacity of the computing node and avoid resource contention problems within the task, for each computing node w l , the number of subtasks assigned cannot exceed 1, expressed as:

[0050]

[0051] Since the number of subtasks trained on each computing node cannot exceed 1, and the present invention is aimed at distributed model training, for each task j i , the number of its subtasks k i The value cannot be less than 2 and cannot exceed the total number of computing nodes in the cluster, m:

[0052] 2≤k i ≤m.

[0053] For each task j i , its completion time C i is the maximum value of its subtask training time:

[0054]

[0055] Problem Solving: In distributed training, each compute node or device (such as a GPU) should have the most balanced computation time possible when executing training tasks. However, due to factors such as hardware differences, network congestion, and resource competition, subtasks on some compute nodes may take longer to complete than those on other compute nodes. These long delays can affect the entire training process because training must wait for the slowest subtask to complete before continuing. The maximum value of the subtask completion time across multiple compute nodes is called tail latency.

[0056] In large-scale distributed training, communication bottlenecks are often more significant than computing bottlenecks. Therefore, the present invention only focuses on the communication part of distributed training, ignoring the optimization problem of the computing time of the task. The overall optimization goal is to reduce the communication time, especially to reduce the communication time in the tail delay. For the sake of rigor, it is assumed that the computing nodes use a data parallel method to distribute the training of each task in the task set J. The training task is divided into multiple small blocks of the same size, each block contains a subset of the data set; these subsets are then assigned to different computing nodes, and the same computing operations are performed in parallel; each computing node processes its assigned data subset and calculates the corresponding gradient update. Therefore, for the same task, the difference in the computing time of each subtask can be ignored, so when comparing the training time of each subtask k, When , only the communication time needs to be compared.

[0057] Furthermore, in large-scale distributed training, communication traffic often peaks, and network bandwidth cannot meet the parallel data transmission needs of all compute nodes, which exacerbates network contention between tasks. In particular, when communication flows from multiple tasks are transmitted simultaneously within the same time period, the slowest flow becomes the bottleneck for the entire training process. Therefore, optimizing the communication time within the tail latency—that is, reducing the impact of communication contention on tail latency—is key to improving the efficiency of distributed training. The aforementioned optimization problem is thus completely transformed into optimizing the communication time within the tail latency through path selection, that is, optimizing the communication time of the slowest flow within a task.

[0058] As can be seen, the path selection problem for the slowest flow is a combinatorial optimization problem with a vast combinatorial space. Especially for large-scale clusters and large-model training scenarios (e.g., with thousands of task communication flows and multiple paths), the search space for the problem grows exponentially. This problem is a typical non-deterministic polynomial (NP-hard) problem, and the optimal solution cannot usually be found within polynomial time. Furthermore, nonlinear objective function optimization requires calculating the maximum communication delay for each task, making the time complexity of the optimal solution very high. Furthermore, the path must be selected in advance before each data exchange. Therefore, to obtain a good solution within an acceptable timeframe, the use of heuristic algorithms, as an approximate solution based on experience and rules, is a viable solution. By performing local optimization and exploration within the search space, the quality of the solution is gradually improved. Compared to exact solution methods, heuristic algorithms have the advantage of being able to find suboptimal solutions within a limited timeframe, thus providing a viable solution. In dynamic path optimization problems, the problem scale increases with the number of computing nodes and tasks. The present invention utilizes heuristic algorithm techniques such as parallel computing, heuristic rule design, and local search to find a good solution within an acceptable timeframe, effectively addressing large-scale problems while maintaining high solution efficiency. This approach can provide a feasible solution to the dynamic path optimization problem, while allowing dynamic adjustments to achieve better results when tasks are completed or new tasks arrive.

[0059] Algorithm design: such as Figure 3 As shown, the proposed method for selecting traffic paths using a link-state-aware neural network model includes three steps: collecting path information, prioritizing tasks, and dynamically allocating paths. The proposed method will be described in detail below.

[0060] like Figure 1 As shown, the present invention provides a link state-aware neural network model training traffic path selection method, comprising the following steps:

[0061] S1 collects path information between each pair of hosts by sending probe packets carrying in-band network telemetry fields. The path information includes information such as the link utilization of each link in the path;

[0062] S2, calculating a priority index based on the number of bytes sent for the training task: the priority index is determined by the ratio of the number of bytes sent for the training task to the total number of bytes to be sent and the number of bytes sent;

[0063] S3 uses a greedy algorithm to select training tasks as the current training task in descending order of priority. The path with the least congestion is assigned to the current training task first. If the path with the least congestion is occupied, the suboptimal path is assigned to the current training task. The congestion level of the path is evaluated by the link utilization of each link in the path.

[0064] S4: When a new training task is added or a training task is completed, steps S1 to S3 are re-executed to update the path allocation.

[0065] In one embodiment, the in-band network telemetry fields include switch egress port, queue depth, timestamp, and link utilization;

[0066] The switch egress port represents the port number from which the detection data packet leaves the switch, and is used to locate the congested egress link;

[0067] The queue depth represents the current queue depth of the switch egress port;

[0068] The timestamp indicates the time when the detection data packet passes through the switch;

[0069] The link utilization rate indicates the utilization rate of the link where the detection data packet is located.

[0070] Specifically, path information collection involves sending probe packets to detect the current level of congestion within the network link. This information is collected between each pair of hosts by sending probe packets. To achieve this, probe packets with different source ports can be sent until all candidate paths are reached. In this algorithm, in-band network telemetry (INT) technology is used to insert per-hop information into the probe packets.

[0071] In-band network telemetry is a type of network bandwidth telemetry technology that embeds metadata (such as switch queue status and timestamps) in the packet header, allowing the packet to carry link status information during transmission. The receiving end (or intermediate node) ultimately feeds back to the sending end. This invention mainly monitors the following indicators through INT technology:

[0072] Link utilization: Directly obtain the instantaneous bandwidth utilization of the switch port through INT technology.

[0073] Queue depth: Real-time monitoring of the backlog of data (bytes or number of packets) in the switch egress queue.

[0074] Hop-by-hop delay: Calculates the delay of each hop the packet takes using the INT timestamp.

[0075] The INT fields contained in the probe packet are shown in Table 1. The switch egress port indicates the port number from which the probe packet leaves the switch and is used to locate congested egress links. The queue depth indicates the current queue depth of the egress port. The timestamp indicates the timestamp (nanosecond accuracy) when the packet passes through the switch and is used to calculate hop-by-hop delay. The link utilization indicates the current link utilization. This information is primarily used in the present invention to determine the level of link congestion.

[0076] Table 1 In-band network telemetry fields

[0077] … Switch outbound port queue depth Timestamp Link utilization …

[0078] Since there may be contention for the same path between different training tasks, the present invention considers setting priorities for tasks and sorting them according to priority, so that training tasks with higher priorities can select paths first. To address this problem, the present invention refers to the Shortest Remaining Processing Time (SRPT) scheduling strategy. The SRPT scheduling strategy is often used to optimize the processing order of tasks or streams. Its core idea is to give priority to tasks with the shortest remaining processing time. This allows short tasks to be completed faster, while long tasks will be postponed. The design goal of the SRPT scheduling strategy is usually to minimize the average waiting time or average completion time, so it can be used in cluster traffic scheduling, especially when multiple network flows share bandwidth. By giving priority to traffic with less remaining transmission data, congestion can be reduced and bandwidth utilization can be improved. For traffic with large amounts of data (such as gradient synchronization in distributed model training), the SRPT scheduling strategy can also be completed by accelerating small traffic to avoid congestion.

[0079] In one embodiment, the priority indicator is determined by the ratio of the number of bytes sent to the total number of bytes to be sent and the number of bytes sent for the training task, specifically including:

[0080]

[0081] P i is the priority index of the i-th training task, byte ratio The number of bytes sent for the training task sent The total number of bytes to be sent total ratio; n is the total number of training tasks.

[0082] It can be seen and byte sent The concept of "remaining processing time" can be quantified from two different quantitative dimensions, and considering these two indicators at the same time can effectively avoid the situation where priorities are difficult to sort when they are repeated.

[0083] Dynamic path allocation: This method draws on the concept of a greedy algorithm. For the slowest flow in each task in a multi-tenant cluster, it selects the least congested path among all currently available paths and greedily finds the optimal solution for the optimization objective for each task. The specific steps are as follows:

[0084] (1) Task priority sorting: Calculate the priority index P based on the number of bytes sent by the task i ;

[0085] (2) Path allocation: Assign paths to tasks in descending order of priority. Select the least congested available path for the current task. If the path is already occupied, select the suboptimal path.

[0086] (3) Dynamic adjustment: When a new task arrives or a task is completed, the priority is recalculated and the path allocation is refreshed.

[0087] The following combination Figure 4 , taking an actual multi-tenant cluster as an example to further explain the technical solution of the present invention. Figure 4 The cluster includes 16 servers as computing nodes, which are connected by a spine-leaf network architecture. Each node participates in at most one model training task.

[0088] Assume that 4 different training tasks arrive at the same time, and according to the formula Calculate their priority indices, P1, P2, P3, and P4, respectively. Assume here that the calculated priority indices are ranked from highest to lowest: Training Task 2 > Training Task 1 > Training Task 3 > Training Task 4. Therefore, Training Task 2 prioritizes paths and sends probe packets carrying the INT field to the cluster, covering all candidate paths through different source ports. The switches fill in the INT field as the probe packets pass through. The receiver aggregates the INT field data for all paths and generates a global view, ultimately selecting idle paths for all nodes in Training Task 2. Subsequently, Training Tasks 1 and 3 select idle paths in the same order.

[0089] However, by the time training task 4 makes its decision, there are no idle paths in the cluster. Therefore, according to the greedy strategy, it must choose the least congested path. To avoid link contention with the higher-priority task, training task 4 is assigned to share the link with the lower-priority task 3. This shows that while this algorithm cannot completely eliminate communication contention, it can effectively mitigate the degree of contention.

[0090] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times. The order of execution of these steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the steps or stages in other steps.

[0091] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0092] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. It is intended that all variations within the meaning and range of equivalents of the claims be embraced herein, and any reference signs in the claims should not be construed as limiting the claims to which they relate.

[0093] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. A link state aware neural network model training traffic path selection method, characterized in that: include: Step 1: Collect path information between each pair of hosts by sending probe packets carrying in-band network telemetry fields. The path information includes the link utilization of each link in the path. Step 2: Calculate the priority index based on the number of bytes sent for the training task: the priority index is determined by the ratio of the number of bytes sent for the training task to the total number of bytes to be sent and the number of bytes sent; Step 3: Using a greedy algorithm, the training tasks are selected as the current training tasks in descending order of priority. The path with the least congestion is assigned to the current training task first. If the path with the least congestion is occupied, the suboptimal path is assigned to the current training task. The congestion level of a path is evaluated by the link utilization of each link in the path; When a new training task is added or a training task is completed, steps 1 to 3 are re-executed to update the path allocation of the training task.

2. The link state-aware neural network model training traffic path selection method according to claim 1, characterized in that: The in-band network telemetry fields include switch egress port, queue depth, timestamp, and link utilization; The switch egress port represents the port number from which the detection data packet leaves the switch, and is used to locate the congested egress link; The queue depth represents the current queue depth of the switch egress port; The timestamp indicates the time when the detection data packet passes through the switch; The link utilization rate indicates the utilization rate of the link where the detection data packet is located.

3. The method for selecting a link state-aware neural network model for training traffic paths according to claim 1, wherein: The priority index is determined by the ratio of the number of bytes sent to the total number of bytes to be sent and the number of bytes sent for the training task, specifically including: P i is the priority index of the i-th training task, byte ratio The number of bytes sent for the training task sent The total number of bytes to be sent total ratio; n is the total number of training tasks.

Citation Information

Cited By

  • Cooperative computing power management method and system based on cross-architecture state sensing engine

    CN120994349A

  • A method and system for computing power coordination and management based on cross-architecture state perception engine

    CN120994349B