Network closed-loop control method and device based on tail delay prediction, medium and program product

By acquiring flow identifiers and training semantic context in distributed AI training, combining telemetry data to predict tail latency indicators, selecting the optimal path, and implementing feedback control, the problem of insufficient identification of tail latency degradation risks in existing technologies is solved, achieving stable network path adjustment and improved training efficiency.

CN122316949BActive Publication Date: 2026-08-04SHANGHAI LINGANG YUANQI INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI LINGANG YUANQI INTELLIGENT TECH CO LTD
Filing Date
2026-06-02
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing network optimization schemes are unable to identify and stably control the tail latency degradation risk of aggregated communication flows in synchronous training of artificial intelligence in advance, resulting in switching lag, false triggering and path oscillation in path adjustment, which affects training efficiency.

Method used

By acquiring the flow identifier information and training semantic context of the target set communication flow, and combining the telemetry data of the current path and candidate paths, the tail delay index of the current path within the preset prediction window is predicted. Using the tail delay index of the current path as a benchmark, the tail delay index of the candidate paths and the path adjustment cost are comprehensively considered to select the optimal path for adjustment. The prediction process is updated by controlling the path to maintain or back off through feedback information.

Benefits of technology

It improves the foresight and effectiveness of path adjustment, reduces false triggering and frequent switching, and enhances the stability and training efficiency of network closed-loop control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122316949B_ABST
    Figure CN122316949B_ABST
Patent Text Reader

Abstract

This application provides a network closed-loop control method, device, medium, and program product based on tail delay prediction. The method includes: acquiring flow identification information and training semantic context of the target set communication flow during the distributed synchronous training process of artificial intelligence; determining the current path and candidate paths based on the flow identification information and collecting corresponding telemetry data; predicting the tail delay index of the current path and each candidate path within a preset prediction window by combining the training semantic context and telemetry data; evaluating the path based on the tail delay index and path adjustment cost of the candidate paths, using the current path as a comparison benchmark, and determining the target path; adjusting the target set communication flow to the target path when the preset path control conditions are met and the suppression conditions are not hit; and executing hold or back control based on the operational feedback information after path adjustment to update the tail delay prediction process. This application can identify the risk of tail delay degradation of the set communication flow in advance, reducing false triggering and path oscillation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of distributed training and communication network optimization technology for artificial intelligence, and in particular to a network closed-loop control method, device, medium and program product based on tail delay prediction. Background Technology

[0002] With the continuous growth of large-scale artificial intelligence model training tasks, distributed synchronous training has become an important way to improve model training efficiency. In distributed synchronous training, multiple computing nodes typically need to perform collective communication operations such as parameter synchronization, gradient merging, or intermediate result exchange at specific stages of each training round. Examples include AllReduce, AllGather, ReduceScatter, and Broadcast. Since synchronous training usually requires waiting for multiple participating nodes to complete their communication operations before proceeding to subsequent training steps, the transmission performance of the collective communication stream directly affects the efficiency of training iterations. Especially in large-scale graphics processing unit (GPU) clusters or intelligent computing center network environments, high-quantum latency anomalies on any path can cause some communication participants to become slow nodes, further triggering synchronization delays for the entire training task.

[0003] Existing network optimization schemes typically collect network status data such as current link latency, queue depth, packet loss rate, and explicit congestion notification marking ratio through in-band telemetry, out-of-band telemetry, and link status monitoring. These are then combined with software-defined network control to perform path adjustments based on current Quality of Service (QoS) metrics, link congestion status, or predictions of general traffic flows. However, these schemes often rely on general traffic flows, end-to-end average latency, link congestion status, or Service Level Agreement (SLA) violation probabilities as control criteria, failing to reflect the specific sensitivity of AI synchronous training ensemble communication to tail latency and synchronization windows. When the high-quantile latency of the ensemble communication flow rises rapidly within a short time window, path adjustments based solely on the current network status or a single threshold can easily lead to problems such as handover lag, false triggering, frequent handovers, or path oscillations. Furthermore, while some predictive routing schemes can perform path adjustments in advance, their control processes often lack continuous verification and feedback correction mechanisms for the path adjustment results, making it difficult to maintain stable control performance during long-term training tasks.

[0004] Therefore, how to predict the risk of tail delay degradation in advance for the aggregated communication flow in synchronous training of artificial intelligence, and achieve stable closed-loop control of network path while avoiding false triggering and path oscillation, has become an urgent technical problem to be solved. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this application provides a network closed-loop control method, device, medium, and program product based on tail delay prediction. This addresses the problem that existing network path optimization schemes struggle to identify and stably control the tail delay degradation risk of synchronous training set communication flows in advance, leading to issues such as switching lag, false triggering, path oscillation, and insufficient long-term control stability in path adjustments.

[0006] To achieve the above objectives and other advantages, some embodiments of this application provide the following aspects:

[0007] In a first aspect, some embodiments of this application provide a network closed-loop control method based on tail delay prediction, including:

[0008] Obtain the stream identifier information of the target set communication stream generated during the distributed synchronous training of artificial intelligence, as well as the training semantic context corresponding to the target set communication stream;

[0009] Based on the flow identification information, determine the current path and at least one candidate path corresponding to the target set communication flow, and collect telemetry data of the current path and each of the candidate paths;

[0010] Based on the training semantic context and the telemetry data, predict the tail latency index of the current path and each of the candidate paths within a preset prediction window;

[0011] Using the tail delay index of the current path as a comparison benchmark, the candidate paths are evaluated based on their tail delay indices and corresponding path adjustment costs, and the target path is determined from among the candidate paths.

[0012] If the tail delay index of the current path meets the preset path control condition and the preset suppression condition is not hit, the target set communication flow is adjusted from the current path to the target path.

[0013] The system acquires operational feedback information after path adjustment, and performs maintain or rollback control on the target path based on the operational feedback information, as well as updates the prediction process for the tail delay index.

[0014] Secondly, some embodiments of this application also provide an electronic device, the electronic device comprising:

[0015] One or more processors; and a memory storing computer program instructions that, when executed, cause the processors to perform the network closed-loop control method based on tail delay prediction as described above.

[0016] Thirdly, some embodiments of this application also provide a computer-readable storage medium having a computer program and / or instructions stored thereon, which, when executed by a processor, implement the network closed-loop control method based on tail delay prediction as described above.

[0017] Fourthly, some embodiments of this application also provide a computer program product, including a computer program and / or instructions, which, when executed by a processor, implement the network closed-loop control method based on tail delay prediction as described above.

[0018] Compared with existing technologies, the solution provided in this application obtains the flow identification information and training semantic context of the target set communication flow during the distributed synchronous training of artificial intelligence, and combines it with the telemetry data of the current path and candidate paths to predict the tail latency indicators of the current path and each candidate path within a preset prediction window. This changes the path control basis from "judgment of the current state" to "judgment of tail latency risk within the prediction window," thereby enabling earlier identification of tail latency deterioration risks in synchronous training set communication. Furthermore, this application uses the tail latency indicator of the current path as a comparison benchmark, and evaluates and determines the target path by combining the tail latency indicators of candidate paths and the path adjustment cost. This makes path selection no longer solely dependent on instantaneous congestion or current latency, but comprehensively considers the tail latency performance of candidate paths in subsequent communication windows and the path adjustment cost, thereby improving the foresight and effectiveness of path adjustment decisions. Simultaneously, this application only performs path adjustment when the tail latency indicator of the current path meets the preset path control conditions and does not hit the preset suppression conditions, which can reduce false triggering, frequent switching, and path oscillations caused by short-term fluctuations or single threshold triggering. After path adjustment, the target path is maintained or rolled back based on operational feedback information, and the prediction process of tail delay index is updated. This allows for timely rollback when the target path performance is substandard or the adjustment fails, shortening the impact time on the training process. Furthermore, the prediction process is continuously corrected through feedback updates, thereby improving the accuracy of tail delay prediction and the stability of network closed-loop control in long-term operation. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other implementation methods can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is one of the flowcharts of a network closed-loop control method based on tail delay prediction provided in the embodiments of this application;

[0021] Figure 2 This is a second schematic flowchart of a network closed-loop control method based on tail delay prediction provided in an embodiment of this application;

[0022] Figure 3 This is a schematic diagram of the path adjustment triggering determination process based on tail delay prediction provided in the embodiments of this application;

[0023] Figure 4 This is a schematic diagram of the cooling detection and rollback control process after path adjustment provided in the embodiments of this application;

[0024] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] The following terms are used in this document.

[0027] Tail latency refers to the high quantile latency near the tail of a communication latency distribution, used to characterize the impact of a small number of high-latency communication events on the overall communication process. For example, the 99th percentile tail latency refers to the latency value located at the 99th percentile after sorting multiple communication latency samples by numerical value within a statistical window. In distributed synchronous training of artificial intelligence, since the training task typically requires waiting for multiple participating nodes to complete their collective communication operations before proceeding to the next training phase, tail latency can be used to characterize the risk of slow nodes or synchronization waiting during synchronous training.

[0028] Path reconstruction refers to the process of reselecting, switching, or adjusting the network forwarding paths of target set communication flows. Predictive path reconstruction means performing path adjustments in advance based on tail latency prediction results before path performance degradation actually affects training synchronization, rather than reactively processing after link congestion, increased tail latency, or impaired training iteration efficiency have occurred.

[0029] False trigger suppression refers to constraining path adjustment actions by setting preset path control conditions and preset suppression conditions before performing path adjustment, in order to reduce invalid path adjustments caused by short-term fluctuations, single threshold exceeding limits, insufficient gains, or unsuitable switching during key training phases.

[0030] A cooling window is a time window set after a target set communication flow completes a path adjustment. Within this cooling window, the system can suppress non-urgent path adjustments for the same target set communication flow to avoid triggering another path adjustment due to short-term observation fluctuations after the initial adjustment, thereby reducing the risk of frequent switching and path oscillation.

[0031] Some embodiments of this application relate to a network closed-loop control method based on tail latency prediction, which can be applied to the network communication optimization process in distributed synchronous training scenarios for artificial intelligence. Exemplarily, this method can be applied to environments such as large-scale graphics processing unit clusters, intelligent computing centers, artificial intelligence training data centers, or high-performance computing clusters. In these environments, training tasks typically employ parallel training methods such as data parallelism, tensor parallelism, pipelined parallelism, or optimizer state partitioning. Multiple computing nodes or training nodes need to periodically perform aggregated communication operations such as parameter synchronization, gradient merging, or intermediate result exchange during each round of training, including full reduction communication, full collection communication, reduction distributed communication, or broadcast communication. Since distributed synchronous training typically requires waiting for multiple participating nodes to complete the corresponding aggregated communication operations before proceeding to subsequent training phases, any high quantile latency anomaly on any path can cause some training nodes to become slow nodes, further triggering synchronization waiting throughout the entire training iteration. Therefore, in distributed synchronous training scenarios for artificial intelligence, the tail latency variation of the aggregated communication flow is a better indicator of communication bottleneck risks during training than the general average latency.

[0032] In some embodiments, the method can be executed by a network control system deployed on the training cluster management side or network control side. This network control system can communicate with the training framework, ensemble communication library, training scheduling platform, network telemetry acquisition component, and path control component to obtain the training semantic context of the target ensemble communication flow, telemetry data of the current path and candidate paths, and perform path evaluation, path adjustment, post-adjustment active probing, hold or rollback control, and prediction process updates based on the tail delay prediction results. For example, the training framework can provide training phase, synchronization window identifier, or iteration information; the ensemble communication library can provide ensemble communication operation type, participating node range, or communication message size; the training scheduling platform can provide training task identifier, parallel group information, or node topology information; the network telemetry acquisition component can collect latency, queue depth, packet loss rate, explicit congestion notification marker ratio, or path stability information of the current path and candidate paths; and the path control component can perform flow table updates, path switching, or path rollback according to the path adjustment strategy issued by the network control system. Thus, the network control system can associate the business semantics during the training process with the network path state, providing a data foundation for tail delay prediction and path closed-loop control of the target ensemble communication flow.

[0033] In some embodiments, the network control system can be deployed in a software-defined network controller, a training network controller, a cluster management node, or a standalone predictive control node, and its specific deployment does not constitute a limitation on the embodiments of this application. For example, in a training cluster using software-defined network control, the network control system can determine its current path and at least one candidate path based on the flow identification information of the target set communication flow, and issue path control rules to switches, routers, or other forwarding devices through the network controller to adjust the target set communication flow from the current path to the target path. After the path adjustment is completed, the network control system can enter a cooling window and perform active probing on the target path according to a preset probing cycle within the cooling window; if the probing results and measured transmission performance indicate that the target path is stable, the target set communication flow is kept to be transmitted through the target path; if there are continuous probing failures, measured latency exceeding the threshold, or packet loss rate exceeding the threshold, the target set communication flow is rolled back to the current path before adjustment, and the running feedback information, rollback reason, or prediction error information is used to update the tail latency prediction process. In this way, the embodiments of this application can perform path control before tail latency degradation actually affects training synchronization, and form a closed-loop network control through adjusted feedback verification and prediction update.

[0034] Reference Figure 1 , Figure 2 As shown, the method may include the following steps:

[0035] Step S1: Obtain the stream identifier information of the target set communication stream generated during the distributed synchronous training of artificial intelligence, as well as the training semantic context corresponding to the target set communication stream.

[0036] In one specific embodiment, the network control system can communicate with the training framework, ensemble communication library, and training scheduling platform used by the AI ​​distributed synchronous training task to obtain the ensemble communication streams generated during training and their corresponding training semantic context. The target ensemble communication stream can be a communication stream generated by ensemble communication operations during AI distributed synchronous training, such as a communication stream used for gradient synchronization, parameter synchronization, or intermediate result exchange. The network control system can identify the target ensemble communication stream currently requiring tail delay prediction and path control based on ensemble communication library call information, training framework events, traffic fingerprint information, or communication stream tags issued by the training scheduling platform, and obtain the stream identification information of the target ensemble communication stream. The stream identification information may include one or more of the following: training task identifier, communication stream quintuple information, parallel group identifier, ensemble communication operation identifier, participating node identifier, or stream tag, used to distinguish communication streams generated by different training tasks, different parallel groups, or different ensemble communication operations on the network side.

[0037] In this embodiment, the training semantic context may include: the current training stage of the distributed synchronous training process of artificial intelligence, the communication primitive type corresponding to the target set communication stream, the communication message size corresponding to the target set communication stream, and the synchronization window identifier associated with the target set communication stream. Specifically, the training stage characterizes the execution stage of the training task when the target set communication stream is generated, such as the forward computation stage, backpropagation stage, gradient synchronization stage, or parameter update stage; the communication primitive type characterizes the type of set communication operation that generates the target set communication stream, such as full reduction communication, full collection communication, reduction distributed communication, or broadcast communication; the communication message size characterizes the amount of data transmitted in this set communication operation; and the synchronization window identifier characterizes the training synchronization timing position to which the target set communication stream belongs, such as the synchronization window number in the current training iteration or the synchronization round identifier corresponding to the current parallel group.

[0038] For example, in a large-scale distributed synchronous training scenario, when the data parallel group enters the gradient synchronization phase, the ensemble communication library initiates a full reduction communication operation to synchronize the gradient data of multiple training nodes. At this time, the network control system can identify the communication stream generated by this full reduction communication operation as the target ensemble communication stream and obtain its stream identifier information. Simultaneously, the network control system can determine the gradient synchronization phase as the current training phase, the full reduction communication as the communication primitive type, the number of gradient data bytes corresponding to this full reduction communication as the communication message size, and the gradient synchronization window number in the current training iteration as the synchronization window identifier. By binding the above training semantic context with the target ensemble communication stream, subsequent tail latency prediction can be performed by combining the training phase, communication operation type, communication load scale, and synchronization timing position of the ensemble communication stream to more accurately predict the tail latency indicators of the current path and candidate paths within the preset prediction window.

[0039] Step S2: Based on the flow identifier information, determine the current path and at least one candidate path corresponding to the target set communication flow, and collect telemetry data of the current path and each candidate path.

[0040] In one specific embodiment, after obtaining the flow identifier information of the target set communication flow, the network control system can query the network forwarding status, path mapping relationship, or flow table matching relationship based on the flow identifier information to determine the network forwarding path that the target set communication flow is currently actually traversing. Specifically, based on the aforementioned flow identifier information, the system searches for forwarding records matching the target set communication flow in the network controller, switch flow table, path database, or telemetry acquisition component, thereby determining the current path corresponding to the target set communication flow. The current path may include the switching devices, ports, links, and intermediate forwarding nodes traversed by the target set communication flow from the source training node to the destination training node or destination node group.

[0041] Based on the current path, the network control system can further determine at least one candidate path according to network topology information, equivalent path information, link availability status, parallel group node distribution, and path policy constraints. A candidate path refers to an alternative path that can replace the current path to carry the target set of communication flows when path adjustments are needed. For example, in a graphics processing unit cluster or intelligent computing center network, there may be multiple equivalent paths between the source and destination nodes via different switching nodes or different uplinks. The network control system can select paths that meet connectivity, bandwidth, policy, and isolation requirements from these equivalent or reachable paths as candidate paths. Therefore, subsequent tail latency prediction is not only performed on the current path but also on both the current path and each candidate path, enabling the system to determine whether the current path has a risk of tail latency degradation and further determine which candidate path is more suitable as the target path.

[0042] In some embodiments, the network control system can collect telemetry data of the current path and each candidate path. The telemetry data is used to characterize the delay state, congestion state, packet loss state, and stability state of the path during the transmission of the target set communication stream. Exemplarily, the telemetry data may include one or more of the following: path delay statistics, delay growth rate, delay jitter, explicit congestion notification marker ratio, queue depth, queue depth change rate, packet loss rate, port utilization, most recent probe round-trip time, or path historical stability index. The path delay statistics may include median delay, high quantile delay, or tail delay, etc.; the explicit congestion notification marker ratio can be used to reflect the degree of congestion marking in the path; queue depth and its change rate can be used to reflect the accumulation of buffer queues in the switching equipment; and the path historical stability index can be used to reflect the delay fluctuations and availability of candidate paths within a historical window.

[0043] In some embodiments, telemetry data can be obtained through one or more of in-band network telemetry, independent channel telemetry, active probing, or network device status acquisition. For the current path, the network control system can acquire information such as hop-by-hop delay, queue status, congestion markers, and packet loss status based on the actual forwarded packets of the target set communication flow. For candidate paths, since the target set communication flow has not yet actually switched to the candidate path, the network control system can acquire telemetry data related to the candidate path based on the link status, port queue status, historical stability data, probe packet return results, or path status information maintained by the network controller. By simultaneously acquiring telemetry data for the current path and candidate paths, the network control system can provide a data foundation for subsequently predicting the tail delay indicators of the current path and each candidate path within a preset prediction window.

[0044] In some embodiments, telemetry data can be collected according to a preset telemetry sampling period. For example, the preset telemetry sampling period can be set to 1ms to 10ms to ensure that changes in network state can be detected in a timely manner while avoiding excessive overhead on network devices or the control plane due to excessively high telemetry collection frequency. For training scenarios where the latency of aggregated communication stream tails fluctuates rapidly, a shorter telemetry sampling period can be used; for scenarios where the network state is relatively stable or the training task is small in scale, a longer telemetry sampling period can be used.

[0045] For example, in a training cluster of a two-layer leaf-spine network topology, if the target set communication flow is currently transmitted via the path "first leaf switch → second spine switch → second leaf switch", this path can be determined as the current path based on the flow identification information of the target set communication flow. Simultaneously, based on the available equivalent links in the topology, "first leaf switch → first spine switch → second leaf switch" is identified as a candidate path. Subsequently, telemetry data such as port latency, queue depth, explicit congestion notification marking ratio, packet loss rate, and most recent probe round-trip time are collected on both the current path and the candidate paths. If the current path exhibits an increase in queue depth, tail latency, or congestion marking ratio, while the candidate paths have high historical stability and low most recent probe latency, these telemetry data can serve as important inputs for subsequent tail latency prediction and candidate path evaluation.

[0046] In the above manner, step S2 maps the target set communication flow from the flow identifier on the training side to the current path and candidate path on the network side, and collects telemetry data that can characterize future transmission risks for these paths, so that subsequent tail delay prediction no longer depends only on the instantaneous state of the current path, but can be compared and judged by combining the network state and stability of the candidate path.

[0047] Step S3: Based on the training semantic context and telemetry data, predict the tail latency index of the current path and each candidate path within the preset prediction window.

[0048] In one specific embodiment, after acquiring the training semantic context corresponding to the target set communication flow, as well as the telemetry data of the current path and each candidate path, the network control system can associate the training-side information with the network-side information to form path-level prediction input data for tail delay prediction. That is, for the current path, the network control system can combine the training semantic context of the target set communication flow, such as the training stage, communication primitive type, communication message size, and synchronization window identifier, with the telemetry data corresponding to the current path, such as delay statistics, delay growth rate, queue status, packet loss rate, and explicit congestion notification marker ratio. Similarly, for each candidate path, the network control system can combine the same training semantic context with the telemetry data corresponding to that candidate path to form prediction input data for the current path and each candidate path, respectively. In this way, the tail delay prediction process not only considers the congestion state and transmission stability of the network path itself, but also the communication triggering position, communication load scale, and synchronization timing characteristics of the target set communication flow during the training process.

[0049] In this embodiment, the preset prediction window refers to a time window relative to the current prediction time, used to reflect the range of communication performance changes that the network control system needs to predict in advance. For example, the preset prediction window can be set according to the time scale of ensemble communication operations in distributed synchronous training of artificial intelligence, such as a time range that can cover one or more ensemble communication operations. Since ensemble communication flows in synchronous training are usually periodic and bursty, if the prediction window is too short, it may not be able to fully reflect the tail delay change trend; if the prediction window is too long, the prediction result may be affected by the uncertainty of subsequent network states. The prediction window is preferably 0.5 seconds to 2 seconds, which is well matched with the time scale of the training synchronization point. Therefore, the network control system can set the preset prediction window according to the synchronization period of the training task, the duration of ensemble communication, and the telemetry sampling frequency, so that the tail delay prediction result can match the training synchronization stage.

[0050] In this embodiment, the tail latency metric is used to characterize the tail latency risk of the target set communication flow within a preset prediction window. The tail latency metric can be a high quantile communication latency, such as the 95th percentile latency, the 99th percentile latency, or other statistical indicators of higher quantile latency. Preferably, the tail latency metric can be the 99th percentile tail latency, used to characterize the slow node risk caused by a small number of high-latency communication events during distributed synchronous training of artificial intelligence. Compared to the average latency, the tail latency metric better reflects the sensitivity of the set communication flow to high quantile latency anomalies in synchronous training scenarios, preventing high-latency samples from being masked by the average. In practical applications, the appropriate tail latency metric can be selected based on the training cluster size, the number of telemetry samples, and the sensitivity to slow node risk.

[0051] In one example, the network control system can input the latency distribution, queue depth changes, explicit congestion notification marking ratio, and packet loss rate of the current path within the most recent observation window, along with the training phase, set communication operation type, communication message size, and synchronization window identifier of the target set communication flow, into the tail latency prediction model to obtain the tail latency index of the current path within a preset prediction window. Simultaneously, for each candidate path, telemetry data such as its historical stability, most recent probe latency, queue status, and link utilization, along with the same training semantic context, can be input into the tail latency prediction model to obtain the tail latency index of each candidate path within a preset prediction window. Thus, the network control system can simultaneously obtain both the potential tail latency risks of continuing to use the current path and the potential tail latency performance after switching to different candidate paths.

[0052] In some embodiments, the tail delay prediction model can be a lightweight time-series prediction model, a recurrent neural network model, a gated recurrent network model, an attention-based time-series model, or other prediction models capable of processing time-series telemetry data and training semantic features. The network control system can construct a telemetry sequence from the telemetry data in chronological order and use the training semantic context as the context feature corresponding to the telemetry sequence, enabling the tail delay prediction model to output the tail delay indices of the current path and each candidate path within a preset prediction window. The specific structure of the above prediction model does not constitute a limitation on the embodiments of this application, as long as it can output the tail delay prediction result corresponding to the path based on the training semantic context and path telemetry data.

[0053] For example, in the synchronization phase of a fully reduced communication, the training semantic context corresponding to the target set communication flow includes the fact that it is currently in the gradient synchronization phase, the set communication operation type is fully reduced communication, the communication message size is the data volume of this gradient synchronization, and the synchronization window identifier is the gradient synchronization window in the current training iteration. The network control system collects data showing that the tail latency of the current path within the most recent observation window is increasing, the queue depth is increasing, and the proportion of explicit congestion notification markers is increasing; at the same time, the recent detection latency of the candidate paths is low and the historical stability is high. After inputting the above training semantic context and the telemetry data of each path into the tail latency prediction model, the tail latency indices of the current path and each candidate path within the preset prediction window can be obtained. If the predicted tail latency index of the current path is significantly higher than that of the candidate paths, it indicates that the current path has a higher risk of tail latency degradation in subsequent synchronization windows, while the candidate paths may have better transmission performance.

[0054] In the above manner, step S3 uses the training semantic context and network path telemetry data from distributed synchronous training in artificial intelligence together for tail delay prediction, enabling the prediction process to simultaneously perceive the training semantic features and path operation status of the target set communication flow. Compared to methods that only determine paths based on the current link status or average delay, this embodiment of the application can pre-judge the future transmission risks of the current path and candidate paths before tail delay degradation actually affects synchronous training, providing a predictive basis for subsequent candidate path evaluation, path adjustment triggering, and closed-loop control.

[0055] Step S4: Using the tail delay index of the current path as a comparison benchmark, evaluate each candidate path based on the tail delay index of each candidate path and the corresponding path adjustment cost, and determine the target path from each candidate path.

[0056] In distributed synchronous training environments for artificial intelligence, there are typically multiple parallel groups, multiple training tasks, or multiple synchronization windows corresponding to aggregated communication flows. Different aggregated communication flows may share some link, switching port, or queue resources. If a switch is performed solely based on the lower predicted tail latency of a candidate path, other aggregated communication flows on that candidate path may be subject to new congestion, or the control plane overhead and connection migration overhead of the path switch itself may outweigh the tail latency improvement benefits. Therefore, this embodiment determines the target path by comprehensively considering the path cost and expected communication benefits, so that path selection simultaneously considers the tail latency improvement of the target aggregated communication flow itself and the overall network transmission stability.

[0057] In a preferred embodiment, step S4 specifically includes:

[0058] Step S401: Determine the path switching cost required to adjust from the current path to each candidate path, and the disturbance cost of each candidate path to other concurrent set communication flows.

[0059] In one specific embodiment, path switching cost can be used to characterize the control and transmission overhead introduced by adjusting the target set of communication flows from the current path to a candidate path. The network control system can determine the path switching cost required to adjust from the current path to each candidate path based on the path control method, the number of flow table updates, network device response latency, potential short-term interruptions during path switching, connection migration overhead, and the degree of path difference between the candidate path and the current path. For example, in a scenario using software-defined network control, if adjusting the target set of communication flows to a candidate path requires issuing or modifying flow table rules on multiple switching devices, the path switching cost corresponding to that candidate path can be determined based on the number of flow table issuances, switching device processing latency, interaction latency between the controller and the switching devices, and the estimated completion time of the path switching. If a candidate path differs from the current path only in some links and requires fewer forwarding rule updates, its path switching cost can be lower than that of a candidate path that requires completely changing intermediate forwarding nodes.

[0060] In this embodiment, the perturbation cost can be used to characterize the potential impact on other concurrently transmitted aggregated communication flows in the network after the target aggregated communication flow is adjusted to a candidate path. The perturbation cost for each candidate path can be determined based on the number of existing aggregated communication flows on the candidate path, the priority of concurrent aggregated communication flows, the link utilization of the candidate path, queue occupancy, the degree to which the candidate path shares links or ports with other aggregated communication flows, and the expected increase in bandwidth usage after the switch. For example, if a candidate path currently carries fully regulated communication flows from other data parallel groups, and the queue depth of that path is already high, switching the target aggregated communication flow to that candidate path may further increase queue congestion and tail latency risks. Therefore, the perturbation cost of that candidate path can be set higher. Conversely, if a candidate path has low link utilization, few concurrent aggregated communication flows, and high historical stability, the perturbation cost of that candidate path can be set lower.

[0061] Step S402: Determine the comprehensive path cost of each candidate path based on the tail delay index, path switching cost, and disturbance cost of each candidate path.

[0062] In one specific embodiment, after obtaining the tail delay index of each candidate path within a preset prediction window, the path switching cost required to adjust from the current path to each candidate path, and the disturbance cost of each candidate path to other concurrent aggregated communication flows in the network, the network control system can determine the comprehensive path cost of each candidate path based on the above three factors. This comprehensive path cost characterizes the comprehensive communication cost corresponding to a candidate path as a target path; the lower the comprehensive path cost, the lower the tail delay risk, the smaller the path switching overhead, and the smaller the disturbance to other aggregated communication flows within the preset prediction window, and therefore the more suitable it is as a target path.

[0063] In this embodiment, the tail latency index, path switching cost, and disturbance cost of the candidate path are all unfavorable factors in path selection. Specifically, a higher tail latency index indicates a greater likelihood of tail latency degradation within a preset prediction window after the target set of communication flows switches to that candidate path; a higher path switching cost indicates higher control plane overhead, flow table update overhead, or connection migration overhead required to adjust from the current path to the candidate path; and a higher disturbance cost indicates a more significant potential impact of the candidate path on other concurrent set of communication flows in the network. Therefore, when determining the comprehensive path cost, these unfavorable factors can be accumulated as positive cost terms, resulting in a lower predicted tail latency, lower path switching cost, and lower disturbance cost for the candidate path, leading to a lower comprehensive path cost.

[0064] In some embodiments, since the tail delay index, path switching cost, and disturbance cost may have different dimensions or value ranges, the above data can be normalized first to bring them within a comparable numerical range. Then, the comprehensive path cost of the candidate path is determined based on the tail delay index and the weighted path switching cost and disturbance cost. The tail delay index can serve as the basic cost term for candidate path evaluation, directly reflecting the tail delay risk of the candidate path within a preset prediction window. The path switching cost and disturbance cost can be adjusted using preset weight coefficients to adapt to different network control strategies or cluster operating states. For example, in network environments with high path switching overhead or sensitive control surfaces, the impact weight corresponding to the path switching cost can be increased to reduce control surface jitter caused by frequent switching. In scenarios with multiple concurrent training tasks or multiple concurrent parallel group communication flows, the impact weight corresponding to the disturbance cost can be increased to reduce the impact of target group communication flow path adjustments on other group communication flows in the network. Through the above normalization and weight configuration methods, candidate path evaluation can adaptively adjust according to the actual training task and network operating environment, thereby improving the stability and rationality of target path selection.

[0065] In a preferred embodiment, for any candidate path pᵢ, its comprehensive path cost can be determined by the following formula:

[0066]

[0067] Among them, Score p ᵢ represents the comprehensive path cost of candidate path pᵢ; PredP99 p ᵢ represents the predicted delay value of candidate path pᵢ at the 99th percentile within the preset prediction window; SwitchCost p ᵢ represents the path switching cost required to adjust from the current path to the candidate path pᵢ; Disturbance p ᵢ represents the disturbance cost generated by other concurrent set communication flows in the candidate path pᵢ network; λ and μ represent preset weight coefficients, which are used to adjust the influence of path switching cost and disturbance cost on the comprehensive path cost.

[0068] Based on the above calculations, a lower overall path cost indicates a better candidate path. In other words, if a candidate path has a low prediction tail delay but a high path switching cost, or if it causes significant disturbance to other concurrent network traffic, its overall path cost may still be high, making it unsuitable as the target path. Conversely, if another candidate path has a slightly higher prediction tail delay but a lower path switching cost and less disturbance to other network traffic, its overall path cost may be even lower, making it more suitable as the target path. This calculation method unifies the prediction tail delay, path switching cost, and disturbance cost of candidate paths into a single cost scale, thus avoiding using only the lowest tail delay as the criterion for target path selection.

[0069] Step S403: Based on the difference between the tail delay metric of the current path and the comprehensive path cost of each candidate path, determine the expected communication benefit of each candidate path relative to the current path.

[0070] In one specific embodiment, the current path is the actual transmission path used by the target set communication flow before path adjustment. If path adjustment is not performed, the target set communication flow will continue to be transmitted via the current path. Therefore, the tail latency index of the current path within a preset prediction window can be used as the benchmark communication cost when path adjustment is not performed, characterizing the tail latency risk that the target set communication flow may face when continuing to transmit along the current path. Since the current path does not involve the control overhead of switching from the current path to other paths, in this embodiment, the tail latency index of the current path can be directly used as the comparison benchmark.

[0071] For each candidate path, step S402 has determined the corresponding comprehensive path cost based on the candidate path's tail delay metric, path switching cost, and disturbance cost. This comprehensive path cost characterizes the overall communication cost after the target set's communication flow is adjusted to the candidate path. The network control system can compare the tail delay metric of the current path with the comprehensive path cost of each candidate path to determine the expected communication benefit of each candidate path relative to the current path. The expected communication benefit can be understood as: the degree of tail delay improvement that the candidate path can achieve compared to continuing to use the current path, after factoring in the path switching cost and disturbance cost based on the predicted tail delay of the candidate path.

[0072] In some embodiments, the expected communication benefit can be determined based on the difference between the tail delay metric of the current path and the comprehensive path cost of the candidate path. For example, the expected communication benefit of the candidate path relative to the current path can be obtained by subtracting the comprehensive path cost of the candidate path from the tail delay metric of the current path. When the difference is positive and large, it indicates that the candidate path is still superior to the current path after comprehensively considering tail delay, path switching cost, and disturbance cost; when the difference is small or negative, it indicates that the improvement of the candidate path relative to the current path is insufficient, and it has no switching value after taking into account path adjustment costs and concurrent disturbances. In this way, path adjustment can be avoided simply because the predicted tail delay of the candidate path is slightly lower than that of the current path, thereby reducing the risk of invalid switching and path oscillation.

[0073] For example, if the current path has a high tail latency within the preset prediction window, and a candidate path has a low predicted tail latency but high path switching and disturbance costs, then the overall path cost of the candidate path may not be low, and its expected communication gains relative to the current path may be insufficient. Conversely, if a candidate path not only has a low predicted tail latency but also low path switching and disturbance costs, then the overall path cost of the candidate path is low, and its expected communication gains relative to the current path are greater, making it more suitable as a candidate target for subsequent path adjustments.

[0074] Step S404: The candidate path with the expected communication benefits exceeding the preset switching benefit threshold and the lowest overall path cost is determined as the target path.

[0075] In one specific embodiment, after determining the expected communication gains of each candidate path relative to the current path, the network control system can compare the expected communication gains with a preset handover gain threshold. The preset handover gain threshold is used to limit the minimum improvement that a candidate path should achieve relative to the current path, avoiding triggering path adjustments when a candidate path only shows a slight improvement or the improvement is insufficient to cover the path handover overhead and concurrent disturbance effects. In other words, a candidate path is not necessarily designated as the target path simply because its tail latency is lower than the current path; rather, the candidate path must still have sufficient expected communication gains relative to the current path after comprehensively considering the tail latency, path handover cost, and disturbance cost.

[0076] In this embodiment, candidate paths with expected communication benefits exceeding a preset switching benefit threshold can be selected from each candidate path to form a set of switchable candidate paths. If the expected communication benefit of a candidate path does not exceed the preset switching benefit threshold, it indicates that the improvement of the candidate path relative to the current path is insufficient, or that it does not have sufficient switching value after considering path switching costs and disturbance costs. Therefore, the candidate path is not considered as a target path candidate. If there are no candidate paths with expected communication benefits exceeding the preset switching benefit threshold, the network control system can continue to transmit the target set communication flow through the current path and wait for the next round of tail delay prediction and path evaluation results.

[0077] After forming a set of switchable candidate paths, the network control system can compare the comprehensive path costs of each candidate path in the set and determine the candidate path with the lowest comprehensive path cost as the target path. Since the comprehensive path cost has taken into account the tail delay index of the candidate path within the preset prediction window, the path switching cost required to adjust from the current path to the candidate path, and the disturbance cost of the candidate path to other concurrent communication flows in the network, selecting the candidate path with the lowest comprehensive path cost can make the target path have a better overall performance in terms of tail delay risk, path adjustment cost, and concurrent communication disturbance.

[0078] For example, the tail latency of the current path within the preset prediction window is 500 microseconds, and the preset handover benefit threshold is 80 microseconds. If the comprehensive path cost of candidate path p1 is 390 microseconds, its expected communication benefit relative to the current path is 110 microseconds; if the comprehensive path cost of candidate path p2 is 320 microseconds, its expected communication benefit relative to the current path is 180 microseconds; and if the comprehensive path cost of candidate path p3 is 450 microseconds, its expected communication benefit relative to the current path is 50 microseconds. Since the expected communication benefits of candidate paths p1 and p2 both exceed the preset handover benefit threshold, and candidate path p2 has the lowest comprehensive path cost, candidate path p2 can be identified as the target path. Although candidate path p3 shows some improvement compared to the current path, its expected communication benefit does not exceed the preset handover benefit threshold, therefore it is not considered as the target path.

[0079] In the solution provided in this embodiment, through steps S401 to S404, the tail delay index of the candidate path, the path switching cost, and the disturbance cost to the concurrent aggregated communication flow are comprehensively considered during the target path determination process, and the expected communication benefit of the candidate path is determined using the current path as a comparison benchmark. Therefore, invalid switching can be avoided when the improvement of the candidate path is small or insufficient to cover the path adjustment cost, thereby reducing the risk of frequent path adjustments and path oscillations, and improving the rationality of target path selection and the stability of path control.

[0080] Step S5: If the tail delay index of the current path meets the preset path control conditions and the preset suppression conditions are not met, adjust the target set communication flow from the current path to the target path.

[0081] In one specific embodiment, after determining the target path, the network control system can further determine whether the current path meets preset path control conditions and whether it hits preset suppression conditions to determine whether to actually perform path adjustment. The preset path control conditions are used to determine whether there is a tail latency degradation risk requiring intervention in the current path, and the preset suppression conditions are used to determine whether there are still situations unsuitable for path adjustment even if there is a tail latency degradation risk. In other words, determining the target path addresses the question of "which path to switch to," while the preset path control conditions and preset suppression conditions address the question of "whether and when to perform the switch."

[0082] In some embodiments, the tail delay index of the current path within a preset prediction window can be obtained, and combined with the growth rate of the tail delay index of the current path and the number of consecutive windows exceeding the threshold, it can be determined whether the current path meets the preset path control conditions. Specifically, the tail delay index of the current path can be determined to meet the preset path control conditions if at least two of the following conditions are met: the tail delay index of the current path exceeds a preset baseline threshold within the preset prediction window; the growth rate of the tail delay index of the current path within the preset prediction window exceeds a preset growth threshold; the tail delay index of the current path exceeds the preset baseline threshold in multiple consecutive preset prediction windows. By adopting a triggering method where at least two conditions are met simultaneously, path adjustment can be avoided from being triggered solely due to short-term fluctuations within a single prediction window or a single tail delay prediction value exceeding the threshold, making the path adjustment action more reflective of the risk of persistent or trending tail delay degradation in the current path in future periods.

[0083] In some embodiments, multiple consecutive preset prediction windows can be M consecutive preset prediction windows, where M is an integer greater than 1. For example, M can be set to 2 to 5. By setting a consecutive window count, occasional tail delay fluctuations within a single prediction window can be avoided from directly triggering path adjustments; when the current path experiences tail delay indicators exceeding the threshold in all M consecutive preset prediction windows, it can more reliably indicate that the current path has a risk of persistent tail delay degradation.

[0084] After determining that the current path meets the preset path control conditions, it can be further determined whether the preset suppression conditions are met. The preset suppression conditions are used to limit situations where immediate path adjustments are not suitable, thereby reducing unnecessary disturbances to the training process and network state caused by path adjustments. For example, the preset suppression conditions may include at least one of the following: the target set communication flow is within a cooling window after path adjustment; the AI ​​distributed synchronous training process is in a preset critical synchronization phase; the expected communication gain of the target path relative to the current path is lower than a preset gain guarantee threshold; the most recent path adjustment result for the target set communication flow failed. The cooling window is used to prevent the same target set communication flow from repeatedly performing path adjustments within a short period; the preset critical synchronization phase can be a phase during training where communication stability requirements are high and path switching is not advisable; the preset gain guarantee threshold is used to limit the minimum gain margin that the target path should achieve relative to the current path, to avoid triggering invalid switching when the expected communication gain is insufficient or the gain boundary is unstable; a failed path adjustment result indicates that the target set communication flow or the current network state may not be suitable for continuing to perform the same type of path adjustment.

[0085] In this embodiment, the network control system adjusts the target set communication flow from the current path to the target path only if the tail delay index of the current path meets the preset path control conditions and the aforementioned preset suppression conditions are not met. Specifically, a corresponding target path forwarding rule can be generated based on the flow identifier information of the target set communication flow, and this rule can be distributed to the network forwarding devices or path control components traversed by the target set communication flow. This causes subsequent data packets matching the flow identifier information to be switched from the current path to the target path for forwarding. The target path forwarding rule may include one or more of the following: flow matching field, target outgoing port, next-hop node, tunnel identifier, path label, or forwarding priority, to achieve targeted path adjustment of the target set communication flow.

[0086] For example, in a distributed synchronous training task in artificial intelligence, the target set communication flow can be a fully reduced communication flow generated during the gradient synchronization phase. (See reference...) Figure 3 As shown, after the target path is determined, the tail latency index of the current path of the target set communication flow within a preset prediction window can be obtained, and it can be determined whether the tail latency index exceeds a preset baseline threshold. Simultaneously, based on the changes in the tail latency index within adjacent prediction windows, it can be determined whether the growth rate of the current path's tail latency index exceeds a preset growth threshold, and further, it can be determined whether the current path has experienced tail latency index exceeding the threshold in multiple consecutive preset prediction windows. Through the above judgments, it is possible to identify whether the current path has tail latency risks such as short-term anomalies, continuous threshold exceeding, or trend deterioration. Subsequently, it is determined whether the target set communication flow is within a cooling window, whether the training process is in a preset critical synchronization stage, whether the expected communication gain of the target path relative to the current path is lower than a preset gain guarantee threshold, and whether the most recent path adjustment result of the target set communication flow was a failure. If none of the above suppression conditions are met, the network control system can perform path adjustment; if the current path does not meet the preset path control conditions, or although it meets the preset path control conditions, it meets any preset suppression condition, the network control system can temporarily refrain from performing path adjustment and allow the target set communication flow to continue transmitting via the current path.

[0087] Figure 3The illustrated process is merely an exemplary implementation of this application and should not be construed as limiting the order of judgment, combination of judgment conditions, or type of suppression conditions. Through the above method, this embodiment distinguishes between whether the current path has a risk of tail latency degradation and whether the current path adjustment is suitable: preset path control conditions are used to confirm the necessity of path adjustment, and preset suppression conditions are used to constrain the timing of path adjustment. Therefore, invalid path adjustments can be avoided due to single tail latency anomalies, insufficient gains, critical training phases, or repeated switching within a short period, thereby reducing the risks of false triggering, frequent switching, and path oscillation, and improving the stability and effectiveness of target set communication flow path adjustment.

[0088] Step S6: Obtain the operational feedback information after path adjustment, and based on the operational feedback information, perform maintain or rollback control on the target path, and update the prediction process of the tail delay index.

[0089] In one specific embodiment, refer to Figure 4 As shown, the network control system may include a collaborative tail delay prediction controller, a path controller, and a prediction model update module. This module verifies the actual operating state of the target path after the target set communication flow is adjusted from the current path to the target path, rather than directly considering the path adjustment result as the final valid result. Specifically, the tail delay prediction controller performs cooling window management, active detection scheduling, stability maintenance determination, backoff determination, and prediction error information generation based on the operational feedback information after path adjustment. The path controller adjusts or restores the forwarding path of the target set communication flow according to the path adjustment rules or backoff rules issued by the tail delay prediction controller. The prediction model update module receives path control feedback samples, prediction error information, path adjustment results, and backoff reasons, and updates the prediction model used to predict the tail delay index when the model update trigger condition is met. Figure 4 The collaborative processing flow shown can form a closed-loop feedback control flow that connects the operation verification after path adjustment, path maintenance or rollback control, and prediction model update process. This can prevent situations where the target path performs well in the prediction stage but fails to detect, suffers from latency degradation or increased packet loss after actually carrying the target set communication flow. This reduces the impact of path adjustment failure on the distributed synchronous training process of artificial intelligence and improves the continuous correction capability of the tail latency prediction process.

[0090] In a preferred embodiment, step S6, which involves obtaining operational feedback information after path adjustment and controlling the target path to maintain or revert based on the operational feedback information, includes:

[0091] Step S601: After adjusting the target set communication flow from the current path to the target path, enter the cooling window.

[0092] In one specific embodiment, after determining the target path and triggering path adjustment, the tail delay prediction controller can issue target path forwarding rules to the path controller. According to the target path forwarding rules, the path controller adjusts the target set communication flow from the current path to the target path and returns a successful path adjustment result to the tail delay prediction controller. Upon receiving the successful path adjustment result, the tail delay prediction controller records information such as the flow identifier of the target set communication flow, the current path before adjustment, the target path after adjustment, the path adjustment time, the predicted tail delay index, the comprehensive path cost, and the expected communication benefits, and initiates a cooling window for that target set communication flow. Within the cooling window, the tail delay prediction controller can suppress non-urgent path adjustment requests for the same target set communication flow to avoid triggering path switching again due to short-term detection fluctuations or telemetry data jitter immediately after path adjustment.

[0093] In this embodiment, the cooling window refers to a preset time window after the target aggregate communication flow completes a path adjustment. The duration of the cooling window can be configured according to the iteration cycle of the AI ​​distributed synchronous training task, the duration of the aggregate communication operation, the path adjustment stabilization time, or the network control strategy. For example, the duration of the cooling window can be set to 30 seconds to 120 seconds. Within this cooling window, the tail delay prediction controller can continue to collect operational feedback information of the target path, but generally will not perform non-urgent path adjustments again for the same target aggregate communication flow, in order to avoid triggering path switching again due to short-term detection fluctuations, telemetry data jitter, or local instantaneous congestion immediately after the path adjustment is completed. The duration of the cooling window can be configured according to the iteration cycle of the AI ​​distributed synchronous training task, the duration of the aggregate communication operation, the telemetry sampling cycle, or the network control strategy. For example, for training tasks with high aggregate communication frequency and sensitive synchronization waiting, a cooling duration corresponding to several training iteration cycles or several aggregate communication windows can be set to ensure that the target path has sufficient observation time after adjustment.

[0094] Step S602: Within the cooling window, send probe messages along the target path according to the preset probe cycle to obtain operational feedback information including probe success rate, measured latency and packet loss rate.

[0095] In one specific embodiment, within the cooling window, the tail delay prediction controller can generate an active probe task for the target path according to a preset probe period, and send probe messages along the target path based on the flow identification information of the target path, or instruct the path controller to send probe messages along the target path. The probe messages can be configured to have the same or similar forwarding path as the target set communication flow, so as to obtain the actual transmission status of the target path after path adjustment.

[0096] For example, the preset detection period can be set to 1 to 5 seconds. By setting the above detection period, the actual running feedback of the target path can be continuously obtained within the cooling window, while avoiding excessive detection frequency that would consume additional network resources. For scenarios with rapid tail latency fluctuations or sensitivity to training synchronization, a shorter detection period can be used; for scenarios with relatively stable network conditions, a longer detection period can be used.

[0097] In this embodiment, the path controller can feed back the transmission results, return results, or timeout results of probe packets to the tail delay prediction controller. The tail delay prediction controller determines the operational feedback information of the target path based on the above feedback results. The operational feedback information may include probe success rate, measured latency, and packet loss rate. Among them, the probe success rate can be determined according to the ratio between the number of probe packets that successfully return within a preset statistical window and the total number of probe packets sent; the measured latency can be determined according to the transmission time and return time of the probe packets, or it can be expressed in the form of round-trip latency, one-way latency, or high-quantum measured latency; the packet loss rate can be determined according to the number of probe packets that do not return, return after timeout, or return abnormally.

[0098] In some embodiments, the operational feedback information may further include information such as the queue depth of the target path, port utilization, explicit congestion notification marking ratio, latency jitter, retransmission status, or path reachability status, to more comprehensively characterize the operational quality of the target path after path adjustment. This information can be obtained by the path controller from network forwarding devices, telemetry acquisition components, or probe results, and fed back to the tail delay prediction controller.

[0099] It should be noted that the purpose of sending probe messages within the cooling window is not to reselect candidate paths, but to verify the selected and adjusted target path. By periodically acquiring operational feedback information such as probe success rate, measured latency, and packet loss rate within the cooling window, transmission anomalies of the target path after it actually carries the target set of communication flows can be detected in a timely manner, providing a basis for subsequent judgment on whether the target path meets the stability maintenance conditions or the fallback conditions.

[0100] Step S603: Based on the operation feedback information, determine whether the target path meets the stability maintenance condition or the rollback condition.

[0101] In one specific embodiment, after obtaining the operational feedback information of the target path, the tail delay prediction controller can compare the operational feedback information with preset stability maintenance conditions and preset backoff conditions to determine whether the target path is suitable to continue carrying the target set communication flow, or whether it needs to backoff to the current path before adjustment. In other words, this step is used to verify the adjustment effect of the target path based on the actual operational feedback after path adjustment, rather than determining whether the path adjustment is effective solely based on the prediction results before path adjustment.

[0102] The stability maintenance condition characterizes the target path as being in a stable and available state after path adjustment. For example, the stability maintenance condition may include one or more of the following: the detection success rate meets a preset detection success condition, the measured latency does not exceed a corresponding latency threshold, and the packet loss rate does not exceed a corresponding packet loss threshold. The preset detection success condition can be a detection success rate reaching a preset proportion, or obtaining normal detection results for multiple consecutive detection cycles; the latency threshold can be determined based on the predicted tail latency index of the target path, the sensitivity of the current training task to synchronization waiting, the historical path latency distribution, or preset service quality requirements; the packet loss threshold can be configured based on the tolerance of the target set communication flow to packet loss, retransmission, or communication congestion.

[0103] The fallback condition is used to characterize that the target path is no longer suitable for carrying the target set of communication flows after path adjustment. For example, the fallback condition may include: a preset number of consecutive probe execution failures, or at least one of the measured latency and packet loss rate exceeding the corresponding threshold. Probe execution failure may include situations such as probe packets not being returned, return timeout, path unreachable, abnormal probe response, or probe packets not being forwarded according to the target path.

[0104] In some embodiments, a preset number of consecutive probe failures can be understood as K consecutive probe failures, where K is an integer greater than 1. For example, K can be set to 2 to 4. By setting a threshold K for the number of consecutive probe failures, premature triggering of fallback control due to single probe packet loss, return timeout, or momentary network jitter can be avoided. When K consecutive probe failures occur, it is possible to more reliably determine if the target path has issues such as path unreachability, unstable transmission, or abnormal performance, thereby timely reverting the target set communication flow to the current path before adjustment.

[0105] In this embodiment, the tail delay prediction controller can determine whether the target path meets the stability maintenance condition or the backoff condition based on the operational feedback information. If the operational feedback information indicates that the target path meets the stability maintenance condition, the target path forwarding rule can be maintained, allowing the target set communication flow to continue to be transmitted via the target path. If the operational feedback information indicates that the target path meets the backoff condition, a backoff rule can be issued to the path controller, causing the target set communication flow to backoff from the target path to the current path before the adjustment. Therefore, the target path can be verified based on the actual operational state after path adjustment, avoiding the continuous impact of erroneous path adjustments on the distributed synchronous training process of artificial intelligence.

[0106] Step S604: If the detection success rate meets the preset detection success condition, and the measured latency and packet loss rate do not exceed the corresponding threshold, then the target path is determined to meet the stability maintenance condition, and the target set communication flow is maintained through the target path.

[0107] In one specific embodiment, if the operational feedback information obtained by the tail delay prediction controller within the cooling window indicates that the detection success rate of the target path meets the preset detection success condition, and the measured delay and packet loss rate of the target path do not exceed the corresponding thresholds, then it can be determined that the target path meets the stability maintenance condition. At this time, the tail delay prediction controller can generate a path maintenance instruction, or maintain the issued target path forwarding rules in a valid state, allowing the target set communication flow to continue to be transmitted via the target path. Specifically, the path controller can maintain the target path forwarding rules already issued to the network forwarding device, allowing subsequent packets matching the target set communication flow to continue to be forwarded according to the target path.

[0108] For example, if the detection results of the target path remain normal within the cooling window, and the measured latency and packet loss rate are both within acceptable ranges, it indicates that the target path not only has a low tail latency risk during the prediction phase, but also can stably carry the target set communication flow after path adjustment. In this case, maintaining the target set communication flow through the target path can reduce unnecessary backoff actions and additional control overhead, and enable the target set communication flow to achieve more stable communication performance.

[0109] Step S605: If the detection fails for a preset number of consecutive times, or if at least one of the measured latency and packet loss rate exceeds the corresponding threshold, then the target path is determined to meet the rollback condition, and the target set communication flow is rolled back from the target path to the current path before adjustment.

[0110] In one specific embodiment, if the tail delay prediction controller detects a preset number of consecutive probe execution failures within the cooling window, or if at least one of the measured delay and packet loss rate of the target path exceeds the corresponding threshold, it can be determined that the target path meets the backoff condition. At this time, the tail delay prediction controller can trigger backoff control to revert the target set communication flow from the target path to the current path used before the path adjustment. In this way, the path adjustment result can be terminated promptly when the actual operating state of the target path does not meet the requirements, preventing the target set communication flow from continuously transmitting through the abnormal target path and affecting the distributed synchronous training process of artificial intelligence.

[0111] In practice, the tail delay prediction controller can generate fallback rules based on the current path information before adjustment recorded in step S601, and send the fallback rules to the path controller. The path controller, according to the fallback rules, restores or reconfigures the forwarding rules corresponding to the target set communication flow, so that subsequent packets matching the target set communication flow are forwarded again via the current path before adjustment. If target path forwarding rules, path labels, tunnel identifiers, or flow table rules were used during the path adjustment process, the fallback control can include canceling or invalidating the target path forwarding rules, restoring the forwarding rules corresponding to the path before adjustment, or remapping the path labels of the target set communication flow to the current path before adjustment.

[0112] In the solution provided in this embodiment, through steps S601 to S605, a cooling window and a periodic detection mechanism are introduced after the path adjustment is completed to continuously verify the actual operating status of the target path. Based on operational feedback information such as detection success rate, measured latency, and packet loss rate, it is determined whether the target path is suitable to continue carrying the target set communication flow. Therefore, it avoids confirming the effectiveness of the path adjustment solely based on prediction results. When the target path is running stably, transmission is maintained along the target path. If the target path experiences continuous detection failures, latency exceeding the threshold, or abnormal packet loss, it promptly reverts to the path before adjustment. This reduces the continuous impact of path adjustment failures on the distributed synchronous training process of artificial intelligence, and reduces the risk of frequent switching and path oscillation.

[0113] In a preferred embodiment, step S6, the step of updating the prediction process of the tail delay index, includes:

[0114] Step S606: Obtain the tail delay index of the target set communication flow before path adjustment, the measured tail delay after path adjustment, the path adjustment result, the operation feedback information and the reason for rollback, and generate path control feedback samples.

[0115] In one specific embodiment, the prediction model update module can receive path control process data written by the tail delay prediction controller and generate path control feedback samples based on the path control process data. The path control feedback samples are used to record the complete information of the path control process of a target set communication flow, from tail delay prediction, target path selection, path adjustment execution, running feedback verification to stable maintenance or backoff control, thereby providing sample basis for subsequent tail delay index prediction process updates.

[0116] Specifically, the path control feedback sample may include one or more of the following: the tail delay metric of the target set communication flow before path adjustment, the predicted tail delay metric of the target path, the measured tail delay after path adjustment, the path adjustment result, operational feedback information, and the reason for rollback. The tail delay metric before path adjustment may include the predicted tail delay metric of the current path within a preset prediction window, or the predicted tail delay metric of the target path before path adjustment. The measured tail delay after path adjustment can be obtained statistically from the probe packet delay collected within the cooling window, the actual communication packet delay, or the communication completion delay of the target set communication flow. The path adjustment result may include successful path adjustment and maintenance of the target path, triggering rollback control, path adjustment failure, or path adjustment not taking effect. The operational feedback information may include probe success rate, measured delay, packet loss rate, path reachability status, queue depth, port utilization, or explicit congestion notification marking ratio. The reasons for rollback may include continuous probe execution failure, measured delay exceeding the threshold, packet loss rate exceeding the threshold, target path unreachable, or target path performance unstable.

[0117] In some embodiments, the prediction model update module can also write information such as the training semantic context, telemetry data, current path information, target path information, comprehensive path cost, expected communication gains, cooling window duration, and detection cycle when generating the path control feedback sample into the path control feedback sample. By retaining the above information, the path control feedback sample can not only record whether the path adjustment was successful, but also reflect the training stage, the type of aggregated communication operation, the size of the communication message, the synchronization window identifier, and the network operating status when the path adjustment occurred, thereby improving the interpretability and usability of the sample during subsequent prediction model updates.

[0118] Step S607: Based on the deviation between the tail delay index and the measured tail delay, determine the prediction error information corresponding to the path control feedback sample.

[0119] In one specific embodiment, after generating path control feedback samples, the prediction model update module can determine the prediction error information corresponding to the path control feedback sample based on the difference between the predicted tail delay index in the path control feedback sample and the measured tail delay after path adjustment. The prediction error information is used to characterize the estimation deviation of the tail delay prediction process on the actual operating state of the target path, thereby providing a basis for subsequent judgment on whether the prediction model needs to be updated and how to update the prediction model.

[0120] Specifically, the prediction model update module can calculate the absolute error, relative error, or error direction information between the predicted tail delay metric and the measured tail delay. The absolute error represents the magnitude of the difference between the predicted tail delay metric and the measured tail delay; the relative error represents the proportion of this difference relative to the predicted tail delay metric or the measured tail delay; and the error direction information indicates whether the prediction result underestimates or overestimates the actual tail delay risk of the target path. When the measured tail delay is higher than the predicted tail delay metric, it can be determined that the prediction process underestimates the tail delay risk of the target path; when the measured tail delay is lower than the predicted tail delay metric, it can be determined that the prediction process overestimates the tail delay risk of the target path.

[0121] In some embodiments, prediction error information can also be determined by combining path adjustment results and operational feedback information. For example, if the target path triggers backoff control after path adjustment, even if the difference between the predicted tail delay and the measured tail delay at a certain moment is not large, the path control feedback sample can still be marked as a high-risk prediction error sample by combining operational feedback information such as continuous probe execution failure, packet loss rate exceeding the threshold, and target path unreachable. If the target path remains stable within the cooling window, and the deviation between the measured tail delay and the predicted tail delay index is small, the path control feedback sample can be marked as a low-error success sample. Thus, prediction error information can not only reflect numerical deviations but also the consistency between the prediction results and the actual path control effect.

[0122] Step S608: Filter model update samples from path control feedback samples, wherein the model update samples include: successful path adjustment samples, path rollback samples that trigger rollback control, and samples whose prediction error information exceeds a preset error threshold.

[0123] In one specific embodiment, the prediction model update module can write path control feedback samples into a feedback sample cache or sample library, and select model update samples from the path control feedback samples according to preset filtering rules. The model update samples are used to participate in subsequent prediction model updates, enabling the prediction model to correct its prediction capability for tail delay risk based on the actual operation feedback after path adjustment.

[0124] Specifically, model update samples can include successful path adjustment samples, path rollback samples that trigger rollback control, and samples where prediction error information exceeds a preset error threshold. Successful path adjustment samples represent positive feedback samples where the prediction result is relatively consistent with the actual running result, enabling the prediction model to learn feature combinations that can bring stable tail latency improvement to candidate paths under the corresponding training semantic context and network telemetry status. Path rollback samples represent negative feedback samples that were judged as superior during the prediction stage but failed to meet the requirements after actual operation, enabling the prediction model to learn feature combinations that lead to path adjustment failure or unstable target path performance. Samples where prediction error information exceeds a preset error threshold are used to focus on correcting prediction biases of the prediction model under specific paths, specific training stages, specific set communication operation types, or specific congestion states.

[0125] In some embodiments, the prediction model update module can also perform sample deduplication, abnormal sample filtering, sample equalization, or time decay processing when screening model update samples. For example, for highly similar feedback samples generated by communication flows of the same target set within a short period of time, representative samples can be retained to reduce the impact of duplicate samples on model updates; abnormal samples that are obviously caused by equipment failure, link interruption, or non-training traffic bursts can be marked or isolated to avoid their unstable impact on the prediction model; for path backoff samples that are few in number but have high correction value, their sampling ratio in the model update samples can be increased. This improves the effectiveness and representativeness of the model update samples.

[0126] Step S609: If the model update triggering conditions are met, update the prediction model used to predict the tail delay index based on the model update samples; wherein, the model update triggering conditions include at least one of the following: the number of model update samples reaches a preset number threshold, the time since the last model update reaches a preset update time period, or the proportion of samples with prediction error information exceeding a preset error threshold reaches a preset proportion threshold.

[0127] In one specific embodiment, after selecting model update samples, the prediction model update module can determine whether the model update trigger condition is met. If the model update trigger condition is met, the prediction model used to predict the tail delay index is updated based on the model update samples. By setting the model update trigger condition, the computational overhead or frequent model fluctuations caused by updating the prediction model every time a single or small number of feedback samples are generated can be avoided. At the same time, the prediction process can be corrected in a timely manner when the prediction error continues to increase or the model update samples accumulate to a certain scale.

[0128] Model update triggering conditions may include at least one of the following: the number of model update samples reaches a preset quantity threshold; the time since the last model update reaches a preset update time period; the proportion of samples with prediction error exceeding a preset error threshold reaches a preset proportion threshold. The preset quantity threshold ensures that model updates have a sufficient sample base, avoiding unstable updates based on a small number of sporadic samples; the preset update time period allows the prediction model to periodically absorb new path control feedback to adapt to changes in training tasks, network topology, or traffic load; the preset proportion threshold triggers timely updates when a high-error sample set appears, reducing the possibility that the prediction model continuously underestimates tail latency risk or overestimates candidate path benefits.

[0129] In some embodiments, the prediction model update module can update the prediction model using methods such as incremental training, periodic fine-tuning, parameter calibration, or adaptive threshold adjustment. For example, the training semantic context, path telemetry data, candidate path features, predicted tail delay indicators, and measured tail delay in the model update samples can be used as training data to incrementally train the prediction model; the output calibration parameters of the prediction model can also be adjusted based on prediction error information to correct systematic biases in specific paths or specific set communication scenarios; and the risk penalty parameters in the candidate path evaluation process can be adjusted based on path backtracking samples to make the subsequent path selection process more cautious about similar high-risk paths.

[0130] After the model update is completed, the prediction model update module can record the model update version, update time, number of samples involved in the update, sample type distribution, change in prediction error before and after the update, and corresponding training task or network state information. The updated tail delay prediction model can be deployed to the tail delay prediction controller for predicting tail delay metrics for the current path and candidate paths within the subsequent prediction window. This allows the tail delay prediction results, path adjustment execution results, target path execution verification results, and feedback sample update process to be interconnected, enabling the tail delay prediction process to continuously correct based on path control feedback, thus forming a closed-loop network control mechanism oriented towards the target set communication flow.

[0131] In the solution provided in this embodiment, through steps S606 to S609, the predicted tail delay before path adjustment, the measured tail delay after path adjustment, the path adjustment result, the operation feedback information, and the reason for backoff are associated to generate path control feedback samples. Based on the deviation between the predicted tail delay and the measured tail delay, prediction error information is determined, thereby enabling the actual operation results during the path adjustment process to be used in reverse to evaluate the accuracy of the tail delay prediction process. Furthermore, by selecting successful path adjustment samples, path backoff samples, and prediction error exceeding the threshold samples from the path control feedback samples, and updating the prediction model when the model update trigger condition is met, the prediction model can continuously absorb path control feedback, correct tail delay prediction deviations under different training semantic contexts and network states, improve the accuracy of subsequent tail delay predictions and the reliability of path adjustment decisions, thereby enhancing the long-term adaptability and stability of the network closed-loop control.

[0132] In summary, this embodiment provides a network closed-loop control method based on tail delay prediction. By acquiring the flow identification information and training semantic context of the target set communication flow during distributed synchronous training of artificial intelligence, and combining the telemetry data of the current path and candidate paths, the tail delay indicators of the current path and each candidate path within a preset prediction window are predicted. This transforms the path control basis from "judgment of the current state" to "judgment of tail delay risk within the prediction window," thereby enabling earlier identification of tail delay degradation risks in synchronous training set communication. Furthermore, this application uses the tail delay indicator of the current path as a comparison benchmark, and evaluates and determines the target path by combining the tail delay indicators of candidate paths and the path adjustment cost. This ensures that path selection no longer depends solely on the instantaneous congestion state or the current delay magnitude, but comprehensively considers the tail delay performance of candidate paths in subsequent communication windows and the path adjustment cost, thereby improving the foresight and effectiveness of path adjustment decisions. Simultaneously, this application only performs path adjustment when the tail delay indicator of the current path meets the preset path control conditions and does not hit the preset suppression conditions, which can reduce false triggering, frequent switching, and path oscillation caused by short-term fluctuations or single threshold triggering. After path adjustment, the target path is maintained or rolled back based on operational feedback information, and the prediction process of tail delay index is updated. This allows for timely rollback when the target path performance is substandard or the adjustment fails, shortening the impact time on the training process. Furthermore, the prediction process is continuously corrected through feedback updates, thereby improving the accuracy of tail delay prediction and the stability of network closed-loop control in long-term operation.

[0133] In one specific embodiment, the network closed-loop control method based on tail latency prediction provided in this application can be deployed in a typical distributed synchronous training cluster for artificial intelligence. The training cluster can consist of 1024 NVIDIA A100 GPUs, each configured with 80GB of HBM2e memory. The training cluster uses eight switches to form a two-layer Fat-Tree topology, including four leaf switches and four spine switches. Two 400GbE links are configured between each pair of leaf switches and spine switches as primary / backup paths or equivalent candidate paths, thus forming at least two selectable transmission paths between the source and destination training nodes of the target set communication flow. Intra-node communication between GPUs can be achieved through NVLink and NVSwitch, while cross-node communication can be achieved via an InfiniBand HDR network or a RoCE v2-based 400GbE network. Furthermore, the DCQCN congestion control mechanism can be combined to control congestion in the data center network.

[0134] In this embodiment, the training framework can use PyTorch 2.x and DeepSpeed ​​0.14, with a data parallelism of 1024. The ZeRO Stage 2 optimizer state partitioning method is used to execute the large model training task. The training task can be a full parameter fine-tuning task of the LLaMA-7B model, with a single batch gradient synchronization communication volume of approximately 1.2GB, a data format of FP16, and an AllReduce communication cycle of approximately 350ms. When the data parallel group enters the gradient synchronization phase, the network control system can identify the upcoming large-scale AllReduce communication and determine the communication stream generated by this AllReduce communication as the target set communication stream. Because the target set communication stream has a large data volume, a large number of participants in the ranking, and the training task needs to wait for the AllReduce communication to complete before entering the subsequent training phase, the P99 tail latency of the target set communication stream can directly affect the training synchronization waiting time and the training step throughput.

[0135] In this embodiment, to obtain the training semantic context corresponding to the target set communication flow, the network control system can set up training semantic acquisition interfaces at the training framework layer and the set communication library layer. Specifically, for the set communication operation type and communication message size, the network control system can intercept set communication calls at the entry points of functions such as allreduce, allgather, reducescatter, and broadcast in the NCCL or RCCL communication library through LD_PRELOAD dynamic hooking, communication library callback registration mechanism, or other user-space interface listening methods, and record the set communication operation type, the number of participating ranks, the data type, and the number of communication message bytes. This training semantic acquisition interface can be located in the user-space communication library layer without modifying the operating system kernel or graphics processor driver.

[0136] For identification during the training phase, key time points such as the start of backpropagation and the start of optimizer updates can be marked by hooking the entry points of PyTorch's `torch.autograd.backward` and `optimizer.step` functions, respectively. When the gradient synchronization process before the `optimizer.step` call is detected, it can be determined that the training task is about to enter or is currently executing the gradient synchronization phase, and the corresponding AllReduce communication operation can be identified. When the next forward call is detected, the current gradient synchronization process can be marked as complete. In this way, the tail delay prediction controller can obtain the type and size of the upcoming collective communication operation before or at the start of AllReduce communication, thus providing pre-training input for subsequent tail delay prediction.

[0137] To obtain the synchronization window number, the current training iteration number can be determined by hooking the iteration counter in the training loop or by reading training task metadata from training scheduling planes such as Kubernetes and MPI Launcher. This synchronization window number can be used to construct historical synchronization identifiers in time series features, enabling the tail latency prediction model to learn the periodic or phased patterns of tail latency changes with training iterations during the training process. For obtaining the parallel group type and topology, the topology information of parallel groups such as data parallelism, tensor parallelism, and pipeline parallelism can be parsed from the parallel configuration files of DeepSpeed ​​or Megatron-LM. This includes the rank list of each parallel group and cross-node communication peer relationships, to determine the cross-node path range of the target set's communication flow and narrow down the candidate path search space.

[0138] In some embodiments, the tail delay prediction controller can encapsulate the training semantic features, such as the type of aggregate communication operation, message size, training stage, synchronization window number, and parallel group topology, into structured metadata and push it to the path controller via gRPC. The path controller can use OpenFlow Group Table, VXLAN NetworkIdentifier, or other flow labeling mechanisms to write the flow labels corresponding to the training semantic context into the matching field of the switch, enabling the switch to perform training semantic-aware flow classification and path policy execution for the target aggregate communication flow. Thus, the network control system can associate the semantic information on the training framework side with the flow matching and path adjustment process on the network control plane.

[0139] In this embodiment, network telemetry data can be collected based on the switch's INT capability. Specifically, a telemetry header can be injected into each hop during packet forwarding via the P4 program. The telemetry header can carry information such as ingress port timestamp, egress port timestamp, queue depth, explicit congestion notification flag count, and port utilization. The telemetry sampling period can be configured to 5ms to obtain multiple sets of path telemetry data samples within the AllReduce synchronization cycle, providing input data for P99 tail delay statistics, trend identification, and prediction of the current path and candidate paths. The aforementioned telemetry data can be aggregated by the telemetry acquisition component and provided to the tail delay prediction controller.

[0140] In this embodiment, the tail latency prediction controller can be deployed on an independent management node, which can be configured with a 64-core CPU, 256GB of memory, and an NVIDIA T4 GPU for inference acceleration. The tail latency prediction controller can communicate with switches, telemetry acquisition components, or path controllers via the gRPC / protobuf protocol to receive training semantic context, path telemetry data, and path control feedback information, and predict the P99 tail latency of the current path and candidate paths within a preset prediction window. The control plane can use the ONOS open-source SDN controller as the path controller, supporting OpenFlow 1.5 flow table distribution and path switching. When the tail latency prediction controller determines that the target set communication flow needs to be adjusted from the current path to the target path, it can distribute path adjustment rules to the ONOS controller. The ONOS controller generates corresponding OpenFlow flow table entries based on these rules and distributes them to the corresponding switches to achieve path switching of the target set communication flow.

[0141] In the specific environment described above, when a data parallel group enters the gradient synchronization phase, the tail delay prediction controller can identify the upcoming large-scale AllReduce communication and determine the communication flow generated by this AllReduce communication as the target set communication flow. Subsequently, the tail delay prediction controller can collect the path telemetry sequence of the target set communication flow within the last 200ms and perform tail delay prediction in conjunction with the corresponding training semantic context. The training semantic context may include information such as the AllReduce communication operation type, communication message size, data parallel group identifier, number of participating ranks, training stage, and synchronization window number. In one example, the tail delay prediction controller can predict that the P99 tail delay of the current path will exceed a preset baseline threshold within the next second, for example, reaching 1.35 times the baseline tail delay; at the same time, the predicted P99 tail delay value of the candidate path within the same prediction window is lower than that of the current path, and the expected communication gain of the candidate path relative to the current path exceeds a preset gain threshold. At this point, the tail delay prediction controller can further determine the comprehensive path cost based on the P99 tail delay prediction value, path switching cost, and disturbance cost of the candidate path. After confirming the suppression conditions such as missing the cooling window, critical synchronization stage, insufficient benefits, or historical failure, it triggers predictive path adjustment to switch the target set communication flow from the current path to the target path corresponding to the candidate path.

[0142] After path adjustment is completed, the tail delay prediction controller can enter a 45-second cooling window. During this window, it performs an active probe along the target path every 2 seconds to obtain operational feedback information such as the probe success rate, measured latency, and packet loss rate. If the probe results within the cooling window are all normal, and the measured latency and packet loss rate of the target path do not exceed the corresponding thresholds, the target path can be determined to meet the stability maintenance condition. The target set communication flow continues to be transmitted via the target path, and this path adjustment is recorded as a successful path adjustment sample for subsequent updates to the tail delay prediction model. If continuous probe failures, measured latency exceeding the threshold, or packet loss rate exceeding the threshold occur within the cooling window, the tail delay prediction controller can use the path controller to revert the target set communication flow to the current path before adjustment, and record the reason for the reversion and prediction error information.

[0143] In some embodiments, the network control system can also generate a unified tracking identifier for the same path closed-loop control process, and record events at each stage of the path control process based on this unified tracking identifier. Specifically, the network control system can sequentially record events such as target set communication flow identification, path telemetry data acquisition, tail delay prediction and trigger determination, candidate path evaluation and target path determination, path adjustment rule issuance, cooling window activation, target path active detection, path hold or rollback, and model update sample writing. These events can be stored in the form of structured logs, event log tables, or feedback sample records, and are linked throughout the same path control closed loop using a unified tracking identifier, enabling end-to-end retrieval, process backtracking, and root cause analysis of a path adjustment decision in the production environment.

[0144] In a specific implementation, after the target set of communication flows is identified, the network control system can record the flow identification information, set communication operation type, communication message size, training stage, and synchronization window number of the target set of communication flows; after completing telemetry acquisition, it can record the latency statistics, queue depth, explicit congestion notification marking ratio, packet loss rate, and path stability information of the current path and candidate paths; after the tail delay prediction model outputs the prediction results, it can record the P99 tail delay prediction values ​​of the current path and candidate paths within the preset prediction window, whether the preset path control conditions are met, and whether the preset suppression conditions are hit. Control conditions: After the candidate path evaluation is completed, the comprehensive path cost, expected communication benefit, and candidate path determined as the target path can be recorded for each candidate path; after the path controller performs path adjustment, the path adjustment rules, path before adjustment, target path, and path adjustment results can be recorded; after entering the cooling window, the start and end time of the cooling window, the detection cycle, and the measured delay, packet loss status, and detection success status obtained from each active detection can be recorded; after the path hold or rollback control is completed, the path adjustment results, rollback reasons, prediction error information, and whether they are written into the model update samples can be recorded.

[0145] By using the aforementioned unified tracking identifier-based association recording method, the stages of target set communication flow identification, telemetry acquisition, tail delay prediction, triggering and suppression determination, candidate path evaluation, path adjustment execution, cooling detection verification, and prediction model update can be linked into a traceable closed-loop control link. On the one hand, it can provide a complete data source for the generation of path control feedback samples, giving successful path adjustment samples, path rollback samples, and prediction error samples clear contextual basis. On the other hand, it can also, when the path adjustment effect does not meet expectations, trace back the prediction input, path evaluation results, control execution results, and operational feedback information before and after the path adjustment based on the unified tracking identifier, thereby improving the efficiency of anomaly localization and subsequent model correction.

[0146] In a validation example, the LLaMA-7B model underwent 1000 consecutive training steps for full parameter fine-tuning on the aforementioned 1024-card training cluster. During the test, the network control system predicted the tail delay of the target set communication flow based on the training semantic context and path telemetry data, and performed predictive path adjustments when the path control conditions were met but the suppression conditions were not met. During the test, the system triggered multiple predictive path adjustments, most of which were confirmed as stable adjustments through active probing within the cooling window. A small number of path adjustments triggered backoff control due to abnormal target path probing or sudden increases in measured delay. Simultaneously, the system was able to suppress path adjustments in cases where the predicted value briefly exceeded the threshold but the growth trend was insufficient, the gain was insufficient, or the suppression conditions were met, thereby avoiding unnecessary disturbances to the training process.

[0147] During a typical path adjustment process, before the adjustment, the high-quantile latency of the current path within the preset observation window before triggering the adjustment showed an upward trend. The P50 latency, P95 latency, and P99 latency were all higher than the normal operating state, with the P99 latency continuously rising from a low level to around 325µs. Simultaneously, the proportion of explicit congestion notification markers exceeded the preset threshold, the queue depth was higher than the historical average, the AllReduce single-run completion time deteriorated to around 387ms, and the training step throughput decreased to around 112.3 TFLOPS. These operational conditions indicate that the current path has already shown a risk of tail latency degradation that affects the synchronization efficiency of the target set communication flow.

[0148] After path adjustment, the target set communication flow is redirected to the target path, and active probing continues within a cooling window. During path adjustment, the network control system can simultaneously record the control plane latency and short-term data plane impacts caused by the path switch. Control plane latency includes the time required for the path controller to issue path adjustment rules and for the switch to refresh corresponding forwarding rules; data plane interruption time includes the short-term impacts introduced by connection state transitions, forwarding rule activation, and probe confirmation during the target set communication flow's switch from the current path to the target path. In one example, the path switch control plane latency is approximately 8ms, the data plane interruption time is approximately 45ms, and the total response time from prediction trigger to path switch completion is approximately 18ms. After path switch completion, the system enters a 45-second cooling window and continuously performs active probing within this window to verify the actual operating status of the target path.

[0149] After verification via the cooling window, the P50, P95, and P99 latencies of the target path were all lower than before path adjustment, with the P99 latency decreasing to approximately 248µs. The explicit congestion notification marking ratio decreased, the queue depth decreased, the single-run completion time of AllReduce was shortened to approximately 328ms, and the training step throughput recovered to approximately 128.1 TFLOPS. This demonstrates that after the target path actually carries the target set communication flow, it can reduce high-quantile latency and improve the communication completion time of the gradient synchronization phase.

[0150] In further continuous training tests, the LLaMA-7B model underwent 1000 consecutive training steps for full parameter fine-tuning on the aforementioned 1024-card training cluster. During the test, the network control system triggered 12 predictive path adjustments, of which 11 were successful (91.7% success rate) and 1 was rolled back (8.3% rollback rate). The rollback was caused by a sudden increase in measured latency during the active detection of the target path; the network control system promptly rolled back the path after confirming the abnormality. During the test, the network control system also suppressed 7 false triggers of path adjustments, mainly corresponding to situations where the predicted value briefly exceeded the threshold but was blocked by preset suppression conditions.

[0151] In terms of overall performance, after adopting the closed-loop path control scheme of this application, the average throughput of the training step increased from 118.2 TFLOPS without closed-loop control to 125.8 TFLOPS, an improvement of approximately 6.4%; the standard deviation of the training step throughput decreased from 12.7 TFLOPS to 4.2 TFLOPS, a reduction in fluctuation of approximately 66.9%; and the average P99 AllReduce latency decreased from 342 ms to 285 ms, a reduction of approximately 16.7%. These results demonstrate that this application can reduce the P99 tail latency of the target set communication stream during continuous training, reduce training throughput fluctuations, and improve the communication stability of distributed synchronous training tasks in artificial intelligence.

[0152] In a comparative verification example, the proposed scheme can be compared with different path control methods. Comparative scheme A uses a flow size-based classification and path allocation method, identifying AllReduce communication flows as large flows based on data volume or duration and assigning them high-throughput paths. While this method can improve the bandwidth availability of the target set communication flows to some extent, its path control is primarily based on flow size or general traffic characteristics. It does not incorporate training semantic context such as the training phase, set communication operation type, communication message size, and synchronization window number to predict future tail latency degradation risks, nor does it form a closed-loop control after path adjustment through cooling windows, active probing, backoff control, and feedback sample updates. Therefore, when the current path tail latency shows a trend of degradation but has not yet manifested as a significant decrease in throughput, this comparative scheme struggles to promptly implement predictive path adjustments for P99 tail latency risks.

[0153] Comparative scheme B employs a general probabilistic predictive routing approach, which predicts service quality risk or the probability of service quality violation based on network conditions and selects paths accordingly. While this approach can utilize network conditions for a certain degree of predictive routing, its predictions primarily focus on service quality risks of general business flows. It fails to introduce training semantic context for the aggregate communication flows in distributed synchronous training of artificial intelligence, and it does not consider factors such as the AllReduce communication cycle, training stage, communication message size, parallel group topology, or synchronization window identifier as inputs for tail latency prediction. Therefore, this comparative scheme cannot accurately reflect the sensitivity of synchronous training aggregate communication to P99 tail latency and slow node waiting.

[0154] Under the same training environment and testing task, compared with the aforementioned comparative scheme, the scheme of this application, by jointly inputting the training semantic context and path telemetry data into the tail latency prediction process, can identify the risk of P99 tail latency degradation in the AllReduce ensemble communication flow within the future prediction window earlier; by comprehensively considering path cost and expected communication benefits for target path selection, it can avoid selecting paths solely based on high-throughput paths or general service quality risks; and by using preset path control conditions, preset suppression conditions, and cooldown detection and rollback control, it can reduce the continuous impact of erroneous switching, insufficient benefit switching, and path adjustment failures on the training process. Therefore, the scheme of this embodiment can achieve control effects more suitable for distributed synchronous training ensemble communication scenarios in artificial intelligence in terms of P99 tail latency control, training throughput fluctuation suppression, and stability assurance after path adjustment.

[0155] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this application.

[0156] Furthermore, some embodiments of this application also provide an electronic device. The electronic device can be various forms of digital computer, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device can also be various forms of mobile devices, such as cellular phones, smartphones, wearable devices, and other similar computing devices.

[0157] The electronic device includes: one or more processors; and a memory storing computer program instructions, which, when executed, cause the processor to perform a network closed-loop control method based on tail delay prediction as provided in any one or more of the above embodiments. Figure 5 An exemplary structural diagram of the electronic device is disclosed. The electronic device includes one or more processors 1101, a memory 1102, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0158] The electronic device may further include an input device 1103 and an output device 1104. The processor 1101, memory 1102, input device 1103, and output device 1104 may be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.

[0159] Input device 1103 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 1104 may include a display device, auxiliary lighting device (e.g., LED), and haptic feedback device (e.g., vibration motor). The display device may include, but is not limited to, a liquid crystal display, a light-emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.

[0160] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device (e.g., a cathode ray tube or LCD monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback); and input from the user can be received in any form (e.g., voice input or tactile input).

[0161] In this embodiment, a computer-readable medium stores a computer program / instruction, which, when executed by a processor, implements a network closed-loop control method based on tail delay prediction provided in any one or more of the above embodiments. The computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable medium carries one or more computer-readable instructions.

[0162] The memory 1102 can serve as a non-transitory computer-readable storage medium, used to store non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, thereby implementing the program instructions / modules corresponding to the methods provided in any one or more of the embodiments described above in this application.

[0163] The memory 1102 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 1102 may optionally include memory remotely located relative to the processor 1101, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0164] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0165] Computer-readable media include permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technologies, read-only optical discs, digital versatile optical discs or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0166] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0167] In the above embodiments, all or part of the implementation can be achieved through software, hardware, firmware, or any combination thereof. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, magnetic or optical drives, floppy disks, and similar devices. In addition, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.

[0168] The computer program product provided in this application includes one or more computer programs / instructions. When executed by a processor, these computer programs / instructions generate, in whole or in part, the processes or functions described in this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0169] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0170] The scope of this application is defined by the appended claims rather than the foregoing description, and is therefore intended to encompass all variations falling within the meaning and scope of equivalents of the claims. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. Terms such as "first," "second," etc., are used only to distinguish descriptions and do not indicate any particular order, nor should they be construed as indicating or implying relative importance.

[0171] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily made by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.

Claims

1. A network closed-loop control method based on tail delay prediction, characterized in that, include: Obtain the stream identifier information of the target set communication stream generated during the distributed synchronous training of artificial intelligence, as well as the training semantic context corresponding to the target set communication stream; Based on the flow identification information, determine the current path and at least one candidate path corresponding to the target set communication flow, and collect telemetry data of the current path and each of the candidate paths; Based on the training semantic context and the telemetry data, predict the tail latency index of the current path and each of the candidate paths within a preset prediction window; Using the tail delay index of the current path as a comparison benchmark, the candidate paths are evaluated based on their tail delay indices and corresponding path adjustment costs, and the target path is determined from among the candidate paths. If the tail delay index of the current path meets the preset path control condition and the preset suppression condition is not hit, the target set communication flow is adjusted from the current path to the target path. The system acquires operational feedback information after path adjustment, and performs maintain or rollback control on the target path based on the operational feedback information, as well as updates the prediction process for the tail delay index.

2. The network closed-loop control method based on tail delay prediction according to claim 1, characterized in that, The step of evaluating each candidate path based on its tail latency index as a comparison benchmark and the corresponding path adjustment cost, and determining the target path from among the candidate paths, includes: Determine the path switching cost required to adjust from the current path to each of the candidate paths, and the disturbance cost of each candidate path to other concurrent aggregated communication flows; Based on the tail delay index of each candidate path, the path switching cost, and the disturbance cost, the comprehensive path cost of each candidate path is determined. Based on the difference between the tail latency metric of the current path and the comprehensive path cost of each candidate path, the expected communication benefit of each candidate path relative to the current path is determined. The candidate path that has the expected communication benefits exceeding a preset switching benefit threshold and has the lowest overall path cost is determined as the target path.

3. The network closed-loop control method based on tail delay prediction according to claim 1, characterized in that, The tail delay index of the current path is determined to meet the preset path control condition if at least two of the following conditions are met: The tail delay index of the current path exceeds the preset baseline threshold within the preset prediction window; The growth rate of the tail delay index of the current path within the preset prediction window exceeds the preset growth threshold. The current path exceeds the baseline threshold in the tail delay index of multiple consecutive preset prediction windows.

4. The network closed-loop control method based on tail delay prediction according to claim 1, characterized in that, The preset suppression condition includes at least one of the following: The target set communication flow is within the cooling window after path adjustment; The artificial intelligence distributed synchronous training process is in a preset key synchronization phase; The expected communication benefit of the target path relative to the current path is lower than a preset benefit guarantee threshold; The most recent path adjustment attempt for the target set communication flow failed.

5. The network closed-loop control method based on tail delay prediction according to claim 1, characterized in that, The step of obtaining the operational feedback information after path adjustment and controlling the target path to be maintained or rolled back based on the operational feedback information includes: After adjusting the target set communication flow from the current path to the target path, a cooling window is entered; Within the cooling window, probe messages are sent along the target path according to a preset probe cycle to obtain operational feedback information including probe success rate, measured latency and packet loss rate; Based on the operational feedback information, determine whether the target path meets the stability maintenance condition or the rollback condition; If the detection success rate meets the preset detection success condition, and the measured delay and the packet loss rate do not exceed the corresponding threshold, then the target path is determined to meet the stability maintenance condition, and the target set communication flow is maintained to be transmitted through the target path; If the detection fails for a preset number of consecutive times, or if at least one of the measured latency and the packet loss rate exceeds the corresponding threshold, then the target path is determined to meet the rollback condition, and the target set communication flow is rolled back from the target path to the current path before adjustment.

6. The network closed-loop control method based on tail delay prediction according to claim 1, characterized in that, The step of updating the prediction process for the tail delay index includes: The target set communication flow is obtained as follows: tail delay index before path adjustment, measured tail delay after path adjustment, path adjustment result, operation feedback information and backoff reason, and path control feedback sample is generated. Based on the deviation between the tail delay index and the measured tail delay, the prediction error information corresponding to the path control feedback sample is determined; Model update samples are selected from the path control feedback samples, wherein the model update samples include: successful path adjustment samples, path rollback samples that trigger rollback control, and samples whose prediction error information exceeds a preset error threshold. When the model update triggering conditions are met, the prediction model used to predict the tail delay index is updated based on the model update samples; wherein, the model update triggering conditions include at least one of the following: the number of model update samples reaches a preset number threshold, the time since the last model update reaches a preset update time period, or the proportion of samples whose prediction error information exceeds a preset error threshold reaches a preset proportion threshold.

7. The network closed-loop control method based on tail delay prediction according to claim 1, characterized in that, The training semantic context includes: the current training stage of the distributed synchronous training process of artificial intelligence, the communication primitive type corresponding to the target set communication stream, the communication message size corresponding to the target set communication stream, and the synchronization window identifier associated with the target set communication stream.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; and a memory storing computer program instructions, which, when executed, cause the processors to perform the network closed-loop control method based on tail delay prediction as described in any one of claims 1-7.

9. A computer-readable storage medium having a computer program and / or instructions stored thereon, characterized in that, When the computer program and / or instructions are executed by the processor, they implement the network closed-loop control method based on tail delay prediction as described in any one of claims 1-7.

10. A computer program product comprising a computer program and / or instructions, characterized in that, When the computer program and / or instructions are executed by the processor, they implement the network closed-loop control method based on tail delay prediction as described in any one of claims 1-7.